Instructions to use cstr/vibevoice-1.5b-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- VibeVoice
How to use cstr/vibevoice-1.5b-GGUF with VibeVoice:
import torch, soundfile as sf, librosa, numpy as np from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor from vibevoice.modular.modeling_vibevoice_inference import VibeVoiceForConditionalGenerationInference # Load voice sample (should be 24kHz mono) voice, sr = sf.read("path/to/voice_sample.wav") if voice.ndim > 1: voice = voice.mean(axis=1) if sr != 24000: voice = librosa.resample(voice, sr, 24000) processor = VibeVoiceProcessor.from_pretrained("cstr/vibevoice-1.5b-GGUF") model = VibeVoiceForConditionalGenerationInference.from_pretrained( "cstr/vibevoice-1.5b-GGUF", torch_dtype=torch.bfloat16 ).to("cuda").eval() model.set_ddpm_inference_steps(5) inputs = processor(text=["Speaker 0: Hello!\nSpeaker 1: Hi there!"], voice_samples=[[voice]], return_tensors="pt") audio = model.generate(**inputs, cfg_scale=1.3, tokenizer=processor.tokenizer).speech_outputs[0] sf.write("output.wav", audio.cpu().numpy().squeeze(), 24000) - Notebooks
- Google Colab
- Kaggle
Reference wav not supported?
Having trouble getting this to work with a reference wav.
Running the example command:
VIBEVOICE_VOICE_AUDIO=input.wav
./build/bin/crispasr --tts "Hello, how are you today?"
-m vibevoice-1.5b-tts-q4_k.gguf
--tts-output output.wav
crispasr 0.6.0 (git a471a3c6, Release) [backends: cpu]
crispasr: auto-detected backend 'vibevoice' from '/cachecrispasr/vibevoice-1.5b-tts-q4_k.gguf'
vibevoice: d_lm=1536, layers=28, heads=12/2, ffn=8960, vocab=151936
vibevoice: vae_acoustic=64, vae_semantic=128, downsample=3200x
vibevoice: loaded 1204 tensors (backend: CPU)
crispasr[vibevoice-tts]: no voice prompt resolved (pass --voice
When adding the --voice path argument suggested from error output, it still errors out:
gguf_init_from_file_ptr: invalid magic characters: 'RIFF', expected 'GGUF'
vibevoice-voice: failed to load tensor metadata from 'input.wav'
vibevoice: failed to load voice prompt 'input.wav'
crispasr[vibevoice-tts]: voice 'input.wav' could not be loaded; refusing to synthesise without a voice prompt.
crispasr: error: TTS synthesis failed
Seems like the current git version expects precomputed embeddings instead of .wav?
Fixed in commit fbda63b6 (2026-05-01) 'fix(vibevoice-tts): fix 1.5B WAV-clone prompt + normalization + silence trim', shipped in v0.6.2 and later. You're on v0.6.0 (a471a3c6). --voice <path>.wav works on v0.6.3+.