YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

VibeVoice Final Samples β€” the whole story in audio

Joint-finetuned VibeVoice-1.5B (LLM-LoRA + full diffusion head, warm-started chain fullv3 -> emo3 -> emo4). All samples generated with our emotion-steering stack.

Listening order

  1. 01_dialogue_baseline_stock β€” stock VibeVoice (the read-only baseline)
  2. 02_dialogue_our_model_fullv3 β€” OUR model, ear-certified intelligible dialogue (WER 0.36 vs stock 0.03; 5 prior head-only attempts produced babble β€” this is the joint-training proof)
  3. 03_dialogue_polished_joint2 β€” the polished twin (more training: a wash statistically, different samples)
  4. 04_emotion_reference_steering_cremad β€” R1: swap the reference clip's EMOTION (same actors, same dialogue): ANGRY_ANGRY / SAD_SAD / ANGRY_SAD_mixed (the per-speaker money clip). Note: CREMA-D actors carry non-US accents β€” the accent transfer is itself evidence of total reference conditioning.
  5. 05_emotion_us_accent_palette β€” R1 US-accent (RAVDESS): same sentence, same voice, 8 emotions (measured F0 tracks emotion: ANGRY 328 Hz ... SAD lowest)
  6. 06_emotion_3speaker_drama β€” 3-speaker scene: angry / fearful / surprised in one conversation
  7. 07_emotion_per_turn_switch β€” same actor goes ANGRY -> relieved mid-dialogue (dual-slot authoring)
  8. 08_ship_config_emo4_20ddpm β€” THE SHIP CONFIG: emo4 (clean-data scale-up) at 20 DDPM steps, tags + matching references. tag2_tags_matching_refs_BEST.wav = the final product demo.
  9. 09_quality_ab β€” the quality diagnosis: the 'webcam-mic static' was REFERENCE-carried (CREMA-D refs: noise floor 0.006 in generation vs 0.0001-0.0004 with clean refs; 20 DDPM steps = 4x high-band)

Ship recipe

emo4 adapter (vibevoice-lora-joint/emo4) | 20 DDPM steps | script tags //... | per-speaker emotional reference clips (clean recordings) | CFG 1.3

Metrics

  • Intelligibility: our model WER 0.32-0.41 band vs stock 0.03 (single-run; run-to-run +-0.1)
  • Speaker identity: SIM 0.92-0.93 vs stock 0.94
  • Emotion: F0 shifts 300->184 Hz (reference steering); tag channel -87 Hz with refs held constant
  • Speed: architecture-bound ~1.4-1.5x slower than stock (launch-bound; measured ceiling)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support