The Viggle-Animate pipeline
A driving video and one of its own frames, repainted in any image editor, are the
model's only two inputs. Geometry enters from the video, appearance from the repainted frame. No
pose skeleton, segmentation mask, face crop, background plate, depth map or text prompt is ever
computed.
driving.mp4
motion, camera, timing
any image editor
gpt-image or similar,
not this model.
ref.png
that frame, repainted
any frame
geometry
appearance
Viggle-Animate
33.1 B ref2va finetune + DMD2 LoRA
3 forward passes · 26 s per clip on one B200
out.mp4
same motion, new character
NEVER COMPUTED pose skeleton segmentation mask face crop depth text prompt