The Viggle-Animate pipeline A driving video and one of its own frames, repainted in any image editor, are the model's only two inputs. Geometry enters from the video, appearance from the repainted frame. No pose skeleton, segmentation mask, face crop, background plate, depth map or text prompt is ever computed. driving.mp4 motion, camera, timing any image editor gpt-image or similar, not this model. ref.png that frame, repainted any frame geometry appearance Viggle-Animate 33.1 B ref2va finetune + DMD2 LoRA 3 forward passes · 26 s per clip on one B200 out.mp4 same motion, new character NEVER COMPUTED pose skeleton segmentation mask face crop depth text prompt