Instructions to use ZhengmingYu/DMAD with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ZhengmingYu/DMAD with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("MiniMaxAI/MiniMax-H3", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("ZhengmingYu/DMAD") prompt = "A man with short gray hair plays a red electric guitar." output = pipe(prompt=prompt).frames[0] export_to_video(output, "output.mp4") - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
4-step MiniMax-H3 students for joint audio-video generation
Zhengming Yu1,2,
Junkun Yuan2,
Haotian Yang2,
Gordon Guocheng Qian2,
Yizhi Wang2,
Angtian Wang2,
Yiding Yang2,
Bo Liu2,
Xin Li1,
Wenping Wang1,
Chongyang Ma2
1Texas A&M University, 2ByteDance
This repository holds the DMAD students of the paper: 4-step students of MiniMax-H3
(33B, text-to-audio-video) and of Wan2.1-T2V (1.3B and 14B), 4- and 1-step students of SDXL, and 1-step
students of the EDM ImageNet-64 teacher. The inference and training code is in the
code repository (train/h3, train/wan, train/image).
MiniMax-H3 (text-to-audio-video)
Rank-128 LoRAs on the H3 transformer that turn the 50-step teacher into a 4-step generator of 1344x768 video with native stereo audio.
| File | Checkpoint | Size |
|---|---|---|
minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors |
the checkpoint of the paper: EMA of the student at iteration 800 of the main run | 1.4 GB |
minimax_h3/dmad_minimax_h3_4step_full_critic.safetensors |
the student of a run whose critic backbone is fully trained (the paper's run keeps it frozen under a LoRA): iteration 1600, live weights; it scores higher on AVGen-Bench | 1.4 GB |
minimax_h3/dmad_minimax_h3_4step_lora_critic_comfyui.safetensors |
lora_critic in ComfyUI's MiniMax-H3 key layout (exact conversion) |
2.0 GB |
minimax_h3/dmad_minimax_h3_4step_full_critic_comfyui.safetensors |
full_critic in ComfyUI's MiniMax-H3 key layout (exact conversion) |
2.0 GB |
LoRA layout of the first two: Diffusers keys (<module>.lora.down.weight = A [128, in], <module>.lora.up.weight = B
[out, 128]) over attn.to_q/to_k/to_v/to_out.0, ff.net.0.proj, ff.net.2 of all 50 transformer blocks and the 2
token-refiner blocks (312 modules). alpha = rank = 128. The safetensors metadata repeats this.
The inference code lives in the code repository: inference.py with the sampler
the paper used (re-noise step rule) and a Diffusers-pipeline example; its README covers the environment. Sampling
settings: 4 steps, time shift 12 (video) and 2 (audio), no classifier-free guidance, 124 frames at 24 fps.
git clone https://github.com/Yzmblog/DMAD.git && cd DMAD # code + environment setup (see its README)
hf download MiniMaxAI/MiniMax-H3 --local-dir models/MiniMax-H3 --exclude "FL2VA/*" --exclude "Ref2VA/*" --exclude "transformer_ref/*"
hf download ZhengmingYu/DMAD --include "minimax_h3/*_critic.safetensors" --local-dir ckpt
python inference.py --model-dir models/MiniMax-H3 --lora ckpt/minimax_h3/dmad_minimax_h3_4step_lora_critic.safetensors \
--prompt-file prompts/dmad_sweater.txt --seed 42 --output-dir outputs/dmad_sweater
ComfyUI
minimax_h3/dmad_minimax_h3_4step_{lora_critic,full_critic}_comfyui.safetensors are the same two LoRAs converted exactly
to ComfyUI's MiniMax-H3 key layout (rank 128; q/k/v fused into attn.qkv_proj adapters of rank 384, alpha = rank,
scale 1.0 β no rank reduction; 2.0 GB each). Load with LoraLoaderModelOnly at strength 1.0, cfg 1.0,
ModelSamplingMiniMaxH3 with shift 12 / audio shift 2, and sample with ComfyUI's lcm sampler and simple scheduler
(the re-noise multistep rule and sigma grid the students were trained with; ODE samplers such as euler are not their
operating point), or with the equivalent DMAD Sampler + DMAD Sigmas nodes from
comfyui/ComfyUI-DMAD.
hf download ZhengmingYu/DMAD --include "minimax_h3/*_comfyui.safetensors" --local-dir /path/to/ComfyUI/models/loras
Wan2.1 (text-to-video)
Full generators (EMA, spectral norm folded in) in the .pth format of
train/wan: 4 steps, 480p, 81 frames, no classifier-free
guidance.
| File | Model | Size |
|---|---|---|
wan2.1/dmad_wan2pt1_1pt3B.pth |
Wan2.1-T2V-1.3B student, iteration 17k | 2.8 GB |
wan2.1/dmad_wan2pt1_14B.pth |
Wan2.1-T2V-14B student, iteration 19.5k | 29 GB |
hf download ZhengmingYu/DMAD --include "wan2.1/*" --local-dir ckpt
bash experiments/dmad/sample.sh 1.3B ckpt/wan2.1/dmad_wan2pt1_1pt3B.pth my_prompts.json outputs/samples # in train/wan
SDXL (text-to-image)
UNet state dicts in fp16 (spectral norm folded in) that load into the standard SDXL pipeline; see
train/image for the sampling code.
| File | Model | Size |
|---|---|---|
sdxl/dmad_sdxl_4step_unet_fp16.bin |
4-step student (backward simulation), iteration 17k | 5.1 GB |
sdxl/dmad_sdxl_1step_unet_fp16.bin |
1-step student (ODE init), iteration 49k | 5.1 GB |
sdxl/dmad_sdxl_1step_frozencritic_unet_fp16.bin |
1-step student (ODE init, frozen critic backbone), iteration 17.5k | 5.1 GB |
import torch
from diffusers import DiffusionPipeline, LCMScheduler, UNet2DConditionModel
from huggingface_hub import hf_hub_download
base_model_id = "stabilityai/stable-diffusion-xl-base-1.0"
unet = UNet2DConditionModel.from_config(base_model_id, subfolder="unet").to("cuda", torch.float16)
unet.load_state_dict(torch.load(hf_hub_download("ZhengmingYu/DMAD", "sdxl/dmad_sdxl_4step_unet_fp16.bin"), map_location="cuda"))
pipe = DiffusionPipeline.from_pretrained(base_model_id, unet=unet, torch_dtype=torch.float16, variant="fp16").to("cuda")
pipe.scheduler = LCMScheduler.from_config(pipe.scheduler.config)
image = pipe(prompt="a photo of a cat", num_inference_steps=4, guidance_scale=0, timesteps=[999, 749, 499, 249]).images[0]
# 1-step models: num_inference_steps=1, timesteps=[399]
ImageNet-64 (class-conditional)
1-step EMA generators of the EDM ImageNet-64 teacher, one folder per critic setting of the paper, in the
checkpoint_model_<iteration>/pytorch_model_ema.bin layout that train/image's evaluation reads directly.
| Folder | Setting | Size |
|---|---|---|
imagenet64/dmad_imagenet_gaproute/ |
teacher-UNet critic + gap routing, iteration 568k | 1.2 GB |
imagenet64/dmad_imagenet_pgcritic/ |
pretrained-feature critic, iteration 108k | 1.2 GB |
imagenet64/dmad_imagenet_frozencritic/ |
frozen teacher-UNet critic + gap routing, iteration 221k | 1.2 GB |
hf download ZhengmingYu/DMAD --include "imagenet64/dmad_imagenet_gaproute/*" --local-dir ckpt
python main/edm/test_folder_edm.py --folder ckpt/imagenet64/dmad_imagenet_gaproute --run_once ... # in train/image
Citation
@misc{yu2026dmad,
title = {DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation},
author = {Zhengming Yu and Junkun Yuan and Haotian Yang and Gordon Guocheng Qian and Yizhi Wang and
Angtian Wang and Yiding Yang and Bo Liu and Xin Li and Wenping Wang and Chongyang Ma},
year = {2026},
eprint = {2610.02188},
archivePrefix = {arXiv}
}
License
Each family of weights is a derivative of its base model and is distributed under that model's license:
minimax_h3/: Model Derivatives of MiniMax H3, under the MiniMax H3 Community License Agreement (see also NOTICE). The weights were modified from MiniMax-H3 by LoRA fine-tuning.wan2.1/: derived from Wan2.1-T2V-1.3B / 14B, under the Apache License 2.0.sdxl/: derived from Stable Diffusion XL base 1.0, under the CreativeML Open RAIL++-M License, including its use-based restrictions.imagenet64/: derived from the NVIDIA EDM ImageNet-64 model, under CC BY-NC-SA 4.0 (non-commercial).
Images and videos produced with these weights are AI-generated.
- Downloads last month
- -
Model tree for ZhengmingYu/DMAD
Base model
MiniMaxAI/MiniMax-H3