Instructions to use akatz-ai/MiniMax-H3-Person-Remover-LoRA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusion Single File
How to use akatz-ai/MiniMax-H3-Person-Remover-LoRA with Diffusion Single File:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
MiniMax H3 Person Remover LoRA V1
Original → Green mask → Person removed. Eight selected examples, shown with their full frames and original timing. Watch or download the 1080p comparison with original audio.
Remove a selected person from a video and reconstruct the background with MiniMax H3 Ref2VA. The accompanying ComfyUI workflow tracks the person with SAM 3.1, fills the mask green, and generates the replacement background in overlapping windows. Supply the original video and a clean version of its first frame.
Download the LoRA · Download the ComfyUI workflow · Training dataset
This is an experimental adapter trained by Akatz Labs. The showcase contains selected successful results, not a benchmark of all inputs. Background details, camera movement, shadows, and hard cuts can still change.
Quick start
- Use a current ComfyUI with native MiniMax H3 and SAM 3.1 support. Install or update H3 Relay to a version with the Person Remover window preview and reroll controls.
- Put
H3-Person-Remover-V1.safetensorsinComfyUI/models/loras/. - Drag examples/Person-Remover-Window-Reroll-V1.json into ComfyUI. The workflow includes model links, folder locations, and notes beside the relevant nodes.
- Load a 24 fps video with width and height divisible by 32. Start with a short, continuous shot of roughly five seconds.
- Make a clean first frame in an image editor or image-editing model: remove the person while retaining the original size and framing. Load it in the green Load Image node.
- Describe the person in the green CLIP Text Encode node, such as
man in gray shirt. Check the saved green-mask preview before running the full removal. - Run the workflow. Inspect the window previews and reroll a weak section as needed. See the workflow guide for preview, reroll, and audio instructions.
The release workflow starts with no saved window locks or cache references. Its layout, connections, and sampling settings follow the supplied Window Reroll V1 workflow.
Required models
Base models are separate downloads. Select the matching filenames in the workflow loaders.
| Model | ComfyUI folder |
|---|---|
| H3 Ref2VA pruned INT8 ConvRot | models/diffusion_models/ |
| Qwen3VL 32B NVFP4 AWQ text encoder | models/text_encoders/ |
| H3 video VAE FP16 | models/vae/ |
| H3 audio VAE FP32 | models/vae/ |
| SAM 3.1 multiplex FP16, including its text encoder | models/checkpoints/ |
| Person Remover V1 LoRA | models/loras/ |
The supplied workflow selects comfy kitchen attention and pins the H3 text encoder to gpu:0 through Select CLIP Device. These options require a compatible ComfyUI/runtime installation. Prior workflow validation used an RTX 4090, ComfyUI 0.37.0, frontend 1.53.6, and PyTorch 2.14.0+cu130; this is a tested configuration, not a minimum hardware specification.
Starting settings
| Setting | Value |
|---|---|
| LoRA strength | 1.0 |
| Window length | 22 frames |
| Sampling steps | 12 |
| Sampler / scheduler | er_sde / simple |
| CFG | 1 |
| Video / audio sigma shift | 12 / 3 |
| Seed | 904234 |
| Mask expansion | 5 pixels |
| Frame rate | 24 fps |
No Turbo or VFX LoRA is required. The workflow uses H3 frame lengths 17n + 5: 22, 39, 56, 73, and so on. A longer window uses more memory and is not necessarily better. Start with 22.
Each continuation uses the generated boundary frame as its next reference and carries 18 frames of generated video/audio history. The relay removes overlap and trims the final result to the source frame count. The adapter itself does not segment people or create the initial clean reference; SAM and the external first-frame edit provide those inputs.
Default removal prompt:
Remove the green-masked person and reconstruct the background. Preserve the rest of the video, including its camera motion and frame timing.
The workflow's clean output is silent by default. Connect the original audio from Get Video Components to the final Create Video node to retain the soundtrack. The showcase MP4 uses this original-audio approach in postprocessing; it does not demonstrate generated audio fidelity or voice removal.
Training record
V1 was trained for 2,000 optimizer steps, including regularization updates.
| Setting | Recorded value |
|---|---|
| Architecture | minimax_h3_ref2va, pruned Ref2VA |
| Total optimizer updates | 2,000, including regularization updates |
| Rank / alpha | 16 / 16; exclude adaln_proj |
| Tensor format | BF16, 416 tensors across 208 adapter targets |
| Training base | Pruned BF16 Ref2VA, with the frozen v2 training assistant |
| Optimizer / learning rate | AdamW8bit / 5e-5 |
| Batch / accumulation | 1 / 1 |
| Text encoder | NVFP4 Qwen3VL |
| Task / regularization resolution budget | 1152 / 256, using aspect-ratio buckets |
| Task frame lengths | 5, 56, 73, 90, 124 at 24 fps |
| Regularization length | 107 frames at 24 fps, with audio |
| Runtime | AI Toolkit 0.13.23 with the recorded aligned-video-guide patch |
The first 250 updates used static five-frame pairs and preservation regularization. Training then continued to 2,000 steps with balanced task-duration scheduling and resumed optimizer state. The mixed phase's available training pool contains 128 static pairs, 80 moving-mask excerpts from 20 scenes, and six preservation clips. Related excerpts are not independent scenes, and pool size is not update frequency.
The dataset card explains the construction, split, and review status. training/ contains the recorded configurations, schedules, model hashes, and runtime patch. The base model and frozen assistant are not merged into or included in this LoRA. Training used BF16 base weights; the workflow's INT8 base is an inference choice.
- File size: 155,110,584 bytes.
- SHA-256:
b01dc9fe888c2f4e455a8eb3b4878dfd4038ee0e2320a7b186c7935e65c93ba5.
Limitations
- H3 regenerates the full scene. Unmasked areas are not locked to their original pixels, even though training pairs preserve pixels outside their masks.
- A missed hand, hair edge, shadow, reflection, or occluded body part can remain. A larger mask asks the model to invent more background.
- A poor clean first frame can propagate through later windows. Strong camera motion and hard cuts can break continuity.
- Rerolling a window changes its continuation. Inspect the rebuilt suffix as well as the rerolled window.
- Technical checks and successful training do not equal final human approval of every dataset example. The recorded final dataset-acceptance flag remains false.
License and attribution
Powered by MiniMax H3. This is an Akatz Labs modification, not an official MiniMax release. The model is distributed under the MiniMax H3 Community License Agreement; see NOTICE. Its territorial, use, redistribution, and commercial provisions apply. The standard territorial grant excludes the US, EU, UK, and Republic of Korea and describes separate authorization. This repository does not expand those terms.
H3 Relay is a separate GPL-3.0 project. Its code license does not relicense these model weights. Dataset component terms are documented separately in the dataset repository.
- Downloads last month
- 554
