mirror of
https://github.com/storytold/cloud-worker.git
synced 2026-10-09 00:09:43 +00:00
MiniMax H3 on RunPod: install/download scripts, ComfyUI wiring, benchmark harness, research docs
- scripts/pod: weight downloads (workspace + tmpfs overflow for volume quota), extra_model_paths.yaml, headless ComfyUI launcher, test image generator, one-shot pod installer - scripts/local: auto-reconnecting port forward for the ComfyUI panel - bench: API-based harness for t2v/i2v/ref2v across durations, resolutions, and weight families; CSV/JSONL results - docs: WEIGHTS.md (what's downloaded where), RESEARCH.md (ComfyUI guides, no-Comfy options via SGLang/vLLM/diffusers, concurrency model, serverless) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,163 @@
|
||||
# MiniMax H3 — research findings
|
||||
|
||||
Compiled 2026-08-06. Sources linked inline; benchmark numbers we measured
|
||||
ourselves live in `bench/results.csv` and `docs/BENCHMARKS.md`.
|
||||
|
||||
## The model
|
||||
|
||||
[MiniMax H3](https://www.minimax.io/blog/minimax-h3) is an open-weights
|
||||
omni-modal video model: one 33B dense "Omni-Transformer" DiT (50 layers, hidden
|
||||
5376) jointly denoises video + **native stereo audio** latents in a single
|
||||
forward pass, conditioned on hidden states from a **Qwen3-VL-32B** text encoder
|
||||
(layer 50). CFG-distilled — no negative prompt / CFG pass. Output is 24 fps
|
||||
fixed, 4–15 s, native canvas ~1 MP (2K comes from a separate regenerate/upscale
|
||||
module in the hosted API). Video VAE: 16x spatial / 4x temporal. Audio VAE:
|
||||
stereo 32 kHz, 40 latent fps.
|
||||
|
||||
Two task checkpoints, each shipped in several precisions:
|
||||
|
||||
- **fl2va** — text-to-video and first/last-frame conditioning (covers t2v + i2v)
|
||||
- **ref2va** — reference-to-video: up to 9 ref images, 3 ref videos (with
|
||||
soundtracks), 3 standalone audio clips, woven in via `<Picture i>` /
|
||||
`<Video k>` / `<Audio j>` prompt tags
|
||||
|
||||
**License caveat**: the MiniMax H3 Community License restricts open-weight use
|
||||
to the **EU, UK, South Korea, and US**; elsewhere requires a license from
|
||||
MiniMax ([Q&A](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/QA-about-License.md)).
|
||||
|
||||
## Weight families ("bf16 / int8 / int8 pruned")
|
||||
|
||||
- **bf16** — full precision, 66 GB per model line. Reference quality.
|
||||
- **int8_convrot** — 34 GB. [ConvRot](https://arxiv.org/pdf/2512.03673) is
|
||||
rotation-based post-training quantization: a group-wise Hadamard rotation
|
||||
(expressed as a conv, hence the name) redistributes activation outliers so
|
||||
int8 quantizes near-losslessly. Community consensus: "near-lossless".
|
||||
- **pruned_int8_convrot** — 21 GB. "Pruned" here is *not* layer dropping: the
|
||||
DiT's adaLN modulation weights (~40% of params) are replaced by a
|
||||
functionally-equivalent lookup table. Comfy/MiniMax claim no quality loss;
|
||||
HF discussions call the pruned line "a bit more experimental".
|
||||
- (Also upstream: pruned_bf16 40 GB, pruned_fp8_scaled 21 GB, and community
|
||||
INT4/NVFP4/W4A8 quants — see [awesome-minimax-H3](https://github.com/wildminder/awesome-minimax-H3).)
|
||||
|
||||
Full stack VRAM classes (community): bf16 → 80 GB+; int8 → 48 GB; pruned int8 →
|
||||
fits 32 GB cards; INT4-class community quants reach 12–16 GB.
|
||||
|
||||
## ComfyUI integration
|
||||
|
||||
- Support merged day-0: [PR #15224](https://github.com/Comfy-Org/ComfyUI/pull/15224);
|
||||
int8_convrot VAE decode (~1.5x faster) in [PR #15334](https://github.com/Comfy-Org/ComfyUI/pull/15334).
|
||||
- Official tutorial: <https://docs.comfy.org/tutorials/video/minimax/minimax-h3>;
|
||||
blog: <https://blog.comfy.org/p/minimax-h3-day-0-support-in-comfyui>;
|
||||
community wiki: <https://comfyui-wiki.com/en/tutorial/advanced/video/minimax/minimax-h3>.
|
||||
- Nodes (`comfy_extras/nodes_minimax_h3.py`):
|
||||
- `MiniMaxH3ImageToVideo` — **both t2v and i2v**: optional `first_frame` /
|
||||
`last_frame`; no frames connected = pure t2v. Outputs conditioning + packed
|
||||
AV latent.
|
||||
- `MiniMaxH3ReferenceToVideo` — ref2v; `ref_image_size` `match` (fast) vs
|
||||
`max` (2048px refs, several times slower — ref tokens ride every step).
|
||||
- `EmptyMiniMaxH3LatentAV`, `MiniMaxH3SigmaShift` (defaults video 12.0 /
|
||||
audio 3.0 — already model defaults; node exists for overrides).
|
||||
- Template sampling config: `res_multistep` sampler + `simple` scheduler,
|
||||
**20 steps**, denoise 1.0, `BasicGuider` (no CFG), `CreateVideo` fps=24.
|
||||
Wiki: quality degrades <15 steps, marginal gains >25; `beta`/`normal`
|
||||
scheduler reportedly beats `simple` for reference-heavy r2v prompts.
|
||||
- **Frame length rule**: `n % 17 == 5`, i.e. 5, 22, …, 124 (≈5 s), 243 (≈10 s),
|
||||
362 (≈15 s). Trained range 124–362; longer untested. Templates compute
|
||||
`max(5, round(sec*24))` snapped **up** to the grid.
|
||||
- **Resolution rule**: multiples of 32 per axis, native area ≤ 768×1344 ≈ 1 MP.
|
||||
Practical floor ~384p. fps is fixed at 24 (baked into the model).
|
||||
- Known pitfalls (Aug 2026): global `--use-sage-attention` produces noise on H3
|
||||
([#15263](https://github.com/Comfy-Org/ComfyUI/issues/15263)) — use the KJNodes
|
||||
"Patch Sage Attention" node instead (~2x speedup reported); EasyCache degrades
|
||||
the audio stream ([#15326](https://github.com/Comfy-Org/ComfyUI/issues/15326));
|
||||
tiled-VAE and ref-video-encode OOM issues (#15312/#15274/#15246).
|
||||
|
||||
## Concurrency: what actually maximizes throughput
|
||||
|
||||
The finding that shapes everything: **one H3 generation saturates the GPU**
|
||||
(DiT denoise ≈ 88% of request wall time per vLLM's own profiling). Consequences:
|
||||
|
||||
- Running N generations concurrently on one GPU ≈ N× the latency of one — no
|
||||
throughput win (SGLang's 2-outputs-per-prompt numbers scale near-linearly).
|
||||
- Neither SGLang nor vLLM-Omni does cross-request batching for H3
|
||||
("one generation per diffusion batch"), and ComfyUI's H3 latent node doesn't
|
||||
expose batch_size.
|
||||
- ComfyUI executes its queue **serially** — one workflow at a time per instance.
|
||||
|
||||
So "as many concurrent generations as possible" = **one resident worker per
|
||||
GPU, horizontal scale-out**, plus multi-GPU parallelism to cut per-video
|
||||
latency where it matters. On a multi-GPU pod you can run one ComfyUI instance
|
||||
per GPU with `--cuda-device N` pinning and separate ports.
|
||||
|
||||
## Running H3 without ComfyUI
|
||||
|
||||
Official repo [MiniMax-AI/MiniMax-H3](https://github.com/MiniMax-AI/MiniMax-H3)
|
||||
documents four supported paths (original weights:
|
||||
[MiniMaxAI/MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)):
|
||||
|
||||
1. **SGLang — recommended resident server.**
|
||||
[Cookbook](https://docs.sglang.io/cookbook/diffusion/MiniMax/MiniMax-H3).
|
||||
`uv pip install "sglang[diffusion]" --prerelease=allow`, then e.g.
|
||||
`sglang serve --model-path MiniMaxAI/MiniMax-H3 --model-variant fl2va
|
||||
--num-gpus 4 --ulysses-degree 4 --performance-mode speed`.
|
||||
Model stays in VRAM; async REST job API (`POST /v1/videos` → poll →
|
||||
download). Published perf: 4x H200, 1344×768, 124 fr, 50 steps = **75 s**
|
||||
(54 s with Cache-DiT); 4x H100 (TP2+Ulysses2) is the fastest verified combo;
|
||||
online FP8 on B200/B300. Supports `num_outputs_per_prompt`.
|
||||
2. **vLLM-Omni** — [recipe](https://recipes.vllm.ai/MiniMaxAI/MiniMax-H3);
|
||||
sync endpoint `POST /v1/videos/sync`; ROCm docker for MI300X. FP8 not yet.
|
||||
3. **diffusers — merged Aug 5 2026** ([PR #14355](https://github.com/huggingface/diffusers/pull/14355)).
|
||||
Modular Diffusers only (no classic `DiffusionPipeline`):
|
||||
`ModularPipeline.from_pretrained("MiniMaxAI/MiniMax-H3")` →
|
||||
`MiniMaxH3ModularPipeline`; classes `MiniMaxH3Transformer3D`,
|
||||
`AutoencoderKLMiniMaxH3`, `MiniMaxH3Scheduler`. Install from git main until
|
||||
released. Best for custom logic; you manage offload/parallelism yourself.
|
||||
4. **ComfyUI API mode** — what this repo uses today: headless `main.py`,
|
||||
models resident between jobs, `POST /prompt` + `/history` (see
|
||||
`bench/minimax_bench.py` for a zero-dependency client). Serial queue.
|
||||
|
||||
Also: [DiffSynth-Studio](https://github.com/modelscope/DiffSynth-Studio)
|
||||
(supports H3 incl. NF4, low-VRAM Python path). No TensorRT/xDiT/FastVideo
|
||||
support found.
|
||||
|
||||
**Keep-loaded vs serverless**: keeping the model resident is what ComfyUI/SGLang
|
||||
already do. For scale-to-zero economics:
|
||||
|
||||
- **RunPod Serverless has B200** (180 GB, $0.00240/s ≈ $8.64/hr; H200 $0.00155/s,
|
||||
H100 $0.00116/s — [docs](https://docs.runpod.io/serverless/endpoints/endpoint-configurations)).
|
||||
- Weights on a network volume load at ~200–400 MB/s → **cold start ~3–8 min**
|
||||
for 40–115 GB of weights. FlashBoot only helps with steady traffic, not true
|
||||
scale-from-zero ([RunPod blog](https://www.runpod.io/blog/serverless-gpu-cold-starts-flashboot)).
|
||||
- Practical setup: [worker-comfyui](https://github.com/runpod-workers/worker-comfyui)
|
||||
or a custom `runpod` SDK handler wrapping SGLang/diffusers; **min-workers 1**
|
||||
turns it into a resident server with burst capacity — usually cheaper than
|
||||
eating multi-minute cold starts given 1–5 min generation times.
|
||||
- Managed alternative: [fal.ai hosts H3](https://fal.ai/minimax-h3).
|
||||
|
||||
## Published performance reference points
|
||||
|
||||
| Setup | Config | Time |
|
||||
|---|---|---|
|
||||
| 4x H100 80GB (SGLang TP2+Ulysses2) | 1344×768, 124 fr, 50 steps | 13.3 s |
|
||||
| 8x B300 (SGLang, bf16 / online FP8) | same | 19.0 / 18.0 s |
|
||||
| 4x H200 (SGLang Ulysses4) | same | 75.1 s (53.7 s Cache-DiT) |
|
||||
| 2x RTX 5090 (layerwise offload, ~380 GB host RAM) | same | 560 s |
|
||||
| RTX 4090 Laptop 16 GB (SageAttention) | 960×540, 5 s, 20 steps | 182 s |
|
||||
| RTX 3060 12 GB (heavy offload) | 864×480, 124 fr, 20 steps | <9 min |
|
||||
|
||||
Our single-B200 ComfyUI numbers: see `docs/BENCHMARKS.md`.
|
||||
|
||||
## Recommendations for storyteller
|
||||
|
||||
1. **Today (single B200 pod)**: ComfyUI headless + API is fine for dev and
|
||||
already keeps weights resident; use `bench/minimax_bench.py` patterns for
|
||||
programmatic generation. Pruned int8 quality is strong and it's the
|
||||
fastest/smallest; keep bf16 for quality A/Bs.
|
||||
2. **Production serving**: SGLang serve is the purpose-built path (async job
|
||||
API, multi-GPU, Cache-DiT). Evaluate 4x H100/H200 vs 1x B200 per-video
|
||||
economics — multi-GPU Ulysses cuts latency 4–6x for ~4x the GPU cost.
|
||||
3. **Concurrency**: scale horizontally (1 worker/GPU). Don't co-schedule two
|
||||
generations on one GPU.
|
||||
4. **Serverless**: viable on RunPod with B200 + network volume + min-workers 1;
|
||||
pure scale-from-zero pays 3–8 min cold starts. Consider baking pruned-int8
|
||||
weights (~42 GB stack) into the image to cut that.
|
||||
@@ -0,0 +1,57 @@
|
||||
# MiniMax H3 — weights log (B200 pod `c306014998a3`, 216.243.220.136)
|
||||
|
||||
All files from **[Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3)** (ComfyUI repack of
|
||||
[MiniMaxAI/MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)). Two model lines:
|
||||
**fl2va** (text-to-video + first/last-frame i2v) and **ref2va** (reference-to-video).
|
||||
ComfyUI finds both storage locations via `/ComfyUI/extra_model_paths.yaml`
|
||||
(installed from `scripts/pod/extra_model_paths.yaml`).
|
||||
|
||||
## Storage locations
|
||||
|
||||
| Path | Medium | Notes |
|
||||
|---|---|---|
|
||||
| `/workspace/minimax-h3/models/` | RunPod network volume | **Persistent.** Volume has a ~100 GB quota — it is full; don't add files without removing others. |
|
||||
| `/dev/shm/minimax-h3-models/` | tmpfs (RAM, 176 GB) | **EPHEMERAL** — wiped on pod restart. Refill with `bash /workspace/minimax-h3/download-weights-tmpfs.sh`. |
|
||||
|
||||
## Downloaded files
|
||||
|
||||
### Persistent — `/workspace/minimax-h3/models/`
|
||||
|
||||
| File | Family | Size |
|
||||
|---|---|---|
|
||||
| `diffusion_models/minimax_h3_ref2va_pruned_int8_convrot.safetensors` | **int8 pruned** (ref2va) | 21 GB |
|
||||
| `diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors` | **int8 pruned** (fl2va) | 21 GB |
|
||||
| `diffusion_models/minimax_h3_ref2va_int8_convrot.safetensors` | **int8** (ref2va) | 34 GB |
|
||||
| `text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors` | text encoder (Qwen3-VL-32B, nvfp4 AWQ) | 16 GB |
|
||||
| `vae/minimax_h3_video_vae_fp16.safetensors` | video VAE fp16 | 5.2 GB |
|
||||
| `vae/minimax_h3_audio_vae_fp32.safetensors` | audio VAE fp32 | 0.6 GB |
|
||||
|
||||
### Ephemeral (tmpfs) — `/dev/shm/minimax-h3-models/`
|
||||
|
||||
| File | Family | Size |
|
||||
|---|---|---|
|
||||
| `diffusion_models/minimax_h3_fl2va_int8_convrot.safetensors` | **int8** (fl2va) | 34 GB |
|
||||
| `diffusion_models/minimax_h3_ref2va_bf16.safetensors` | **bf16** (ref2va) | 66 GB |
|
||||
| `diffusion_models/minimax_h3_fl2va_bf16.safetensors` | **bf16** (fl2va) | 66 GB |
|
||||
|
||||
The int8 family is split across both locations purely because of the volume
|
||||
quota: ref2va_int8 landed on the volume before it filled, fl2va_int8 overflowed
|
||||
to tmpfs.
|
||||
|
||||
## Not downloaded (exist upstream)
|
||||
|
||||
| File | Size | Why skipped |
|
||||
|---|---|---|
|
||||
| `diffusion_models/minimax_h3_{fl2va,ref2va}_pruned_bf16.safetensors` | 40 GB each | 4th family (pruned but unquantized); not in the 3 families requested |
|
||||
| `diffusion_models/minimax_h3_{fl2va,ref2va}_pruned_fp8_scaled.safetensors` | 21 GB each | 5th family; same size class as pruned int8 |
|
||||
| `text_encoders/qwen3vl_32b_minimax_h3_bf16.safetensors` | 52 GB | nvfp4_awq TE is the one every official template uses |
|
||||
| `text_encoders/qwen3vl_32b_minimax_h3_int8_convrot.safetensors` | 27 GB | same |
|
||||
|
||||
## History / gotchas
|
||||
|
||||
- 2026-08-06: original manual download of ref2va_pruned_int8 had a broken
|
||||
filename (`...safetensors?download=true` from a wget of the HF web URL); byte
|
||||
size matched HF exactly so it was renamed + moved to the volume, not re-downloaded.
|
||||
- 2026-08-06: bf16 + fl2va_int8 downloads to the volume failed with
|
||||
`Disk quota exceeded` at ~93 GB used → tmpfs overflow scheme added.
|
||||
- Downloads log: `/workspace/minimax-h3/download.log`.
|
||||
Reference in New Issue
Block a user