- bench/failures.py: per-GPU success rates and error-class breakdown by config - bench/validate_outputs.py: ffprobe every recorded output via /view - finding: 24GB cards fail on multi-ref (TE vision OOM), not duration/size; 100% of successful runs produce valid videos - outputs now written to /workspace (survives pod resets) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
8.5 KiB
MiniMax H3 — measured benchmarks
Measured 2026-08-06 with bench/minimax_bench.py through the ComfyUI API.
Raw data: bench/results.jsonl (B200), bench/4090/, bench/4000/.
Regenerate the tables with python3 bench/summarize.py.
Method. ComfyUI 0.30 headless; template-reference sampling everywhere:
res_multistep sampler, simple scheduler, 20 steps, no CFG, 24 fps,
nvfp4_awq Qwen3-VL-32B text encoder, fp16 video VAE + fp32 audio VAE.
Times are ComfyUI execution time (execution_start→execution_success), which
excludes client/queue overhead (wall time tracked separately in the raw data;
typically +1–2 s). "Warm" = model already resident; first-load runs are
recorded but excluded from the means. Repeated configs use distinct seeds so
ComfyUI's node cache can't short-circuit. Reference inputs are real images
(film stills, anime characters, forests/volcanos, up to 23040×3840 —
scripts/pod/fetch-bench-images-v2.py).
Hardware: B200 180 GB (US, $6.94/hr) · RTX 4090 24 GB (EU, $0.74/hr) ·
RTX PRO 4000 Blackwell 24 GB (EU, $0.57/hr). The EU cards run pruned int8
only (bf16/int8 don't fit 24 GB usefully) with --reserve-vram 2.5; the RTX
4000 additionally needs --cache-none --disable-pinned-memory because its
container has a 29 GB RAM cap that otherwise OOM-kills weight staging.
B200: weight-family comparison (warm, 20 steps)
The headline: bf16 is the fastest family on the B200 — quantization costs speed here (dequant overhead) and only pays when VRAM is scarce.
| Config | bf16 | int8 | pruned int8 | bf16 VRAM | int8 VRAM | pruned VRAM |
|---|---|---|---|---|---|---|
| t2v 864×480 5 s | 51 s | 63 s | 63 s | 159 GB | 63 GB | 69 GB |
| t2v 864×480 15 s | 240 s | 270 s | 270 s | 182 GB | 101 GB | 75 GB |
| t2v 1344×768 5 s | 169 s | 197 s | 194 s | 177 GB | 99 GB | 73 GB |
| ref2v 1 ref 864×480 5 s | 54 s | 68 s | ~68 s¹ | 186 GB | 95 GB | 68 GB |
| ref2v 4 refs match | 63 s | 86 s | 81 s | 154 GB | 95 GB | 70 GB |
| ref2v 8 refs match | 79 s | 99 s | 95 s | 185 GB | 96 GB | 68 GB |
| ref2v 8 refs match 1344×768 | 282 s | 335 s | 339 s | 158 GB | 103 GB | 77 GB |
| ref2v 4 refs max | 91 s | 117 s | 112 s | 156 GB | 96 GB | 70 GB |
| ref2v 8 refs max | 785 s | 891 s | 841 s | 177 GB | 112 GB | 92 GB |
¹ pruned 1-ref mean is 90±41 s over 3 runs because the first run absorbed a model swap; the steady-state runs are ~68 s, matching int8.
bf16 peaks at 155–187 GB — it only runs on ≥180 GB cards. int8 fits under 112 GB (H200-class); pruned int8 under ~95 GB worst-case, ~70 GB typical (H100-class without offload, 24 GB cards with offload).
Scaling behavior (B200, pruned int8)
Duration (864×480 t2v): 5 s → 63 s, 10 s → 154 s, 15 s → 270 s. Cost per output-second rises from 12.6→15.4→18.0 s — mildly superlinear (attention).
Resolution (5 s t2v): 608×352 → 32 s, 864×480 → 63 s, 1344×768 → 194 s. Roughly ∝ pixels^1.2. The corner case 1344×768×15 s = 1108 s (18.5 min).
Reference count (864×480 5 s, match sizing): 1 ref ≈ 68 s, 4 refs ≈ 81 s,
8 refs ≈ 95 s — ~+4 s per extra reference image. Cheap.
Reference sizing max is the expensive lever: with 8 large refs it goes
95 s → 841 s (~9×), because 2048px reference tokens ride through every
sampling step. Use match unless identity fidelity demands otherwise.
Cross-GPU (pruned int8, t2v 864×480 5 s, warm)
| GPU | $/hr | time/video | videos/hr | $/video | notes |
|---|---|---|---|---|---|
| B200 (bf16) | $6.94 | 51 s | 70 | $0.098 | fastest family on this card |
| B200 (pruned) | $6.94 | 63 s | 57 | $0.121 | |
| RTX 4090 | $0.74 | ~222 s | ~16 | ~$0.046 | suite still running — interim |
| RTX PRO 4000 | $0.57 | pending | pending | pending | suite still running |
(EU numbers will be finalized when their suites complete; ref2v numbers for both cards land then too.)
Concurrency (2 ComfyUI instances, one B200)
See docs/CONCURRENCY.md for the model. Experiment: two resident instances,
sequential pair vs concurrent pair of identical warm t2v jobs.
Measured (pruned int8, t2v 864×480 5 s, both instances warm):
| Mode | Total for 2 videos | Per-video effective |
|---|---|---|
| Sequential (one instance) | 128 s | 64 s |
| Concurrent (two instances) | 145 s | 72.5 s |
Concurrent was 13% slower for the same work — each co-scheduled job ran 142 s vs 63 s alone (2.27×). There is no pipeline-overlap win; cache and scheduler contention make co-scheduling strictly worse. One worker per GPU, queue in front, scale horizontally. (This also removes any argument for MPS/MIG partitioning for throughput.)
Reliability & failure statistics
Every run's end state and error is recorded in the JSONL files; regenerate the
failure tables with python3 bench/failures.py, and validate output files
with python3 bench/validate_outputs.py (downloads each MP4 and ffprobes
streams + duration). Snapshot as of 2026-08-06 ~09:00 UTC:
| GPU | success rate | dominant failure |
|---|---|---|
| B200 180 GB | 96/96 (100%) | none |
| RTX 4090 24 GB | 8/42 (19%) | VRAM OOM in text-encoder vision path on ref2v |
| RTX PRO 4000 24 GB | 6/7 so far | one container-RAM OOM kill before mitigation flags |
Output validation: 100% of successful runs produced valid videos (h264 +
stereo AAC at exactly the requested duration) — every success on all three
GPUs checks out. (Three early B200 outputs 404 on re-download because they
were written to the container disk that the pod reset wiped — outputs now go
to the persistent volume via --output-directory.)
The 24 GB failure pattern is reference images, not duration/resolution.
t2v succeeds reliably on both consumer cards at every size tried; ref2v OOMs
in torch._int_mm during Qwen3-VL image encoding: the 15 GB text encoder plus
multi-image vision activations exceeds 24 GB (observed 22.1 GB peak at OOM).
With aggressive flags (--cache-none --disable-pinned-memory --reserve-vram 2.5) a single reference works (4090: 392 s); 4-ref and 8-ref
configurations fail consistently, and repeated OOMs eventually killed the
server process, cascading failures across the rest of that suite (harness now
survives this and records server_died).
A second environment trap on cheap tiers: the RTX 4000 pod's container has a
29 GB RAM cap that OOM-killed weight staging until --cache-none --disable-pinned-memory was applied; both EU pods also shipped torch 2.4.1,
too old for H3 (needs ≥2.5; Blackwell needs ≥2.7+cu128).
Fleet implication: consumer 24 GB cards are viable for t2v/i2v and single-ref ref2v only. Multi-reference work — the priority modality — needs ≥32 GB (5090-class, unverified) and realistically ≥48 GB for headroom, or it stays on big-GPU pools.
Takeaways
- On the B200, run bf16. Fastest and highest quality; quantized families exist for smaller cards, not for speed.
- Many references are cheap; big references are not. 8 refs at
matchcosts ~1.4× a single ref.maxsizing costs up to ~9×. Defaultmatch. - Duration and resolution both scale superlinearly; the trained envelope corner (15 s @ 1344×768) costs ~18 min even on a B200 — treat full-res long clips as premium jobs.
- 24 GB consumer cards work but pay a heavy offload tax (~3.5× B200
latency rather than the ~3× raw-compute ratio), and cheap pod tiers bring
operational traps: container RAM caps (29 GB on the RTX 4000 pod) and
stock torch builds too old for H3 (needs ≥2.5 for
enable_gqa; Blackwell needs ≥2.7+cu128).