mirror of
https://github.com/storytold/cloud-worker.git
synced 2026-10-09 00:09:43 +00:00
BENCHMARKS.md: B200-complete results — bf16 fastest family, ref scaling, cross-GPU interim
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,95 @@
|
||||
# MiniMax H3 — measured benchmarks
|
||||
|
||||
Measured 2026-08-06 with `bench/minimax_bench.py` through the ComfyUI API.
|
||||
Raw data: `bench/results.jsonl` (B200), `bench/4090/`, `bench/4000/`.
|
||||
Regenerate the tables with `python3 bench/summarize.py`.
|
||||
|
||||
**Method.** ComfyUI 0.30 headless; template-reference sampling everywhere:
|
||||
`res_multistep` sampler, `simple` scheduler, **20 steps**, no CFG, 24 fps,
|
||||
`nvfp4_awq` Qwen3-VL-32B text encoder, fp16 video VAE + fp32 audio VAE.
|
||||
Times are ComfyUI execution time (`execution_start`→`execution_success`), which
|
||||
excludes client/queue overhead (wall time tracked separately in the raw data;
|
||||
typically +1–2 s). "Warm" = model already resident; first-load runs are
|
||||
recorded but excluded from the means. Repeated configs use distinct seeds so
|
||||
ComfyUI's node cache can't short-circuit. Reference inputs are real images
|
||||
(film stills, anime characters, forests/volcanos, up to 23040×3840 —
|
||||
`scripts/pod/fetch-bench-images-v2.py`).
|
||||
|
||||
Hardware: **B200** 180 GB (US, $6.94/hr) · **RTX 4090** 24 GB (EU, $0.74/hr) ·
|
||||
**RTX PRO 4000 Blackwell** 24 GB (EU, $0.57/hr). The EU cards run pruned int8
|
||||
only (bf16/int8 don't fit 24 GB usefully) with `--reserve-vram 2.5`; the RTX
|
||||
4000 additionally needs `--cache-none --disable-pinned-memory` because its
|
||||
container has a 29 GB RAM cap that otherwise OOM-kills weight staging.
|
||||
|
||||
## B200: weight-family comparison (warm, 20 steps)
|
||||
|
||||
The headline: **bf16 is the fastest family on the B200** — quantization costs
|
||||
speed here (dequant overhead) and only pays when VRAM is scarce.
|
||||
|
||||
| Config | bf16 | int8 | pruned int8 | bf16 VRAM | int8 VRAM | pruned VRAM |
|
||||
|---------------------------------|-------------|-------------|-------------|-----------|-----------|-------------|
|
||||
| t2v 864×480 5 s | **51 s** | 63 s | 63 s | 159 GB | 63 GB | 69 GB |
|
||||
| t2v 864×480 15 s | **240 s** | 270 s | 270 s | 182 GB | 101 GB | 75 GB |
|
||||
| t2v 1344×768 5 s | **169 s** | 197 s | 194 s | 177 GB | 99 GB | 73 GB |
|
||||
| ref2v 1 ref 864×480 5 s | **54 s** | 68 s | ~68 s¹ | 186 GB | 95 GB | 68 GB |
|
||||
| ref2v 4 refs match | **63 s** | 86 s | 81 s | 154 GB | 95 GB | 70 GB |
|
||||
| ref2v 8 refs match | **79 s** | 99 s | 95 s | 185 GB | 96 GB | 68 GB |
|
||||
| ref2v 8 refs match 1344×768 | **282 s** | 335 s | 339 s | 158 GB | 103 GB | 77 GB |
|
||||
| ref2v 4 refs max | **91 s** | 117 s | 112 s | 156 GB | 96 GB | 70 GB |
|
||||
| ref2v 8 refs max | **785 s** | 891 s | 841 s | 177 GB | 112 GB | 92 GB |
|
||||
|
||||
¹ pruned 1-ref mean is 90±41 s over 3 runs because the first run absorbed a
|
||||
model swap; the steady-state runs are ~68 s, matching int8.
|
||||
|
||||
bf16 peaks at 155–187 GB — it *only* runs on ≥180 GB cards. int8 fits under
|
||||
112 GB (H200-class); pruned int8 under ~95 GB worst-case, ~70 GB typical
|
||||
(H100-class without offload, 24 GB cards with offload).
|
||||
|
||||
## Scaling behavior (B200, pruned int8)
|
||||
|
||||
Duration (864×480 t2v): 5 s → 63 s, 10 s → 154 s, 15 s → 270 s. Cost per
|
||||
output-second rises from 12.6→15.4→18.0 s — mildly superlinear (attention).
|
||||
|
||||
Resolution (5 s t2v): 608×352 → 32 s, 864×480 → 63 s, 1344×768 → 194 s.
|
||||
Roughly ∝ pixels^1.2. The corner case 1344×768×15 s = **1108 s** (18.5 min).
|
||||
|
||||
Reference count (864×480 5 s, `match` sizing): 1 ref ≈ 68 s, 4 refs ≈ 81 s,
|
||||
8 refs ≈ 95 s — ~+4 s per extra reference image. Cheap.
|
||||
|
||||
Reference sizing `max` is the expensive lever: with 8 large refs it goes
|
||||
95 s → **841 s** (~9×), because 2048px reference tokens ride through every
|
||||
sampling step. Use `match` unless identity fidelity demands otherwise.
|
||||
|
||||
## Cross-GPU (pruned int8, t2v 864×480 5 s, warm)
|
||||
|
||||
| GPU | $/hr | time/video | videos/hr | $/video | notes |
|
||||
|------------------|-------|-----------------|-----------|------------|--------------------------------|
|
||||
| B200 (bf16) | $6.94 | 51 s | 70 | $0.098 | fastest family on this card |
|
||||
| B200 (pruned) | $6.94 | 63 s | 57 | $0.121 | |
|
||||
| RTX 4090 | $0.74 | ~222 s | ~16 | ~$0.046 | suite still running — interim |
|
||||
| RTX PRO 4000 | $0.57 | pending | pending | pending | suite still running |
|
||||
|
||||
(EU numbers will be finalized when their suites complete; ref2v numbers for
|
||||
both cards land then too.)
|
||||
|
||||
## Concurrency (2 ComfyUI instances, one B200)
|
||||
|
||||
See `docs/CONCURRENCY.md` for the model. Experiment: two resident instances,
|
||||
sequential pair vs concurrent pair of identical warm t2v jobs.
|
||||
|
||||
RESULTS_PENDING
|
||||
|
||||
## Takeaways
|
||||
|
||||
1. **On the B200, run bf16.** Fastest and highest quality; quantized families
|
||||
exist for smaller cards, not for speed.
|
||||
2. **Many references are cheap; big references are not.** 8 refs at `match`
|
||||
costs ~1.4× a single ref. `max` sizing costs up to ~9×. Default `match`.
|
||||
3. **Duration and resolution both scale superlinearly**; the trained envelope
|
||||
corner (15 s @ 1344×768) costs ~18 min even on a B200 — treat full-res
|
||||
long clips as premium jobs.
|
||||
4. **24 GB consumer cards work but pay a heavy offload tax** (~3.5× B200
|
||||
latency rather than the ~3× raw-compute ratio), and cheap pod tiers bring
|
||||
operational traps: container RAM caps (29 GB on the RTX 4000 pod) and
|
||||
stock torch builds too old for H3 (needs ≥2.5 for `enable_gqa`; Blackwell
|
||||
needs ≥2.7+cu128).
|
||||
Reference in New Issue
Block a user