BENCHMARKS.md: final cross-GPU tables, failure stats, multi-ref rescue status

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Hanashi
2026-08-06 15:23:36 -04:00
parent c0a875423e
commit a4bd386e52
+46 -12
View File
@@ -60,17 +60,30 @@ Reference sizing `max` is the expensive lever: with 8 large refs it goes
95 s → **841 s** (~9×), because 2048px reference tokens ride through every
sampling step. Use `match` unless identity fidelity demands otherwise.
## Cross-GPU (pruned int8, t2v 864×480 5 s, warm)
## Cross-GPU (pruned int8, warm, 20 steps — final measured)
| GPU | $/hr | time/video | videos/hr | $/video | notes |
|------------------|-------|-----------------|-----------|------------|--------------------------------|
| B200 (bf16) | $6.94 | 51 s | 70 | $0.098 | fastest family on this card |
| B200 (pruned) | $6.94 | 63 s | 57 | $0.121 | |
| RTX 4090 | $0.74 | ~222 s | ~16 | ~$0.046 | suite still running — interim |
| RTX PRO 4000 | $0.57 | pending | pending | pending | suite still running |
t2v 864×480 5 s:
(EU numbers will be finalized when their suites complete; ref2v numbers for
both cards land then too.)
| GPU | $/hr | time/video | videos/hr | $/video | notes |
|------------------|-------|------------------|-----------|---------|------------------------------------|
| B200 (bf16) | $6.94 | 51 s (n=2) | 70 | $0.098 | fastest family on this card |
| B200 (pruned) | $6.94 | 63 s (n=3) | 57 | $0.121 | |
| RTX 4090 | $0.74 | 222 s (n=3, ±0) | 16 | $0.046 | --reserve-vram 2.5 |
| RTX PRO 4000 | $0.57 | 553 s (n=3, ±3) | 6.5 | $0.087 | --cache-none tax (29 GB RAM cap) |
ref2v single ref, 864×480 5 s:
| GPU | time/video | $/video |
|------------------|------------------|---------|
| B200 (bf16) | 54 s (n=3) | $0.104 |
| RTX 4090 | 287–392 s | $0.059–0.081 |
| RTX PRO 4000 | 581 s (n=3, ±6) | $0.092 |
**Verdict:** the 4090 is the only consumer card here with a real cost
advantage (~2.5× cheaper than B200 per t2v video). The RTX PRO 4000's
low price is consumed by its slow, reload-heavy runs — it is NOT more
economical than the 4090, and barely beats the B200. Neither 24 GB card can
do multi-ref, 15 s clips, or full-res today (see reliability section).
## Concurrency (2 ComfyUI instances, one B200)
@@ -97,11 +110,11 @@ failure tables with `python3 bench/failures.py`, and validate output files
with `python3 bench/validate_outputs.py` (downloads each MP4 and ffprobes
streams + duration). Snapshot as of 2026-08-06 ~09:00 UTC:
| GPU | success rate | dominant failure |
| GPU | success rate | dominant failures |
|-------------------|--------------|-------------------------------------------------------|
| B200 180 GB | 96/96 (100%) | none |
| RTX 4090 24 GB | 8/42 (19%) | VRAM OOM in text-encoder vision path on ref2v |
| RTX PRO 4000 24 GB| 6/7 so far | one container-RAM OOM kill before mitigation flags |
| RTX 4090 24 GB | 9/53 (17%) | VRAM OOM: TE vision (multi-ref), DiT sampling (refs/15s/full-res); cascading server death |
| RTX PRO 4000 24 GB| 7/12 (58%) | container-RAM OOM (pre-flags); corrupt ref image (fixed); server death on 15 s t2v |
**Output validation**: 100% of successful runs produced valid videos (h264 +
stereo AAC at exactly the requested duration) — every success on all three
@@ -129,6 +142,27 @@ single-ref ref2v only. Multi-reference work — the priority modality — needs
≥32 GB (5090-class, unverified) and realistically ≥48 GB for headroom, or it
stays on big-GPU pools.
## Multi-ref rescue status (in progress)
Community research (see agent report summarized in commit history) identified
the root cause as **ComfyUI PR #15316**: VRAM is budgeted for TE weights but
not image-encode activations. Two OOM phases were isolated on the 4090:
1. **TE vision encode** — fixed by ref pre-downscaling (`--ref-downscale`,
now in the harness) and/or the #15316 patch + higher `--reserve-vram`.
Confirmed: with 0.1–0.2 MP refs the TE phase completes.
2. **DiT sampling with ref tokens** — the 20 GB DiT + ref-token-enlarged
attention doesn't fit 24 GB. `--lowvram` made things worse (crashes).
Remaining candidates: #15316 patch active + `--reserve-vram 5` (patched
file is applied on both EU pods, pending verification), TE eviction via
ComfyUI-H3-Multishot, or a smaller DiT quant (community INT4/W4A8,
~10–12 GB, quality TBD).
Verification was interrupted when all three pods went offline simultaneously
(apparent RunPod account event ~09:30 UTC). Resume plan lives in the repo:
restart pods → `scripts/local/push-to-pod.sh` → rescue matrix via
`bench/minimax_bench.py --suite refheavy --ref-downscale 0.2`.
## Takeaways
1. **On the B200, run bf16.** Fastest and highest quality; quantized families