From a4bd386e5298f84269be8d5d6a6dcdc43ab80790 Mon Sep 17 00:00:00 2001 From: Hanashi Date: Thu, 6 Aug 2026 15:23:36 -0400 Subject: [PATCH] BENCHMARKS.md: final cross-GPU tables, failure stats, multi-ref rescue status Co-Authored-By: Claude Fable 5 --- docs/BENCHMARKS.md | 58 ++++++++++++++++++++++++++++++++++++---------- 1 file changed, 46 insertions(+), 12 deletions(-) diff --git a/docs/BENCHMARKS.md b/docs/BENCHMARKS.md index 0937549..7320858 100644 --- a/docs/BENCHMARKS.md +++ b/docs/BENCHMARKS.md @@ -60,17 +60,30 @@ Reference sizing `max` is the expensive lever: with 8 large refs it goes 95 s → **841 s** (~9×), because 2048px reference tokens ride through every sampling step. Use `match` unless identity fidelity demands otherwise. -## Cross-GPU (pruned int8, t2v 864×480 5 s, warm) +## Cross-GPU (pruned int8, warm, 20 steps — final measured) -| GPU | $/hr | time/video | videos/hr | $/video | notes | -|------------------|-------|-----------------|-----------|------------|--------------------------------| -| B200 (bf16) | $6.94 | 51 s | 70 | $0.098 | fastest family on this card | -| B200 (pruned) | $6.94 | 63 s | 57 | $0.121 | | -| RTX 4090 | $0.74 | ~222 s | ~16 | ~$0.046 | suite still running — interim | -| RTX PRO 4000 | $0.57 | pending | pending | pending | suite still running | +t2v 864×480 5 s: -(EU numbers will be finalized when their suites complete; ref2v numbers for -both cards land then too.) +| GPU | $/hr | time/video | videos/hr | $/video | notes | +|------------------|-------|------------------|-----------|---------|------------------------------------| +| B200 (bf16) | $6.94 | 51 s (n=2) | 70 | $0.098 | fastest family on this card | +| B200 (pruned) | $6.94 | 63 s (n=3) | 57 | $0.121 | | +| RTX 4090 | $0.74 | 222 s (n=3, ±0) | 16 | $0.046 | --reserve-vram 2.5 | +| RTX PRO 4000 | $0.57 | 553 s (n=3, ±3) | 6.5 | $0.087 | --cache-none tax (29 GB RAM cap) | + +ref2v single ref, 864×480 5 s: + +| GPU | time/video | $/video | +|------------------|------------------|---------| +| B200 (bf16) | 54 s (n=3) | $0.104 | +| RTX 4090 | 287–392 s | $0.059–0.081 | +| RTX PRO 4000 | 581 s (n=3, ±6) | $0.092 | + +**Verdict:** the 4090 is the only consumer card here with a real cost +advantage (~2.5× cheaper than B200 per t2v video). The RTX PRO 4000's +low price is consumed by its slow, reload-heavy runs — it is NOT more +economical than the 4090, and barely beats the B200. Neither 24 GB card can +do multi-ref, 15 s clips, or full-res today (see reliability section). ## Concurrency (2 ComfyUI instances, one B200) @@ -97,11 +110,11 @@ failure tables with `python3 bench/failures.py`, and validate output files with `python3 bench/validate_outputs.py` (downloads each MP4 and ffprobes streams + duration). Snapshot as of 2026-08-06 ~09:00 UTC: -| GPU | success rate | dominant failure | +| GPU | success rate | dominant failures | |-------------------|--------------|-------------------------------------------------------| | B200 180 GB | 96/96 (100%) | none | -| RTX 4090 24 GB | 8/42 (19%) | VRAM OOM in text-encoder vision path on ref2v | -| RTX PRO 4000 24 GB| 6/7 so far | one container-RAM OOM kill before mitigation flags | +| RTX 4090 24 GB | 9/53 (17%) | VRAM OOM: TE vision (multi-ref), DiT sampling (refs/15s/full-res); cascading server death | +| RTX PRO 4000 24 GB| 7/12 (58%) | container-RAM OOM (pre-flags); corrupt ref image (fixed); server death on 15 s t2v | **Output validation**: 100% of successful runs produced valid videos (h264 + stereo AAC at exactly the requested duration) — every success on all three @@ -129,6 +142,27 @@ single-ref ref2v only. Multi-reference work — the priority modality — needs ≥32 GB (5090-class, unverified) and realistically ≥48 GB for headroom, or it stays on big-GPU pools. +## Multi-ref rescue status (in progress) + +Community research (see agent report summarized in commit history) identified +the root cause as **ComfyUI PR #15316**: VRAM is budgeted for TE weights but +not image-encode activations. Two OOM phases were isolated on the 4090: + +1. **TE vision encode** — fixed by ref pre-downscaling (`--ref-downscale`, + now in the harness) and/or the #15316 patch + higher `--reserve-vram`. + Confirmed: with 0.1–0.2 MP refs the TE phase completes. +2. **DiT sampling with ref tokens** — the 20 GB DiT + ref-token-enlarged + attention doesn't fit 24 GB. `--lowvram` made things worse (crashes). + Remaining candidates: #15316 patch active + `--reserve-vram 5` (patched + file is applied on both EU pods, pending verification), TE eviction via + ComfyUI-H3-Multishot, or a smaller DiT quant (community INT4/W4A8, + ~10–12 GB, quality TBD). + +Verification was interrupted when all three pods went offline simultaneously +(apparent RunPod account event ~09:30 UTC). Resume plan lives in the repo: +restart pods → `scripts/local/push-to-pod.sh` → rescue matrix via +`bench/minimax_bench.py --suite refheavy --ref-downscale 0.2`. + ## Takeaways 1. **On the B200, run bf16.** Fastest and highest quality; quantized families