mirror of
https://github.com/storytold/cloud-worker.git
synced 2026-10-09 00:09:43 +00:00
BENCHMARKS.md: final cross-GPU tables, failure stats, multi-ref rescue status
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
+46
-12
@@ -60,17 +60,30 @@ Reference sizing `max` is the expensive lever: with 8 large refs it goes
|
||||
95 s → **841 s** (~9×), because 2048px reference tokens ride through every
|
||||
sampling step. Use `match` unless identity fidelity demands otherwise.
|
||||
|
||||
## Cross-GPU (pruned int8, t2v 864×480 5 s, warm)
|
||||
## Cross-GPU (pruned int8, warm, 20 steps — final measured)
|
||||
|
||||
| GPU | $/hr | time/video | videos/hr | $/video | notes |
|
||||
|------------------|-------|-----------------|-----------|------------|--------------------------------|
|
||||
| B200 (bf16) | $6.94 | 51 s | 70 | $0.098 | fastest family on this card |
|
||||
| B200 (pruned) | $6.94 | 63 s | 57 | $0.121 | |
|
||||
| RTX 4090 | $0.74 | ~222 s | ~16 | ~$0.046 | suite still running — interim |
|
||||
| RTX PRO 4000 | $0.57 | pending | pending | pending | suite still running |
|
||||
t2v 864×480 5 s:
|
||||
|
||||
(EU numbers will be finalized when their suites complete; ref2v numbers for
|
||||
both cards land then too.)
|
||||
| GPU | $/hr | time/video | videos/hr | $/video | notes |
|
||||
|------------------|-------|------------------|-----------|---------|------------------------------------|
|
||||
| B200 (bf16) | $6.94 | 51 s (n=2) | 70 | $0.098 | fastest family on this card |
|
||||
| B200 (pruned) | $6.94 | 63 s (n=3) | 57 | $0.121 | |
|
||||
| RTX 4090 | $0.74 | 222 s (n=3, ±0) | 16 | $0.046 | --reserve-vram 2.5 |
|
||||
| RTX PRO 4000 | $0.57 | 553 s (n=3, ±3) | 6.5 | $0.087 | --cache-none tax (29 GB RAM cap) |
|
||||
|
||||
ref2v single ref, 864×480 5 s:
|
||||
|
||||
| GPU | time/video | $/video |
|
||||
|------------------|------------------|---------|
|
||||
| B200 (bf16) | 54 s (n=3) | $0.104 |
|
||||
| RTX 4090 | 287–392 s | $0.059–0.081 |
|
||||
| RTX PRO 4000 | 581 s (n=3, ±6) | $0.092 |
|
||||
|
||||
**Verdict:** the 4090 is the only consumer card here with a real cost
|
||||
advantage (~2.5× cheaper than B200 per t2v video). The RTX PRO 4000's
|
||||
low price is consumed by its slow, reload-heavy runs — it is NOT more
|
||||
economical than the 4090, and barely beats the B200. Neither 24 GB card can
|
||||
do multi-ref, 15 s clips, or full-res today (see reliability section).
|
||||
|
||||
## Concurrency (2 ComfyUI instances, one B200)
|
||||
|
||||
@@ -97,11 +110,11 @@ failure tables with `python3 bench/failures.py`, and validate output files
|
||||
with `python3 bench/validate_outputs.py` (downloads each MP4 and ffprobes
|
||||
streams + duration). Snapshot as of 2026-08-06 ~09:00 UTC:
|
||||
|
||||
| GPU | success rate | dominant failure |
|
||||
| GPU | success rate | dominant failures |
|
||||
|-------------------|--------------|-------------------------------------------------------|
|
||||
| B200 180 GB | 96/96 (100%) | none |
|
||||
| RTX 4090 24 GB | 8/42 (19%) | VRAM OOM in text-encoder vision path on ref2v |
|
||||
| RTX PRO 4000 24 GB| 6/7 so far | one container-RAM OOM kill before mitigation flags |
|
||||
| RTX 4090 24 GB | 9/53 (17%) | VRAM OOM: TE vision (multi-ref), DiT sampling (refs/15s/full-res); cascading server death |
|
||||
| RTX PRO 4000 24 GB| 7/12 (58%) | container-RAM OOM (pre-flags); corrupt ref image (fixed); server death on 15 s t2v |
|
||||
|
||||
**Output validation**: 100% of successful runs produced valid videos (h264 +
|
||||
stereo AAC at exactly the requested duration) — every success on all three
|
||||
@@ -129,6 +142,27 @@ single-ref ref2v only. Multi-reference work — the priority modality — needs
|
||||
≥32 GB (5090-class, unverified) and realistically ≥48 GB for headroom, or it
|
||||
stays on big-GPU pools.
|
||||
|
||||
## Multi-ref rescue status (in progress)
|
||||
|
||||
Community research (see agent report summarized in commit history) identified
|
||||
the root cause as **ComfyUI PR #15316**: VRAM is budgeted for TE weights but
|
||||
not image-encode activations. Two OOM phases were isolated on the 4090:
|
||||
|
||||
1. **TE vision encode** — fixed by ref pre-downscaling (`--ref-downscale`,
|
||||
now in the harness) and/or the #15316 patch + higher `--reserve-vram`.
|
||||
Confirmed: with 0.1–0.2 MP refs the TE phase completes.
|
||||
2. **DiT sampling with ref tokens** — the 20 GB DiT + ref-token-enlarged
|
||||
attention doesn't fit 24 GB. `--lowvram` made things worse (crashes).
|
||||
Remaining candidates: #15316 patch active + `--reserve-vram 5` (patched
|
||||
file is applied on both EU pods, pending verification), TE eviction via
|
||||
ComfyUI-H3-Multishot, or a smaller DiT quant (community INT4/W4A8,
|
||||
~10–12 GB, quality TBD).
|
||||
|
||||
Verification was interrupted when all three pods went offline simultaneously
|
||||
(apparent RunPod account event ~09:30 UTC). Resume plan lives in the repo:
|
||||
restart pods → `scripts/local/push-to-pod.sh` → rescue matrix via
|
||||
`bench/minimax_bench.py --suite refheavy --ref-downscale 0.2`.
|
||||
|
||||
## Takeaways
|
||||
|
||||
1. **On the B200, run bf16.** Fastest and highest quality; quantized families
|
||||
|
||||
Reference in New Issue
Block a user