final: 5090 tables (incl. bf16-streaming), stats tooling covers all 5 datasets

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Hanashi
2026-08-06 19:21:16 -04:00
parent 056d5785e7
commit 77524a10e6
16 changed files with 232 additions and 8 deletions
+26 -5
View File
@@ -79,11 +79,32 @@ ref2v single ref, 864×480 5 s:
| RTX 4090 | 287–392 s | $0.059–0.081 |
| RTX PRO 4000 | 581 s (n=3, ±6) | $0.092 |
**Verdict:** the 4090 is the only consumer card here with a real cost
advantage (~2.5× cheaper than B200 per t2v video). The RTX PRO 4000's
low price is consumed by its slow, reload-heavy runs — it is NOT more
economical than the 4090, and barely beats the B200. Neither 24 GB card can
do multi-ref, 15 s clips, or full-res today (see reliability section).
**RTX 5090 (32 GB, $0.99/hr) — added later, the fleet workhorse:**
| Config (pruned int8, warm) | time | $/video | notes |
|-----------------------------|------------------|---------|------------------------------------|
| t2v 864×480 5 s | 143 s (n=3, ±1) | $0.039 | cheapest t2v measured |
| ref2v 1 ref | 153 s (n=3, ±0) | $0.042 | |
| ref2v 8 refs full-size | 220 s (n=3, ±1) | $0.060 | vs $0.152 B200-bf16: 2.5× cheaper |
| ref2v 9 refs full-size | 237 s | $0.065 | max ref count |
| t2v 15 s | 582 s | $0.160 | works (killed 24 GB pre-patch) |
| t2v 1344×768 | 455 s | $0.125 | works |
| 9:16 480×864 / 1:1 640×640 | 187 s / 141 s | — | aspects fine |
| ref2v 9 refs `max` sizing | 2031 s (34 min) | $0.559 | works; avoid in production |
**bf16 on the 5090** (66 GB DiT layer-streamed through 32 GB): warm t2v =
**110 s — faster than its own pruned int8 (143 s)**; ref2v 1/4/8 refs =
310/336/369 s (needs `--cache-none` when switching between bf16 models —
two 66 GB stacks break the 124 GB container RAM cap). B200-quality output
at $0.99/hr for everything except 15 s + full-res extremes.
**Verdict:** the 5090 dominates cost AND capability among consumer cards —
cheapest per video on every task shape, full reference envelope, and a
credible bf16 quality tier. The patched 4090 (see MITIGATIONS.md) is a
viable budget multi-ref tier (~8-9 refs at 429–444 s). The RTX PRO 4000's
low price is consumed by its slow, reload-heavy runs — dropped. Fleet
recipe: 5090 pools for volume, B200 pool only for 15 s/full-res premium
jobs and fastest-latency needs.
## Concurrency (2 ComfyUI instances, one B200)
+3 -3
View File
@@ -34,9 +34,9 @@ Blackwell (B200/5090/RTX PRO 4000) torch ≥ 2.7 with cu128 wheels.
| Pod class | VRAM | RAM cap | Required config | What works | What fails |
|----------------------|--------|---------|--------------------------------------------------------------|---------------------------------------------------|-------------------------------|
| B200 ($6.94/hr) | 180 GB | ~3 TB | none (defaults) | everything incl. 9-ref, 15 s, full-res, bf16 | nothing observed (96/96) |
| RTX 5090 ($0.99/hr) | 32 GB | 124 GB | #15316 patch + `--reserve-vram 4` | multi-ref up to 8 full-size refs confirmed; 9-ref + max sizing under test | TBD (matrix running) |
| RTX 4090 ($0.74/hr) | 24 GB | 62 GB | #15316 patch + `--reserve-vram 5` + `--cache-none` + `--disable-pinned-memory` | t2v/i2v ≤10 s @ ≤864×480, 1-ref ref2v; multi-ref under re-test with patch | 15 s clips, 1344×768, multi-ref (pre-patch) |
| RTX PRO 4000 ($0.57/hr) | 24 GB | 29 GB | same as 4090 (mandatory even for t2v) | t2v/i2v, 1-ref (551–588 s — uneconomical) | multi-ref untested post-fix; 15 s died |
| RTX 5090 ($0.99/hr) | 32 GB | 124 GB | #15316 patch + `--reserve-vram 4` (add `--cache-none` when switching bf16 models) | EVERYTHING: 1–9 refs full-size, 9-ref max (34 min), 15 s, full-res, all aspects, bf16 incl. 8-ref (369 s) | bf16 model switching without cache-none (RAM kill) |
| RTX 4090 ($0.74/hr) | 24 GB | 62 GB | #15316 patch + `--reserve-vram 5` + `--cache-none` + `--disable-pinned-memory` | **multi-ref RESCUED**: 8-ref 429 s, 9-ref 444 s, 8-ref@1344×768 1159 s (17–19 GB peaks), all aspects | 9-ref `max` sizing only |
| RTX PRO 4000 ($0.57/hr) | 24 GB | 29 GB | same as 4090 (mandatory even for t2v) | t2v/i2v, 1-ref (551–588 s — uneconomical) | dropped from fleet on economics |
Key procurement insight: **the container RAM cap, not the GPU, decides the
mitigation burden.** A 24 GB card in a ≥96 GB-RAM container would skip