mirror of
https://github.com/storytold/cloud-worker.git
synced 2026-10-09 00:09:43 +00:00
final: 5090 tables (incl. bf16-streaming), stats tooling covers all 5 datasets
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
+3
-3
@@ -34,9 +34,9 @@ Blackwell (B200/5090/RTX PRO 4000) torch ≥ 2.7 with cu128 wheels.
|
||||
| Pod class | VRAM | RAM cap | Required config | What works | What fails |
|
||||
|----------------------|--------|---------|--------------------------------------------------------------|---------------------------------------------------|-------------------------------|
|
||||
| B200 ($6.94/hr) | 180 GB | ~3 TB | none (defaults) | everything incl. 9-ref, 15 s, full-res, bf16 | nothing observed (96/96) |
|
||||
| RTX 5090 ($0.99/hr) | 32 GB | 124 GB | #15316 patch + `--reserve-vram 4` | multi-ref up to 8 full-size refs confirmed; 9-ref + max sizing under test | TBD (matrix running) |
|
||||
| RTX 4090 ($0.74/hr) | 24 GB | 62 GB | #15316 patch + `--reserve-vram 5` + `--cache-none` + `--disable-pinned-memory` | t2v/i2v ≤10 s @ ≤864×480, 1-ref ref2v; multi-ref under re-test with patch | 15 s clips, 1344×768, multi-ref (pre-patch) |
|
||||
| RTX PRO 4000 ($0.57/hr) | 24 GB | 29 GB | same as 4090 (mandatory even for t2v) | t2v/i2v, 1-ref (551–588 s — uneconomical) | multi-ref untested post-fix; 15 s died |
|
||||
| RTX 5090 ($0.99/hr) | 32 GB | 124 GB | #15316 patch + `--reserve-vram 4` (add `--cache-none` when switching bf16 models) | EVERYTHING: 1–9 refs full-size, 9-ref max (34 min), 15 s, full-res, all aspects, bf16 incl. 8-ref (369 s) | bf16 model switching without cache-none (RAM kill) |
|
||||
| RTX 4090 ($0.74/hr) | 24 GB | 62 GB | #15316 patch + `--reserve-vram 5` + `--cache-none` + `--disable-pinned-memory` | **multi-ref RESCUED**: 8-ref 429 s, 9-ref 444 s, 8-ref@1344×768 1159 s (17–19 GB peaks), all aspects | 9-ref `max` sizing only |
|
||||
| RTX PRO 4000 ($0.57/hr) | 24 GB | 29 GB | same as 4090 (mandatory even for t2v) | t2v/i2v, 1-ref (551–588 s — uneconomical) | dropped from fleet on economics |
|
||||
|
||||
Key procurement insight: **the container RAM cap, not the GPU, decides the
|
||||
mitigation burden.** A 24 GB card in a ≥96 GB-RAM container would skip
|
||||
|
||||
Reference in New Issue
Block a user