MITIGATIONS.md: per-machine-class OOM map — RAM cap decides mitigation burden

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Hanashi
2026-08-06 17:09:40 -04:00
parent a88626c0ca
commit 9fc70a0b81
2 changed files with 67 additions and 0 deletions
+1
View File
@@ -29,6 +29,7 @@ Current test rig: 1x B200 (180 GB) pod — connection details in `pod-info.txt`.
(SGLang / vLLM-Omni / diffusers), concurrency model, RunPod serverless (SGLang / vLLM-Omni / diffusers), concurrency model, RunPod serverless
- `docs/BENCHMARKS.md` — measured results on the B200 - `docs/BENCHMARKS.md` — measured results on the B200
- `docs/CONCURRENCY.md` — concurrency model + cost-per-video math across GPU tiers - `docs/CONCURRENCY.md` — concurrency model + cost-per-video math across GPU tiers
- `docs/MITIGATIONS.md` — per-machine-class OOM map and required ComfyUI flags/patches
## Quick start ## Quick start
+66
View File
@@ -0,0 +1,66 @@
# MiniMax H3 — per-machine-class mitigations & OOM map
Compiled 2026-08-06 from measured failures and fixes across five pods.
Companion to `docs/BENCHMARKS.md` (runtimes) and `bench/failures.py` (stats).
## The two resources that fail, and their fixes
**1. Container RAM (cgroup cap — check `/sys/fs/cgroup/memory/memory.limit_in_bytes`,
cgroup v1 on RunPod consumer pods, or `/sys/fs/cgroup/memory.max` for v2).**
ComfyUI stages every loaded model's weights in host RAM ("prepared for dynamic
VRAM loading"). Task-switching between fl2va and ref2va stacks WITHOUT eviction
stages TE (15 GB) + two DiTs (20 GB each) + overhead ≈ 60 GB+. Exceeding the
cap = silent SIGKILL, no traceback, "server died" mid-run.
Fix: `--cache-none` (one model staged at a time) + `--disable-pinned-memory`
(staging stays pageable/reclaimable). Cost: weights reload per run — measured
~2.3× slowdown on the RTX 4000 (551 s vs an est. ~240 s), ~1.4× on the 4090.
**2. VRAM, in two distinct phases:**
- *TE vision encode* (multi-ref): budget bug, fixed by ComfyUI **PR #15316**
(apply: `git fetch origin pull/15316/head:pr15316 && git checkout pr15316 --
comfy/text_encoders/minimax.py`) + `--reserve-vram 4`–`5`. Ref pre-downscale
(`--ref-downscale 0.2` in our harness) also clears this phase alone.
- *DiT sampling with ref tokens / long clips / large canvas*: needs the
activations to fit beside the (partially loaded) 20 GB DiT. On 24 GB this is
the hard wall: 15 s clips, 1344×768, and multi-ref all exceeded it pre-patch.
`--lowvram` made it WORSE (crashes) — do not use for H3.
Hygiene for all capped machines: `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`
(set automatically by `start-comfyui.sh`), torch ≥ 2.5 (`enable_gqa`), and for
Blackwell (B200/5090/RTX PRO 4000) torch ≥ 2.7 with cu128 wheels.
## Machine-class matrix (all measured, not theoretical)
| Pod class | VRAM | RAM cap | Required config | What works | What fails |
|----------------------|--------|---------|--------------------------------------------------------------|---------------------------------------------------|-------------------------------|
| B200 ($6.94/hr) | 180 GB | ~3 TB | none (defaults) | everything incl. 9-ref, 15 s, full-res, bf16 | nothing observed (96/96) |
| RTX 5090 ($0.99/hr) | 32 GB | 124 GB | #15316 patch + `--reserve-vram 4` | multi-ref up to 8 full-size refs confirmed; 9-ref + max sizing under test | TBD (matrix running) |
| RTX 4090 ($0.74/hr) | 24 GB | 62 GB | #15316 patch + `--reserve-vram 5` + `--cache-none` + `--disable-pinned-memory` | t2v/i2v ≤10 s @ ≤864×480, 1-ref ref2v; multi-ref under re-test with patch | 15 s clips, 1344×768, multi-ref (pre-patch) |
| RTX PRO 4000 ($0.57/hr) | 24 GB | 29 GB | same as 4090 (mandatory even for t2v) | t2v/i2v, 1-ref (551–588 s — uneconomical) | multi-ref untested post-fix; 15 s died |
Key procurement insight: **the container RAM cap, not the GPU, decides the
mitigation burden.** A 24 GB card in a ≥96 GB-RAM container would skip
`--cache-none` and its ~1.4–2.3× tax. When renting, prefer pod configs with
RAM ≥ 4× VRAM.
## Reference-image envelope (node schema + measured)
- The node accepts up to **9 reference images** + 3 reference videos (each
with optional soundtrack) + 3 standalone audio clips — "12 refs" = 9 images
+ 3 videos.
- `match` sizing (refs scaled down to canvas area) is the default and the
cheap path; `max` (2048 px short edge) multiplies sampling cost ~9× at
8 refs and is the most VRAM-hungry config in the whole matrix.
- Ref count scaling on B200: +4 s per extra ref (match). Counts 1→9 and the
ref-resolution ladder (0.1 MP → native) are being measured on all three
GPU classes right now (`--suite refmatrix`).
## Open items
- 5090: 9-ref and 8-ref-`max` verdicts (running).
- 4090 with full stack (#15316 + cache-none + reserve 5): multi-ref verdict
(running).
- RTX PRO 4000: dropped from fleet consideration on economics; would need
re-test only if its price class matters.
- Upstream: file the multi-ref OOM issue referencing #15316 (user decision —
outward-facing).