diff --git a/docs/RESEARCH.md b/docs/RESEARCH.md index 412763e..51841fa 100644 --- a/docs/RESEARCH.md +++ b/docs/RESEARCH.md @@ -136,14 +136,14 @@ already do. For scale-to-zero economics: ## Published performance reference points -| Setup | Config | Time | -|---|---|---| -| 4x H100 80GB (SGLang TP2+Ulysses2) | 1344×768, 124 fr, 50 steps | 13.3 s | -| 8x B300 (SGLang, bf16 / online FP8) | same | 19.0 / 18.0 s | -| 4x H200 (SGLang Ulysses4) | same | 75.1 s (53.7 s Cache-DiT) | -| 2x RTX 5090 (layerwise offload, ~380 GB host RAM) | same | 560 s | -| RTX 4090 Laptop 16 GB (SageAttention) | 960×540, 5 s, 20 steps | 182 s | -| RTX 3060 12 GB (heavy offload) | 864×480, 124 fr, 20 steps | <9 min | +| Setup | Config | Time | +|---------------------------------------------------|----------------------------|---------------------------| +| 4x H100 80GB (SGLang TP2+Ulysses2) | 1344×768, 124 fr, 50 steps | 13.3 s | +| 8x B300 (SGLang, bf16 / online FP8) | same | 19.0 / 18.0 s | +| 4x H200 (SGLang Ulysses4) | same | 75.1 s (53.7 s Cache-DiT) | +| 2x RTX 5090 (layerwise offload, ~380 GB host RAM) | same | 560 s | +| RTX 4090 Laptop 16 GB (SageAttention) | 960×540, 5 s, 20 steps | 182 s | +| RTX 3060 12 GB (heavy offload) | 864×480, 124 fr, 20 steps | <9 min | Our single-B200 ComfyUI numbers: see `docs/BENCHMARKS.md`. diff --git a/docs/WEIGHTS.md b/docs/WEIGHTS.md index 9877af2..ee90707 100644 --- a/docs/WEIGHTS.md +++ b/docs/WEIGHTS.md @@ -18,29 +18,29 @@ ComfyUI finds both storage locations below via `/ComfyUI/extra_model_paths.yaml` ## Storage locations on the pod -| Path | Medium | Notes | -|---|---|---| -| `/workspace/minimax-h3/models/` | RunPod network volume | **Persistent.** The volume has a ~100 GB quota and is effectively full. | -| `/dev/shm/minimax-h3-models/` | tmpfs (RAM, 176 GB) | **EPHEMERAL** — wiped on pod restart. Refill with `bash /workspace/minimax-h3/download-weights-tmpfs.sh` (~10 min). | +| Path | Medium | Notes | +|---------------------------------|-----------------------|---------------------------------------------------------------------------------------------------------------------| +| `/workspace/minimax-h3/models/` | RunPod network volume | **Persistent.** The volume has a ~100 GB quota and is effectively full. | +| `/dev/shm/minimax-h3-models/` | tmpfs (RAM, 176 GB) | **EPHEMERAL** — wiped on pod restart. Refill with `bash /workspace/minimax-h3/download-weights-tmpfs.sh` (~10 min). | ## Downloaded — persistent, `/workspace/minimax-h3/models/` -| File | Family | Size | -|---|---|---| -| `diffusion_models/minimax_h3_ref2va_pruned_int8_convrot.safetensors` | **int8 pruned** (ref2va) | 21 GB | -| `diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors` | **int8 pruned** (fl2va) | 21 GB | -| `diffusion_models/minimax_h3_ref2va_int8_convrot.safetensors` | **int8** (ref2va) | 34 GB | -| `text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors` | text encoder (Qwen3-VL-32B, nvfp4 AWQ) | 16 GB | -| `vae/minimax_h3_video_vae_fp16.safetensors` | video VAE fp16 | 5.2 GB | -| `vae/minimax_h3_audio_vae_fp32.safetensors` | audio VAE fp32 | 0.6 GB | +| File | Family | Size | +|----------------------------------------------------------------------|----------------------------------------|--------| +| `diffusion_models/minimax_h3_ref2va_pruned_int8_convrot.safetensors` | **int8 pruned** (ref2va) | 21 GB | +| `diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors` | **int8 pruned** (fl2va) | 21 GB | +| `diffusion_models/minimax_h3_ref2va_int8_convrot.safetensors` | **int8** (ref2va) | 34 GB | +| `text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors` | text encoder (Qwen3-VL-32B, nvfp4 AWQ) | 16 GB | +| `vae/minimax_h3_video_vae_fp16.safetensors` | video VAE fp16 | 5.2 GB | +| `vae/minimax_h3_audio_vae_fp32.safetensors` | audio VAE fp32 | 0.6 GB | ## Downloaded — ephemeral tmpfs, `/dev/shm/minimax-h3-models/` -| File | Family | Size | -|---|---|---| -| `diffusion_models/minimax_h3_fl2va_int8_convrot.safetensors` | **int8** (fl2va) | 34 GB | -| `diffusion_models/minimax_h3_ref2va_bf16.safetensors` | **bf16** (ref2va) | 66 GB | -| `diffusion_models/minimax_h3_fl2va_bf16.safetensors` | **bf16** (fl2va) | 66 GB | +| File | Family | Size | +|--------------------------------------------------------------|-------------------|-------| +| `diffusion_models/minimax_h3_fl2va_int8_convrot.safetensors` | **int8** (fl2va) | 34 GB | +| `diffusion_models/minimax_h3_ref2va_bf16.safetensors` | **bf16** (ref2va) | 66 GB | +| `diffusion_models/minimax_h3_fl2va_bf16.safetensors` | **bf16** (fl2va) | 66 GB | The int8 family is split across the two locations only because of the volume quota: ref2va_int8 landed on the volume before it filled; fl2va_int8 and both @@ -51,12 +51,12 @@ byte-verified against the HF API). ## Not downloaded (exist upstream in the same repo) -| File | Size | Why skipped | -|---|---|---| -| `diffusion_models/minimax_h3_{fl2va,ref2va}_pruned_bf16.safetensors` | 40 GB each | a 4th family (pruned, unquantized) — not in the 3 requested families | -| `diffusion_models/minimax_h3_{fl2va,ref2va}_pruned_fp8_scaled.safetensors` | 21 GB each | a 5th family — same size class as pruned int8 | -| `text_encoders/qwen3vl_32b_minimax_h3_bf16.safetensors` | 52 GB | every official template uses the nvfp4_awq TE; bf16 TE would also not fit the volume quota | -| `text_encoders/qwen3vl_32b_minimax_h3_int8_convrot.safetensors` | 27 GB | same | +| File | Size | Why skipped | +|----------------------------------------------------------------------------|------------|--------------------------------------------------------------------------------------------| +| `diffusion_models/minimax_h3_{fl2va,ref2va}_pruned_bf16.safetensors` | 40 GB each | a 4th family (pruned, unquantized) — not in the 3 requested families | +| `diffusion_models/minimax_h3_{fl2va,ref2va}_pruned_fp8_scaled.safetensors` | 21 GB each | a 5th family — same size class as pruned int8 | +| `text_encoders/qwen3vl_32b_minimax_h3_bf16.safetensors` | 52 GB | every official template uses the nvfp4_awq TE; bf16 TE would also not fit the volume quota | +| `text_encoders/qwen3vl_32b_minimax_h3_int8_convrot.safetensors` | 27 GB | same | ## History / gotchas