Files
2026-08-06 20:08:03 -04:00

252 lines
14 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# MiniMax H3 serving plan — pools, pricing, capacity scenarios
Prepared 2026-08-06. Every generation time in this document is **measured** on
our pods (see `docs/BENCHMARKS.md`); derived numbers state their formula.
GPU prices are what we actually pay: B200 $6.94/hr, RTX 5090 $0.99/hr,
RTX 4090 $0.74/hr. All jobs are 20 steps, 24 fps, template sampler settings.
Two structural facts drive everything:
1. **One generation saturates one GPU** (measured: co-scheduling is 13%
*slower*). Throughput = number of workers; latency = service time + queue.
2. **bf16 is the fastest family wherever it fits** — including, surprisingly,
layer-streamed on the RTX 5090 (110 s vs 143 s pruned for the standard
clip). Quantized families are for VRAM-constrained multi-model workers,
not for speed or cost.
## The product surface (what we can offer)
**Modalities**: text-to-video, keyframe-to-video (first and/or last frame),
reference-to-video (1–9 images; also up to 3 ref videos + 3 audio clips,
untested by us). All produce video + native stereo audio.
**Resolutions** (canvas must be multiples of 32, area ≤ ~1.03 MP):
| Tier | 16:9 | 9:16 | 1:1 | Notes |
|---------------|-----------|-----------|-----------|--------------------------|
| Draft (352p) | 608×352 | 352×608 | 448×448 | fastest, preview quality |
| SD (480p) | 864×480 | 480×864 | 640×640 | the sweet spot |
| HD (768p) | 1344×768 | 768×1344 | 992×992 | native max; ~3.3× cost |
Intermediate 16:9 rungs exist every 32 px (736×416, 960×544, 1056×608,
1152×640, 1216×672, 1280×736) — we recommend selling the three tiers above.
Other aspects (4:3, 21:9, …) work too — anything multiple-of-32 under the
area cap. 2K does not exist locally (hosted-API upscale only).
**Durations**: frame count must satisfy n ≡ 5 (mod 17); trained range 124–362
frames → **~5.2 s to ~15.1 s in ~0.7 s steps** (5.2, 5.9, 6.6, 7.3, 8.0, 8.7,
9.4, 10.1, 10.8, 11.5, 12.2, 12.9, 13.7, 14.4, 15.1). Sell "5 / 10 / 15
seconds"; internally snap up to the grid. fps is fixed at 24.
## Pool 1 — Paid (bf16, everything)
### Worker type
**B200 (180 GB) is the only worker class that runs the entire paid surface
on bf16** — HD, 15 s, and heavy reference jobs all need its VRAM. However,
55–70 % of paid traffic (≤480p, ≤10 s) runs **4× cheaper on 5090-bf16
workers** at identical weights/quality. Recommended deployment:
- **5090 bf16 pool** — handles t2v/i2v/ref2v at ≤480p, ≤10 s (dedicated
single-model workers; no model switching, so no cache-none tax).
- **B200 bf16 pool (small)** — handles HD, >10 s, and `max`-sizing reference
jobs; also absorbs overflow for latency.
Route by job shape. This hybrid cuts blended cost ~2.5–3× vs all-B200.
### Cost per video (B200 bf16, measured/derived)
t2v baseline; keyframe (i2v) ≈ +5 %; reference adder below.
| Resolution | 5 s | 10 s | 15 s |
|------------|------------------|------------------|------------------|
| Draft 352p | 26 s → $0.050 | 63 s → $0.122 | 111 s → $0.214 |
| SD 480p | 51 s → $0.098 | 125 s → $0.240 | 240 s → $0.463 |
| HD 768p | 169 s → $0.326 | 490 s → $0.945 | 965 s → $1.861 |
(10 s and 352p/HD-longer cells derived from measured duration/resolution
scaling on identical hardware; 480p-5/15 s, HD-5 s are directly measured.)
Reference-to-video adder (match sizing): **+$0.006 (1 ref) to +$0.054
(8–9 refs)** at ≤480p; **+$0.22** at HD with 8 refs. `max` reference sizing
multiplies job cost ~5–10× — if offered at all, price it as its own SKU at
~5× the base price, or restrict to ≤4 refs (+$0.08).
On the 5090-bf16 pool the same ≤480p jobs cost: 5 s $0.030 · 10 s ~$0.074 ·
1-ref +~$0.01. Blend accordingly; the pricing below uses **B200 costs** as
the conservative basis, so hybrid routing only improves your real margin.
### Pricing — per output-second (the unit to reason in)
All prices are **$ per second of delivered video**. Margin means gross margin
on price (price = cost ÷ (1 − margin)). Note the honest wrinkle: our cost per
output-second RISES with clip length (compute is superlinear in frames), so a
truly flat per-second price quietly gives long clips a discount — the tables
show exact per-duration cost so you can decide flat vs tiered.
**Our cost per output-second (B200 bf16 basis), and price at each margin:**
| Tier | Clip | Cost/s | 0% | 15% | 25% | 50% | 70% |
|-------|-------|----------|---------|---------|---------|---------|---------|
| 352p | 5 s | $0.0097 | $0.0097 | $0.0114 | $0.0129 | $0.0194 | $0.0323 |
| 352p | 10 s | $0.0120 | $0.0120 | $0.0141 | $0.0160 | $0.0240 | $0.0400 |
| 352p | 15 s | $0.0142 | $0.0142 | $0.0167 | $0.0189 | $0.0284 | $0.0473 |
| 480p | 5 s | $0.0190 | $0.0190 | $0.0224 | $0.0254 | $0.0380 | $0.0634 |
| 480p | 10 s | $0.0238 | $0.0238 | $0.0280 | $0.0317 | $0.0476 | $0.0793 |
| 480p | 15 s | $0.0307 | $0.0307 | $0.0361 | $0.0409 | $0.0614 | $0.1023 |
| 768p | 5 s | $0.0630 | $0.0630 | $0.0741 | $0.0840 | $0.1260 | $0.2101 |
| 768p | 10 s | $0.0932 | $0.0932 | $0.1097 | $0.1243 | $0.1865 | $0.3108 |
| 768p | 15 s | $0.1234 | $0.1234 | $0.1451 | $0.1645 | $0.2467 | $0.4112 |
**Hybrid-routing true cost** (≤480p/≤10 s jobs on 5090-bf16 workers):
480p drops to **$0.0059/s (5 s)** and **$0.0073/s (10 s)** — roughly 3×
cheaper than the B200 basis above. If you price from the B200 table and route
hybrid, realized margin is far above nominal (e.g. a "25 %" 480p price
carries ~77 % true margin on the 5090 share).
**Flat-rate menu option** (blended 60/25/15 duration mix, hybrid routing):
| Public flat price | True blended margin |
|---------------------|----------------------------------|
| SD 480p @ $0.02/s | ~50 % |
| SD 480p @ $0.03/s | ~67 % |
| HD 768p @ $0.10/s | ~20 % (HD stays on B200 — price HD at $0.13–0.17/s for 40–50 %) |
| Draft 352p @ $0.01/s | ~0 % on B200, ~60 % if drafts route to 5090s |
Per-video equivalents of the B200-basis table (for invoice display):
| Duration | Cost | 0% | 15% | 25% | 50% | 70% |
|----------|--------|--------|--------|--------|--------|--------|
| 5 s | $0.098 | $0.098 | $0.116 | $0.131 | $0.197 | $0.328 |
| 10 s | $0.240 | $0.240 | $0.283 | $0.320 | $0.481 | $0.801 |
| 15 s | $0.463 | $0.463 | $0.545 | $0.617 | $0.926 | $1.543 |
| SKU modifiers | ×cost | example @25% margin, 10 s |
|------------------------------|-------|--------------------------------|
| Draft 352p | ×0.51 | $0.163 |
| HD 768p | ×3.6+ | $1.260 (5 s: $0.434, 15 s: $2.481) |
| + references (≤9, match) | +$0.01–0.06 | round: +$0.05 |
| + references at HD | +$0.22 | round: +$0.30 |
| keyframe (first/last frame) | +5 % | fold into base |
Simple public menu suggestion (at ~50 % blended margin, knowing hybrid
routing makes true margin higher): **Draft $0.10 · SD $0.20/5 s · HD $0.65/5 s;
+$0.05 per clip with references; duration billed per 5 s block.**
## Pool 2 — Free (480p, 5 s, ≤3 cropped refs)
Constraint set: 864×480 (any of the three aspects), 5 s, up to 3 reference
images pre-cropped/downscaled (our `--ref-downscale 0.2` path — measured to
also *improve* stability and speed).
### Cost per 1,000 free generations (measured)
| Worker + weights | time/job | jobs/hr/worker | $/1k videos |
|-------------------------------|----------|----------------|-------------|
| **5090 bf16 (t2v/i2v only)** | 110 s | 32.7 | **$30** |
| 5090 pruned int8, t2v | 143 s | 25.2 | $39 |
| **5090 pruned int8, 3 refs** | ~176 s | 20.5 | **$48** |
| 4090 patched, 3 refs | ~400 s | 9.0 | $82 |
| B200 bf16, 3 refs | 60 s | 60.0 | $116 |
int8 (non-pruned) sits between pruned and bf16 in speed on every card and
fits nothing extra — it has no serving niche; skip it. **Recommendation:
free pool = 5090s running pruned int8 with the ref2va model resident**
(one model handles ref jobs; t2v works via 0-ref… no — t2v needs fl2va).
Practical split: ~⅓ of free workers hold fl2va (t2v/i2v jobs, $39/1k), ~⅔
hold ref2va (ref jobs, $48/1k) — no model switching, no cache-none tax,
blended **~$45 per 1,000 free videos**. bf16-on-5090 is competitive for
t2v-only but its ref2v carries the switch tax; revisit if free tier drops
references.
Quality note: pruned int8's output is strong (community: "near-lossless
lineage"), and free users get 480p/5 s — the quality delta vs bf16 at this
tier is minimal. If you want bf16 everywhere anyway, free costs rise ~25 %.
## Capacity & demand scenarios
Model: workers sized so that **peak-hour utilization ≤ 70 %** (queueing knee).
Peak hour assumed 3× the daily average rate. Wait ≈ service_time × ρ/(1−ρ)
per M/M/1 worker; at 70 % that's ~2.3× service time queued ahead of you —
the sizing keeps typical waits under ~1 job-length even at peak.
Wait-time intuition (176 s free ref job):
| Utilization | Avg queue wait |
|-------------|----------------|
| 50 % | ~3 min |
| 70 % | ~7 min |
| 85 % | ~17 min |
| 95 % | ~56 min |
### Free pool (5090 pruned, blended 165 s/job, 21.8 jobs/hr/worker)
| Demand (videos/day) | Peak rate/hr | Workers (70 % peak) | $/day (24/7) | $/day (autoscaled*) | Mean-hour wait | Peak wait |
|---------------------|--------------|----------------------|--------------|----------------------|----------------|-----------|
| 1,000 | 125 | 9 | $214 | ~$95 | <1 min | ~6 min |
| 5,000 | 625 | 41 | $974 | ~$430 | <1 min | ~6 min |
| 20,000 | 2,500 | 164 | $3,897 | ~$1,700 | <1 min | ~6 min |
| 100,000 | 12,500 | 820 | $19,483 | ~$8,600 | <1 min | ~6 min |
*Autoscaled ≈ 44 % of 24/7 (integrate a 3:1 peak:trough sinusoidal day at
70 % target util). Requires image-baked or volume-attached workers that cold
start in ~4 min (measured: container rebuild ~4 min + first model load
~40 s on a warm volume).
Rule of thumb: **free tier costs ~$45 per 1,000 videos served, i.e.
~$0.0087 per output-second** — the worker math above just determines how
fast you serve them. (For comparison: every paid SD second sold at $0.02
funds ~2.3 free seconds.)
### Paid pool (hybrid; blended job assumed 60 % SD-5s / 20 % SD-10s / 10 % HD-5s / 10 % ref-heavy)
Blended: ~$0.11 cost per job on hybrid routing (~$0.16 all-B200), ~150 s
blended service time on the 5090 share, 84 s on the B200 share.
| Demand (jobs/day) | 5090 workers | B200 workers | $/day (24/7) | Revenue/day @25 % margin | @50 % |
|--------------------|--------------|--------------|--------------|---------------------------|--------|
| 500 | 3 | 1 | $238 | $73 rev / $55 cost* | $110 |
| 2,500 | 12 | 2 | $618 | $367 / $275 | $550 |
| 10,000 | 46 | 5 | $1,926 | $1,467 / $1,100 | $2,200 |
| 50,000 | 229 | 22 | $9,105 | $7,333 / $5,500 | $11,000 |
*Cost column is per-job cost × volume (what you actually burn on GPU-seconds);
the $/day worker column is capacity cost at 24/7 — the gap is idle headroom,
recovered by autoscaling (×0.44) or by letting free-tier jobs soak idle paid
workers (recommended: one queue, two priorities — paid preempts free).
**Break-even insight**: at 25 % margin the paid pool only covers its 24/7
capacity above ~4,000 jobs/day; below that, either autoscale, raise margin,
or (best) run paid and free as one fleet with priority scheduling so paid
headroom serves free demand instead of idling.
## Concrete deployment recommendation
1. **One fleet, two queues** (paid priority, free backfill) on **RTX 5090
workers in one region** sharing a network volume: bf16-fl2va workers for
paid t2v/i2v ≤480p/≤10 s, pruned-ref2va workers for all reference jobs
and free tier.
2. **Small B200 pool** (1–2 + burst) for HD / 15 s / max-sizing paid SKUs.
3. Config per `docs/MITIGATIONS.md`: 5090s need the #15316 patch +
`--reserve-vram 4`; keep one model per worker (no switching) to avoid
every RAM-cap failure mode we found.
4. Autoscale on queue depth; workers cold-start in ~5 min from a shared
volume (bake the container image to cut pip-install; then ~1 min).
5. Watch for ComfyUI merging #15316 and the KJNodes SageAttention patch
(~2× potential speedup, currently broken for H3) — either would shift
every number in this doc favorably.
## Assumptions register (audit before betting real money)
- GPU prices: our current RunPod rates; community-cloud 5090s can be ~30 %
cheaper, secure-cloud availability at scale unverified (provider/worker-
availability research is a pending follow-up).
- 10 s and 352p/HD-long cells derived from measured scaling laws, not
directly timed on bf16; ±15 % error bars.
- 3-ref free-tier time interpolated between measured 1-ref and 4-ref runs.
- Peak:average = 3:1 assumed; measure your real diurnal curve and re-run
the sizing (formulas inline above).
- No egress/storage/queue-infra costs included (small: ~1–3 MB per video).
- Quality parity of pruned int8 for the free tier is a judgment call —
A/B the actual outputs (they're all on the pods' volumes).