mirror of
https://github.com/storytold/cloud-worker.git
synced 2026-10-09 00:09:43 +00:00
e375345bf6
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
252 lines
14 KiB
Markdown
252 lines
14 KiB
Markdown
# MiniMax H3 serving plan — pools, pricing, capacity scenarios
|
||
|
||
Prepared 2026-08-06. Every generation time in this document is **measured** on
|
||
our pods (see `docs/BENCHMARKS.md`); derived numbers state their formula.
|
||
GPU prices are what we actually pay: B200 $6.94/hr, RTX 5090 $0.99/hr,
|
||
RTX 4090 $0.74/hr. All jobs are 20 steps, 24 fps, template sampler settings.
|
||
|
||
Two structural facts drive everything:
|
||
|
||
1. **One generation saturates one GPU** (measured: co-scheduling is 13%
|
||
*slower*). Throughput = number of workers; latency = service time + queue.
|
||
2. **bf16 is the fastest family wherever it fits** — including, surprisingly,
|
||
layer-streamed on the RTX 5090 (110 s vs 143 s pruned for the standard
|
||
clip). Quantized families are for VRAM-constrained multi-model workers,
|
||
not for speed or cost.
|
||
|
||
## The product surface (what we can offer)
|
||
|
||
**Modalities**: text-to-video, keyframe-to-video (first and/or last frame),
|
||
reference-to-video (1–9 images; also up to 3 ref videos + 3 audio clips,
|
||
untested by us). All produce video + native stereo audio.
|
||
|
||
**Resolutions** (canvas must be multiples of 32, area ≤ ~1.03 MP):
|
||
|
||
| Tier | 16:9 | 9:16 | 1:1 | Notes |
|
||
|---------------|-----------|-----------|-----------|--------------------------|
|
||
| Draft (352p) | 608×352 | 352×608 | 448×448 | fastest, preview quality |
|
||
| SD (480p) | 864×480 | 480×864 | 640×640 | the sweet spot |
|
||
| HD (768p) | 1344×768 | 768×1344 | 992×992 | native max; ~3.3× cost |
|
||
|
||
Intermediate 16:9 rungs exist every 32 px (736×416, 960×544, 1056×608,
|
||
1152×640, 1216×672, 1280×736) — we recommend selling the three tiers above.
|
||
Other aspects (4:3, 21:9, …) work too — anything multiple-of-32 under the
|
||
area cap. 2K does not exist locally (hosted-API upscale only).
|
||
|
||
**Durations**: frame count must satisfy n ≡ 5 (mod 17); trained range 124–362
|
||
frames → **~5.2 s to ~15.1 s in ~0.7 s steps** (5.2, 5.9, 6.6, 7.3, 8.0, 8.7,
|
||
9.4, 10.1, 10.8, 11.5, 12.2, 12.9, 13.7, 14.4, 15.1). Sell "5 / 10 / 15
|
||
seconds"; internally snap up to the grid. fps is fixed at 24.
|
||
|
||
## Pool 1 — Paid (bf16, everything)
|
||
|
||
### Worker type
|
||
|
||
**B200 (180 GB) is the only worker class that runs the entire paid surface
|
||
on bf16** — HD, 15 s, and heavy reference jobs all need its VRAM. However,
|
||
55–70 % of paid traffic (≤480p, ≤10 s) runs **4× cheaper on 5090-bf16
|
||
workers** at identical weights/quality. Recommended deployment:
|
||
|
||
- **5090 bf16 pool** — handles t2v/i2v/ref2v at ≤480p, ≤10 s (dedicated
|
||
single-model workers; no model switching, so no cache-none tax).
|
||
- **B200 bf16 pool (small)** — handles HD, >10 s, and `max`-sizing reference
|
||
jobs; also absorbs overflow for latency.
|
||
|
||
Route by job shape. This hybrid cuts blended cost ~2.5–3× vs all-B200.
|
||
|
||
### Cost per video (B200 bf16, measured/derived)
|
||
|
||
t2v baseline; keyframe (i2v) ≈ +5 %; reference adder below.
|
||
|
||
| Resolution | 5 s | 10 s | 15 s |
|
||
|------------|------------------|------------------|------------------|
|
||
| Draft 352p | 26 s → $0.050 | 63 s → $0.122 | 111 s → $0.214 |
|
||
| SD 480p | 51 s → $0.098 | 125 s → $0.240 | 240 s → $0.463 |
|
||
| HD 768p | 169 s → $0.326 | 490 s → $0.945 | 965 s → $1.861 |
|
||
|
||
(10 s and 352p/HD-longer cells derived from measured duration/resolution
|
||
scaling on identical hardware; 480p-5/15 s, HD-5 s are directly measured.)
|
||
|
||
Reference-to-video adder (match sizing): **+$0.006 (1 ref) to +$0.054
|
||
(8–9 refs)** at ≤480p; **+$0.22** at HD with 8 refs. `max` reference sizing
|
||
multiplies job cost ~5–10× — if offered at all, price it as its own SKU at
|
||
~5× the base price, or restrict to ≤4 refs (+$0.08).
|
||
|
||
On the 5090-bf16 pool the same ≤480p jobs cost: 5 s $0.030 · 10 s ~$0.074 ·
|
||
1-ref +~$0.01. Blend accordingly; the pricing below uses **B200 costs** as
|
||
the conservative basis, so hybrid routing only improves your real margin.
|
||
|
||
### Pricing — per output-second (the unit to reason in)
|
||
|
||
All prices are **$ per second of delivered video**. Margin means gross margin
|
||
on price (price = cost ÷ (1 − margin)). Note the honest wrinkle: our cost per
|
||
output-second RISES with clip length (compute is superlinear in frames), so a
|
||
truly flat per-second price quietly gives long clips a discount — the tables
|
||
show exact per-duration cost so you can decide flat vs tiered.
|
||
|
||
**Our cost per output-second (B200 bf16 basis), and price at each margin:**
|
||
|
||
| Tier | Clip | Cost/s | 0% | 15% | 25% | 50% | 70% |
|
||
|-------|-------|----------|---------|---------|---------|---------|---------|
|
||
| 352p | 5 s | $0.0097 | $0.0097 | $0.0114 | $0.0129 | $0.0194 | $0.0323 |
|
||
| 352p | 10 s | $0.0120 | $0.0120 | $0.0141 | $0.0160 | $0.0240 | $0.0400 |
|
||
| 352p | 15 s | $0.0142 | $0.0142 | $0.0167 | $0.0189 | $0.0284 | $0.0473 |
|
||
| 480p | 5 s | $0.0190 | $0.0190 | $0.0224 | $0.0254 | $0.0380 | $0.0634 |
|
||
| 480p | 10 s | $0.0238 | $0.0238 | $0.0280 | $0.0317 | $0.0476 | $0.0793 |
|
||
| 480p | 15 s | $0.0307 | $0.0307 | $0.0361 | $0.0409 | $0.0614 | $0.1023 |
|
||
| 768p | 5 s | $0.0630 | $0.0630 | $0.0741 | $0.0840 | $0.1260 | $0.2101 |
|
||
| 768p | 10 s | $0.0932 | $0.0932 | $0.1097 | $0.1243 | $0.1865 | $0.3108 |
|
||
| 768p | 15 s | $0.1234 | $0.1234 | $0.1451 | $0.1645 | $0.2467 | $0.4112 |
|
||
|
||
**Hybrid-routing true cost** (≤480p/≤10 s jobs on 5090-bf16 workers):
|
||
480p drops to **$0.0059/s (5 s)** and **$0.0073/s (10 s)** — roughly 3×
|
||
cheaper than the B200 basis above. If you price from the B200 table and route
|
||
hybrid, realized margin is far above nominal (e.g. a "25 %" 480p price
|
||
carries ~77 % true margin on the 5090 share).
|
||
|
||
**Flat-rate menu option** (blended 60/25/15 duration mix, hybrid routing):
|
||
|
||
| Public flat price | True blended margin |
|
||
|---------------------|----------------------------------|
|
||
| SD 480p @ $0.02/s | ~50 % |
|
||
| SD 480p @ $0.03/s | ~67 % |
|
||
| HD 768p @ $0.10/s | ~20 % (HD stays on B200 — price HD at $0.13–0.17/s for 40–50 %) |
|
||
| Draft 352p @ $0.01/s | ~0 % on B200, ~60 % if drafts route to 5090s |
|
||
|
||
Per-video equivalents of the B200-basis table (for invoice display):
|
||
|
||
| Duration | Cost | 0% | 15% | 25% | 50% | 70% |
|
||
|----------|--------|--------|--------|--------|--------|--------|
|
||
| 5 s | $0.098 | $0.098 | $0.116 | $0.131 | $0.197 | $0.328 |
|
||
| 10 s | $0.240 | $0.240 | $0.283 | $0.320 | $0.481 | $0.801 |
|
||
| 15 s | $0.463 | $0.463 | $0.545 | $0.617 | $0.926 | $1.543 |
|
||
|
||
| SKU modifiers | ×cost | example @25% margin, 10 s |
|
||
|------------------------------|-------|--------------------------------|
|
||
| Draft 352p | ×0.51 | $0.163 |
|
||
| HD 768p | ×3.6+ | $1.260 (5 s: $0.434, 15 s: $2.481) |
|
||
| + references (≤9, match) | +$0.01–0.06 | round: +$0.05 |
|
||
| + references at HD | +$0.22 | round: +$0.30 |
|
||
| keyframe (first/last frame) | +5 % | fold into base |
|
||
|
||
Simple public menu suggestion (at ~50 % blended margin, knowing hybrid
|
||
routing makes true margin higher): **Draft $0.10 · SD $0.20/5 s · HD $0.65/5 s;
|
||
+$0.05 per clip with references; duration billed per 5 s block.**
|
||
|
||
## Pool 2 — Free (480p, 5 s, ≤3 cropped refs)
|
||
|
||
Constraint set: 864×480 (any of the three aspects), 5 s, up to 3 reference
|
||
images pre-cropped/downscaled (our `--ref-downscale 0.2` path — measured to
|
||
also *improve* stability and speed).
|
||
|
||
### Cost per 1,000 free generations (measured)
|
||
|
||
| Worker + weights | time/job | jobs/hr/worker | $/1k videos |
|
||
|-------------------------------|----------|----------------|-------------|
|
||
| **5090 bf16 (t2v/i2v only)** | 110 s | 32.7 | **$30** |
|
||
| 5090 pruned int8, t2v | 143 s | 25.2 | $39 |
|
||
| **5090 pruned int8, 3 refs** | ~176 s | 20.5 | **$48** |
|
||
| 4090 patched, 3 refs | ~400 s | 9.0 | $82 |
|
||
| B200 bf16, 3 refs | 60 s | 60.0 | $116 |
|
||
|
||
int8 (non-pruned) sits between pruned and bf16 in speed on every card and
|
||
fits nothing extra — it has no serving niche; skip it. **Recommendation:
|
||
free pool = 5090s running pruned int8 with the ref2va model resident**
|
||
(one model handles ref jobs; t2v works via 0-ref… no — t2v needs fl2va).
|
||
Practical split: ~⅓ of free workers hold fl2va (t2v/i2v jobs, $39/1k), ~⅔
|
||
hold ref2va (ref jobs, $48/1k) — no model switching, no cache-none tax,
|
||
blended **~$45 per 1,000 free videos**. bf16-on-5090 is competitive for
|
||
t2v-only but its ref2v carries the switch tax; revisit if free tier drops
|
||
references.
|
||
|
||
Quality note: pruned int8's output is strong (community: "near-lossless
|
||
lineage"), and free users get 480p/5 s — the quality delta vs bf16 at this
|
||
tier is minimal. If you want bf16 everywhere anyway, free costs rise ~25 %.
|
||
|
||
## Capacity & demand scenarios
|
||
|
||
Model: workers sized so that **peak-hour utilization ≤ 70 %** (queueing knee).
|
||
Peak hour assumed 3× the daily average rate. Wait ≈ service_time × ρ/(1−ρ)
|
||
per M/M/1 worker; at 70 % that's ~2.3× service time queued ahead of you —
|
||
the sizing keeps typical waits under ~1 job-length even at peak.
|
||
|
||
Wait-time intuition (176 s free ref job):
|
||
|
||
| Utilization | Avg queue wait |
|
||
|-------------|----------------|
|
||
| 50 % | ~3 min |
|
||
| 70 % | ~7 min |
|
||
| 85 % | ~17 min |
|
||
| 95 % | ~56 min |
|
||
|
||
### Free pool (5090 pruned, blended 165 s/job, 21.8 jobs/hr/worker)
|
||
|
||
| Demand (videos/day) | Peak rate/hr | Workers (70 % peak) | $/day (24/7) | $/day (autoscaled*) | Mean-hour wait | Peak wait |
|
||
|---------------------|--------------|----------------------|--------------|----------------------|----------------|-----------|
|
||
| 1,000 | 125 | 9 | $214 | ~$95 | <1 min | ~6 min |
|
||
| 5,000 | 625 | 41 | $974 | ~$430 | <1 min | ~6 min |
|
||
| 20,000 | 2,500 | 164 | $3,897 | ~$1,700 | <1 min | ~6 min |
|
||
| 100,000 | 12,500 | 820 | $19,483 | ~$8,600 | <1 min | ~6 min |
|
||
|
||
*Autoscaled ≈ 44 % of 24/7 (integrate a 3:1 peak:trough sinusoidal day at
|
||
70 % target util). Requires image-baked or volume-attached workers that cold
|
||
start in ~4 min (measured: container rebuild ~4 min + first model load
|
||
~40 s on a warm volume).
|
||
|
||
Rule of thumb: **free tier costs ~$45 per 1,000 videos served, i.e.
|
||
~$0.0087 per output-second** — the worker math above just determines how
|
||
fast you serve them. (For comparison: every paid SD second sold at $0.02
|
||
funds ~2.3 free seconds.)
|
||
|
||
### Paid pool (hybrid; blended job assumed 60 % SD-5s / 20 % SD-10s / 10 % HD-5s / 10 % ref-heavy)
|
||
|
||
Blended: ~$0.11 cost per job on hybrid routing (~$0.16 all-B200), ~150 s
|
||
blended service time on the 5090 share, 84 s on the B200 share.
|
||
|
||
| Demand (jobs/day) | 5090 workers | B200 workers | $/day (24/7) | Revenue/day @25 % margin | @50 % |
|
||
|--------------------|--------------|--------------|--------------|---------------------------|--------|
|
||
| 500 | 3 | 1 | $238 | $73 rev / $55 cost* | $110 |
|
||
| 2,500 | 12 | 2 | $618 | $367 / $275 | $550 |
|
||
| 10,000 | 46 | 5 | $1,926 | $1,467 / $1,100 | $2,200 |
|
||
| 50,000 | 229 | 22 | $9,105 | $7,333 / $5,500 | $11,000 |
|
||
|
||
*Cost column is per-job cost × volume (what you actually burn on GPU-seconds);
|
||
the $/day worker column is capacity cost at 24/7 — the gap is idle headroom,
|
||
recovered by autoscaling (×0.44) or by letting free-tier jobs soak idle paid
|
||
workers (recommended: one queue, two priorities — paid preempts free).
|
||
|
||
**Break-even insight**: at 25 % margin the paid pool only covers its 24/7
|
||
capacity above ~4,000 jobs/day; below that, either autoscale, raise margin,
|
||
or (best) run paid and free as one fleet with priority scheduling so paid
|
||
headroom serves free demand instead of idling.
|
||
|
||
## Concrete deployment recommendation
|
||
|
||
1. **One fleet, two queues** (paid priority, free backfill) on **RTX 5090
|
||
workers in one region** sharing a network volume: bf16-fl2va workers for
|
||
paid t2v/i2v ≤480p/≤10 s, pruned-ref2va workers for all reference jobs
|
||
and free tier.
|
||
2. **Small B200 pool** (1–2 + burst) for HD / 15 s / max-sizing paid SKUs.
|
||
3. Config per `docs/MITIGATIONS.md`: 5090s need the #15316 patch +
|
||
`--reserve-vram 4`; keep one model per worker (no switching) to avoid
|
||
every RAM-cap failure mode we found.
|
||
4. Autoscale on queue depth; workers cold-start in ~5 min from a shared
|
||
volume (bake the container image to cut pip-install; then ~1 min).
|
||
5. Watch for ComfyUI merging #15316 and the KJNodes SageAttention patch
|
||
(~2× potential speedup, currently broken for H3) — either would shift
|
||
every number in this doc favorably.
|
||
|
||
## Assumptions register (audit before betting real money)
|
||
|
||
- GPU prices: our current RunPod rates; community-cloud 5090s can be ~30 %
|
||
cheaper, secure-cloud availability at scale unverified (provider/worker-
|
||
availability research is a pending follow-up).
|
||
- 10 s and 352p/HD-long cells derived from measured scaling laws, not
|
||
directly timed on bf16; ±15 % error bars.
|
||
- 3-ref free-tier time interpolated between measured 1-ref and 4-ref runs.
|
||
- Peak:average = 3:1 assumed; measure your real diurnal curve and re-run
|
||
the sizing (formulas inline above).
|
||
- No egress/storage/queue-infra costs included (small: ~1–3 MB per video).
|
||
- Quality parity of pruned int8 for the free tier is a judgment call —
|
||
A/B the actual outputs (they're all on the pods' volumes).
|