Commit Graph

14 Commits

Author SHA1 Message Date
Hanashi e375345bf6 SERVING-PLAN: per-output-second cost and pricing at all margins
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 20:08:03 -04:00
Hanashi 111f748c9d SERVING-PLAN.md: paid/free pools, pricing at 0-70% margins, capacity scenarios
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 20:03:21 -04:00
Hanashi 77524a10e6 final: 5090 tables (incl. bf16-streaming), stats tooling covers all 5 datasets
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 19:21:16 -04:00
Hanashi 056d5785e7 MITIGATIONS.md: 24GB multi-ref RESCUED — full envelope verdicts per machine class
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 19:15:36 -04:00
Hanashi 9fc70a0b81 MITIGATIONS.md: per-machine-class OOM map — RAM cap decides mitigation burden
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 17:09:40 -04:00
Hanashi a4bd386e52 BENCHMARKS.md: final cross-GPU tables, failure stats, multi-ref rescue status
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 15:23:36 -04:00
Hanashi a293d0a2c4 reliability: failure-stats + output-validation tooling; persistent output dir
- bench/failures.py: per-GPU success rates and error-class breakdown by config
- bench/validate_outputs.py: ffprobe every recorded output via /view
- finding: 24GB cards fail on multi-ref (TE vision OOM), not duration/size;
  100% of successful runs produce valid videos
- outputs now written to /workspace (survives pod resets)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 14:41:00 -04:00
Hanashi fbc68bb749 concurrency experiment measured: co-scheduling 13% slower than sequential — one worker per GPU confirmed
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 14:27:32 -04:00
Hanashi 24134b9653 BENCHMARKS.md: B200-complete results — bf16 fastest family, ref scaling, cross-GPU interim
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 14:16:13 -04:00
Hanashi 94a94a268f docs: concurrency & cost-per-video report across GPU tiers
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 02:28:24 -04:00
Hanashi c3cda1fb30 WEIGHTS.md: all 9 files persistent on 1TB volume, tmpfs scheme retired
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 02:12:42 -04:00
Hanashi 23a755e394 docs: pad markdown table columns for raw-text readability
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 00:38:51 -04:00
Hanashi 9682b0a084 WEIGHTS.md: final download status — all 9 files present, none in progress
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 00:37:44 -04:00
Hanashi 225aa97620 MiniMax H3 on RunPod: install/download scripts, ComfyUI wiring, benchmark harness, research docs
- scripts/pod: weight downloads (workspace + tmpfs overflow for volume quota),
  extra_model_paths.yaml, headless ComfyUI launcher, test image generator,
  one-shot pod installer
- scripts/local: auto-reconnecting port forward for the ComfyUI panel
- bench: API-based harness for t2v/i2v/ref2v across durations, resolutions,
  and weight families; CSV/JSONL results
- docs: WEIGHTS.md (what's downloaded where), RESEARCH.md (ComfyUI guides,
  no-Comfy options via SGLang/vLLM/diffusers, concurrency model, serverless)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 00:28:16 -04:00