5 Commits

Author SHA1 Message Date
Hanashi 77524a10e6 final: 5090 tables (incl. bf16-streaming), stats tooling covers all 5 datasets
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 19:21:16 -04:00
Hanashi a293d0a2c4 reliability: failure-stats + output-validation tooling; persistent output dir
- bench/failures.py: per-GPU success rates and error-class breakdown by config
- bench/validate_outputs.py: ffprobe every recorded output via /view
- finding: 24GB cards fail on multi-ref (TE vision OOM), not duration/size;
  100% of successful runs produce valid videos
- outputs now written to /workspace (survives pod resets)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 14:41:00 -04:00
Hanashi e90f64b30e bench: results summarizer (mean±std tables); rebuild CSVs from authoritative JSONL
The CSV header predated the ref_count column so later rows were misaligned;
JSONL rows were always complete.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 13:51:17 -04:00
Hanashi 548e4efcaf benchmarking info 2026-08-06 01:49:26 -04:00
Hanashi 225aa97620 MiniMax H3 on RunPod: install/download scripts, ComfyUI wiring, benchmark harness, research docs
- scripts/pod: weight downloads (workspace + tmpfs overflow for volume quota),
  extra_model_paths.yaml, headless ComfyUI launcher, test image generator,
  one-shot pod installer
- scripts/local: auto-reconnecting port forward for the ComfyUI panel
- bench: API-based harness for t2v/i2v/ref2v across durations, resolutions,
  and weight families; CSV/JSONL results
- docs: WEIGHTS.md (what's downloaded where), RESEARCH.md (ComfyUI guides,
  no-Comfy options via SGLang/vLLM/diffusers, concurrency model, serverless)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 00:28:16 -04:00