13 Commits

Author SHA1 Message Date
Hanashi 77524a10e6 final: 5090 tables (incl. bf16-streaming), stats tooling covers all 5 datasets
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 19:21:16 -04:00
Hanashi a88626c0ca bench: refmatrix suite — ref count (1-9) x ref resolution ladder
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 16:34:35 -04:00
Hanashi 4c2635515b bench: --ref-downscale option — pre-shrink refs before the node (24GB OOM mitigation)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 14:44:54 -04:00
Hanashi a293d0a2c4 reliability: failure-stats + output-validation tooling; persistent output dir
- bench/failures.py: per-GPU success rates and error-class breakdown by config
- bench/validate_outputs.py: ffprobe every recorded output via /view
- finding: 24GB cards fail on multi-ref (TE vision OOM), not duration/size;
  100% of successful runs produce valid videos
- outputs now written to /workspace (survives pod resets)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 14:41:00 -04:00
Hanashi e90f64b30e bench: results summarizer (mean±std tables); rebuild CSVs from authoritative JSONL
The CSV header predated the ref_count column so later rows were misaligned;
JSONL rows were always complete.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 13:51:17 -04:00
Hanashi 53b4a133e2 bench: survive server loss during polling (mark run failed, keep suite alive)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 03:48:37 -04:00
Hanashi 1debde7673 bench: v2 reference set — JP stills, Ashitaka/Bebop, forests, volcanos (up to 23k px)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 02:45:34 -04:00
Hanashi cc111c0c33 bench: real reference images (anime/film/panorama) replace synthetic gradients
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 02:10:27 -04:00
Hanashi 4a802d8c3e bench: key-config suite with N=3 repeats for statistical averages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 02:07:28 -04:00
Hanashi 3eb04b254b remove accidentally committed pycache; ignore *.pyc/*.pyo
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 02:03:20 -04:00
Hanashi 548e4efcaf benchmarking info 2026-08-06 01:49:26 -04:00
Hanashi 52f8a79f35 bench: multi-reference ref2v support (1-8 ref images, match/max sizing)
- refheavy suite: ref2v saturation with 1/4/8 refs, the priority modality
- family suite: reduced grid for int8/bf16 family comparisons
- 8 distinct 1344x768 reference scenes generated into /ComfyUI/input

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 00:45:49 -04:00
Hanashi 225aa97620 MiniMax H3 on RunPod: install/download scripts, ComfyUI wiring, benchmark harness, research docs
- scripts/pod: weight downloads (workspace + tmpfs overflow for volume quota),
  extra_model_paths.yaml, headless ComfyUI launcher, test image generator,
  one-shot pod installer
- scripts/local: auto-reconnecting port forward for the ComfyUI panel
- bench: API-based harness for t2v/i2v/ref2v across durations, resolutions,
  and weight families; CSV/JSONL results
- docs: WEIGHTS.md (what's downloaded where), RESEARCH.md (ComfyUI guides,
  no-Comfy options via SGLang/vLLM/diffusers, concurrency model, serverless)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 00:28:16 -04:00