Commit Graph

5 Commits

Author SHA1 Message Date
Hanashi a293d0a2c4 reliability: failure-stats + output-validation tooling; persistent output dir
- bench/failures.py: per-GPU success rates and error-class breakdown by config
- bench/validate_outputs.py: ffprobe every recorded output via /view
- finding: 24GB cards fail on multi-ref (TE vision OOM), not duration/size;
  100% of successful runs produce valid videos
- outputs now written to /workspace (survives pod resets)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 14:41:00 -04:00
Hanashi 427f2878c4 start-comfyui: per-host logs (shared volumes!), COMFY_ARGS passthrough
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 03:31:45 -04:00
Hanashi 58b4755475 start-comfyui: wait out mid-shutdown process before already-running check
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 03:23:09 -04:00
Hanashi 64e488fa49 start-comfyui: nohup fallback when tmux unavailable; parameterize port-forward
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 03:04:28 -04:00
Hanashi 225aa97620 MiniMax H3 on RunPod: install/download scripts, ComfyUI wiring, benchmark harness, research docs
- scripts/pod: weight downloads (workspace + tmpfs overflow for volume quota),
  extra_model_paths.yaml, headless ComfyUI launcher, test image generator,
  one-shot pod installer
- scripts/local: auto-reconnecting port forward for the ComfyUI panel
- bench: API-based harness for t2v/i2v/ref2v across durations, resolutions,
  and weight families; CSV/JSONL results
- docs: WEIGHTS.md (what's downloaded where), RESEARCH.md (ComfyUI guides,
  no-Comfy options via SGLang/vLLM/diffusers, concurrency model, serverless)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 00:28:16 -04:00