Commit Graph

2 Commits

Author SHA1 Message Date
Hanashi a293d0a2c4 reliability: failure-stats + output-validation tooling; persistent output dir
- bench/failures.py: per-GPU success rates and error-class breakdown by config
- bench/validate_outputs.py: ffprobe every recorded output via /view
- finding: 24GB cards fail on multi-ref (TE vision OOM), not duration/size;
  100% of successful runs produce valid videos
- outputs now written to /workspace (survives pod resets)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 14:41:00 -04:00
Hanashi e90f64b30e bench: results summarizer (mean±std tables); rebuild CSVs from authoritative JSONL
The CSV header predated the ref_count column so later rows were misaligned;
JSONL rows were always complete.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 13:51:17 -04:00