- bench/failures.py: per-GPU success rates and error-class breakdown by config
- bench/validate_outputs.py: ffprobe every recorded output via /view
- finding: 24GB cards fail on multi-ref (TE vision OOM), not duration/size;
100% of successful runs produce valid videos
- outputs now written to /workspace (survives pod resets)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The CSV header predated the ref_count column so later rows were misaligned;
JSONL rows were always complete.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>