chore(cost-model): record the second measured review
31 steps, 190s, 2,213,077 in / 9,058 out, cache 0/0 — within 7% of the first run's input on the same tier, which is the first evidence the recalibrated tier defaults hold rather than fitting a single point. Also weakens the step-cap hypothesis from the previous commit: this run used 31 steps (more than the 28 that succeeded, and more than the run that failed) and parsed fine, so hitting `steps: 40` is not on its own what breaks the output format. Leaving the cap alone until the stderr logging catches a real failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B11e8TZZxJyzHW7jj7KWUN
This commit is contained in:
@@ -188,6 +188,22 @@ OBSERVED_RUNS: list[dict] = [
|
||||
"cache_write": 0,
|
||||
"subagents": 0,
|
||||
},
|
||||
{
|
||||
"label": "gitea_admin/pragent#7 (+ cost-model calibration + salvage fix)",
|
||||
"date": "2026-08-18",
|
||||
"tier": "full",
|
||||
"diff_tokens": 21_000, # same PR, two commits later
|
||||
"steps": 31,
|
||||
"duration_s": 189.8,
|
||||
"input": 2_213_077,
|
||||
"output": 9_058,
|
||||
"cache_read": 0,
|
||||
"cache_write": 0,
|
||||
"subagents": 0,
|
||||
},
|
||||
# A third run of the same PR (sha 2613b3e, 31 steps' worth of work in 330s)
|
||||
# ended without a parseable findings block and so reported no usage at all —
|
||||
# the reason `salvage_summary` now keeps the usage section on that path.
|
||||
]
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user