Group review, feedback, evaluation, observability, and entrypoint code into packages. Keep thin top-level compatibility shims for existing scripts and imports, and mirror the structure in the tests.
Three changes from operator feedback:
1. Per-comment � attribution restored on inline comments (operator wants
it back — the PR-level collapsible is collapsed by default, so the
attribution is the visible signal of per-finding cost share).
Hidden only when no _tok_attrib was computed (legacy callers / ollama
path without usage metering).
2. Agent prompt now bounds reads beyond the diff — the single biggest
driver of input-token bloat on long agent loops:
* ≤ 5 file reads beyond the diff for the entire review
* ≤ 80 lines per read (use --offset + --limit)
* ≤ 3 grep calls beyond the diff (prefer rtk grep)
* no re-reads of files already seen
* no directory walks (ls -R, find .)
* honor .pr-review.json:exclude_paths
3. De-generalize cost_model calibration labels. The OBSERVED_RUNS list
referred to `gitea_admin/pragent#7` — a real internal repo path that
blocks commercialization. Replaced with `internal/hardening-PR (16
files, 1020 insertions / 91 deletions)`. The numbers (input/output
tokens, steps, duration) are unchanged — only the labels are
generic.
Tests:
* test_inline_comment_body_with_attribution_line — asserts 🪙 line
shows when _tok_attrib is set
* test_inline_comment_body_no_attribution_no_coin_line — still
verifies the line is hidden when no attribution data
* test_observed_report_prices_every_model — asserts no internal
repo name appears in the rendered report
Co-Authored-By: Claude <noreply@anthropic.com>
31 steps, 190s, 2,213,077 in / 9,058 out, cache 0/0 — within 7% of the first
run's input on the same tier, which is the first evidence the recalibrated tier
defaults hold rather than fitting a single point.
Also weakens the step-cap hypothesis from the previous commit: this run used 31
steps (more than the 28 that succeeded, and more than the run that failed) and
parsed fine, so hitting `steps: 40` is not on its own what breaks the output
format. Leaving the cap alone until the stderr logging catches a real failure.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B11e8TZZxJyzHW7jj7KWUN
PR #7 ran under the AI-USAGE label and reported real numbers: 28 agent steps,
348s, 2,071,025 input / 17,303 output tokens, and zero cache reads or writes.
The model predicted ~$0.73 on Opus 5 for that tier. The measurement prices it
at $10.79 — the model was ~15x low.
Two wrong assumptions:
- Step count and per-step growth. `full` assumed 12 steps and 1,200 tokens per
tool result; the run did 28 steps averaging ~3,300. Cost is roughly quadratic
in steps, so this compounds. Tier defaults are re-derived from the measured
per-step growth rather than from guesses.
- Caching. The model defaulted to prompt caching on. The headroom/glm-5.2 path
reports 0 read / 0 write, so the stable prefix is paid at full input price on
every step. Budget with caching off until that column is nonzero.
Adds OBSERVED_RUNS as an append-only calibration anchor, an observed-runs
section in the report, and a regression test asserting the model stays within
2.5x of the measurement — so the next drift is caught by the suite rather than
by a surprising invoice.
Corrected blended figures at 350 PRs/month: ~$1,740 Opus 5, ~$1,755 GPT-5.6
Sol, ~$696 Sonnet 5, ~$348 Haiku 4.5, ~$70 GPT-5.6 Luna.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B11e8TZZxJyzHW7jj7KWUN
Skills — the primary now loads conditionally (each one is input tokens), per a
load table in pragent.md:
- attention-tiering: classify every PR trivial/lite/full/oversized BEFORE
reading anything, and cap file reads, linter runs and subagent fan-out per
tier. This is the cost governor; the other skills defer to its budget.
- linter-playbook: per-ecosystem detect-and-run commands scoped to changed
files, the never-install rule, and how to turn a diagnostic into a finding
instead of pasting tool output.
- security-lens: the inline security checklist for when @security isn't worth
delegating, built around a source -> sink test each finding must pass.
- malicious-change: hostile-PR detection — injection aimed at the reviewer,
install/CI-time hooks, obfuscated payloads, dependency confusion, logic
backdoors. Complements the runtime containment added in the previous commit:
that stops the agent being hijacked, this makes it report the attempt.
- comment-craft: how to write problem/fix/suggestion so a maintainer can act in
one read, and what to cut.
pilot/cost_model.py — prices a review against published Claude and OpenAI rates
(fetched 2026-08-18). Prompt sizes are measured from the factory files rather
than guessed; per-tier workloads come from the tiering budgets. The model is
explicit about the thing that actually dominates an agent loop: the whole
conversation is resent every step, so caching moves ~2.3x of the bill.
Blended over a 5/35/55/5 mix with caching on: ~$0.61/PR on Opus 5 or GPT-5.6
Sol, ~$0.24 on Sonnet 5 or Terra, ~$0.12 on Haiku 4.5, ~$0.02 on Luna. At 350
PRs/month that's ~$212 / ~$85 / ~$43 / ~$8.50.
Tests: 101 -> 122.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B11e8TZZxJyzHW7jj7KWUN