Adds the evaluation layer on top of the review traces: five deterministic
scores describing how the reviewer behaved, a bridge that turns human reactions
into ground truth, and a dataset seeded from the reviews already run.
The two are kept apart on purpose. feedback.db has recorded 113 reviews and
zero reactions, resolutions or replies — nobody has ever responded to a bot
comment — so an accuracy metric cannot be built yet. The scorers therefore
measure behaviour, which is computable from data in hand, and feedback_scores
turns verdicts into scores the moment any arrive.
eval_scores.py emits finding_rate, severity_info_ratio, severity_max,
dropped_findings and cost_per_finding into the same ingestion batch as the
trace. Undefined values are omitted rather than reported as zero: an info ratio
over a silent review is undefined, and charting it as 0 would read as perfect
calibration.
dropped_findings needed a parser change. Both parsers silently discard findings
with an unusable path/line, which made a model emitting garbage locations
indistinguishable from one that found nothing. last_parse_dropped() exposes the
delta, read at parse time — after apply_repo_config the drops are the config
working as intended, not the model misbehaving.
feedback_scores.py scores the session ("{repo}#{pr}"), because feedback arrives
days later against a PR and nothing records which re-run produced which
comment. review_acceptance is absent rather than 0 when nothing was engaged.
eval_bootstrap.py registers the score configs, seeds the pragent-reviews
dataset, and can backfill scores onto traces that predate the scorers.
expectedOutput is the reviewer's own prior output, flagged
labelled_by_human: false — a regression baseline, not verified truth.
Also fixes a silent telemetry failure: the ingestion endpoint answers 207 when
only some events succeed, so a batch with every event rejected still looked
like success. Score events were missing the required per-event timestamp and
ingested nothing while reporting 207. _warn_on_rejected_events now logs the
per-event errors under LANGFUSE_DEBUG.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
pragent
An AI pull-request reviewer for Gitea that posts inline comments with suggested fixes, not a wall of prose — and reports what each review cost.
A webhook wakes for any PR on a repo whose default branch carries a
.pr-review.json with "enabled": true. The service checks the repo out at the
PR's head commit, reads the changed files and the code around them, runs the
repo's own linters, and posts a review anchored to real lines.
See it work: pragent-demo PR #1 — a PR with
planted defects, and the review it drew: 9 findings, 3 critical, all anchored
inline.
**[CRITICAL]** search_notes builds its SQL by string concatenation: owner and
term come from request['query'] and are spliced directly, so a term like
`' OR 1=1 --` reads every row in the table.
Fix: Parameterise owner and term with placeholders and a real LIKE pattern.
🪙 ~362 tok (11% · attributed output)
Why this exists
CodeRabbit, Greptile and Qodo are good products with fixed review dimensions and
per-seat pricing. pragent targets the case where a platform team needs to add
its own dimensions — an internal compliance rule, a service-catalog ownership
check, a house performance idiom — without forking a vendor's reviewer. The
analyzers, the data, and the analytics are yours.
It is also self-hosted end to end: the model endpoint is a config value, so the code never has to leave your network.
Status
A pilot is live and reviewing real PRs. The full framework (pragent init,
tiering as code, analyzer fan-out, explain / replay) is designed but not
built — see docs/plans/.
What works today:
- a central webhook service, so onboarding a repo is add the bot + commit
.pr-review.json:enabled = true - whole-repo context: the reviewer reads callers and types, not just the hunk
- inline comments with language-highlighted suggested fixes, anchored to post-change lines and validated in Python before posting
- per-commit dedupe, and prior reviews fed back so a re-push synthesises rather than repeats
.pr-review.jsonfor per-repo focus and house rules (also the opt-in flag)- token-usage reporting on every review, measured from opencode
step_finishevents - containment against hostile PR content (see Security)
Not yet: status checks, fail-close, attention tiering enforced in code (it is currently a skill the agent follows), multi-model routing.
How a review runs
PR opened on repo with `.pr-review.json:enabled = true`
│ Gitea webhook (HMAC-verified, body-capped, concurrency-bounded)
▼
review_pr()
1. opt-in .pr-review.json:enabled=true on base? if not, skip.
2. dedupe already reviewed this exact sha? stop.
3. fetch diff + .pr-review.json from the BASE branch
4. checkout repo archive at head sha → temp workdir
5. sanitize delete author-controlled agent-instruction files
6. brief .pragent/brief.md, untrusted parts explicitly fenced
7. review opencode agent: read code, run linters, emit findings JSON
8. anchor validate every line against the diff's post-change lines
9. post inline comments + summary, as pragent-bot
Steps 1, 3, 8 and 9 are deterministic Python. The model's only job is step 7 — producing correct findings. It never talks to Gitea, and a finding whose line does not validate becomes a summary bullet rather than a misplaced comment.
Setup
Onboarding a repo, once the service is running for that owner:
- add
pragent-botas a Write collaborator - commit
.pr-review.json: {"enabled": true}to the repo's default branch - open a PR
Standing up the service itself — the webhook, the image, the Gitea SSRF
allow-list, the per-owner webhook registration — is in
pilot/README-webhook.md. A legacy per-repo CI-step
path is in pilot/README.md.
The model endpoint is supplied at runtime via PRAGENT_MODEL_BASE_URL; the
committed opencode.json carries a placeholder.
Per-review token spend, latency and equivalent cost are shipped to a
self-hosted Langfuse, split into ollama and claude environments so the two
spend stories stay separate: pilot/README-langfuse.md.
Emission is a silent no-op unless LANGFUSE_HOST and the key pair are set.
Extending it
The review "factory" is .opencode/ — agent definitions
and skills as plain Markdown. Adding a review dimension is dropping a file in,
not writing code:
| Add | How |
|---|---|
| A review lens | .opencode/agents/<name>.md + one allow-list line in pragent.md |
| Domain knowledge | .opencode/skills/<name>/SKILL.md, referenced from the load table |
| Per-repo rules | .pr-review.json in the repo being reviewed |
Shipped skills: attention-tiering (the cost governor), review-methodology,
findings-schema, linter-playbook, security-lens, malicious-change,
comment-craft.
Security
The reviewer runs an autonomous agent with shell access over a checkout of the PR author's branch, and its bot account holds a Write credential. Anyone who can open a PR can therefore put arbitrary text in front of the model and arbitrary files on its disk — the setup exploited in the April 2026 disclosures against Claude Code Security Review, Gemini CLI Action and Copilot Agent.
Four controls, none of which rely on the model behaving:
- No credentials in the agent's environment. The subprocess environment is built from an allow-list, not inherited. There is nothing to exfiltrate.
- No author-controlled instruction files on disk. Nested
AGENTS.md,CLAUDE.md,.cursorrules, a repoopencode.json— all deleted before the agent starts, so a PR cannot ship its own system prompt. They are still reviewed, as data. - Untrusted-data framing. PR text and diffs are fenced; the agent reports
injection attempts as
criticalfindings instead of following them. - Reviewer config comes from the base branch, so a PR cannot rewrite the rules it is judged by.
Plus: tar-slip guards on the archive, a non-root container, and bounded
concurrency. Full threat model and residual risks: pilot/README-webhook.md.
What it costs
The pilot runs against a self-hosted model and bills nothing per token, but the
token work is real. pilot/cost_model.py prices it against published API
rates, calibrated against runs measured through the usage telemetry
(OBSERVED_RUNS in that file — append to it, don't guess). Tokens are summed
from opencode step_finish events per review.
Two measured reviews of a ~1100-line PR in this repo: 28 and 31 agent steps, ~2.1M input tokens each, zero cache reads or writes. The demo repo's PR, same tier: 126K tokens.
| Model | this repo, ~1100-line PR | demo repo PR |
|---|---|---|
| Claude Opus 5 | ~$10.79 | ~$0.71 |
| Claude Sonnet 5 | ~$4.32 | ~$0.28 |
| Claude Haiku 4.5 | ~$2.16 | ~$0.14 |
Three things that estimate wrong if you skip them:
- The loop resends its context every step. Cost is roughly quadratic in step
count, not linear in diff size. This is what
attention-tieringexists to cap. - Repository size dominates diff size. The 16x gap above is the same reviewer on the same tier — the difference is how much repo there was to read.
- Prompt caching is worth about a third of the bill and is not currently
happening on this stack. Check
cache_readbefore budgeting.
python3 pilot/cost_model.py --help # other mixes, volumes, models
Development
python3 -m pytest tests -q # 137 tests, stdlib only, no network
The pilot is stdlib-only Python by design — it runs from a bare python:slim
image with the scripts mounted, and has no dependency resolution to go wrong at
review time.
License
Not yet chosen. Until one is added, no reuse rights are granted.