pragent
An AI pull-request reviewer for Gitea that posts inline comments with suggested fixes, not a wall of prose — and reports what each review cost.
A webhook wakes for any PR on a repo whose default branch carries a
.pr-review.json with "enabled": true. The service checks the repo out at the
PR's head commit, reads the changed files and the code around them, runs the
repo's own linters, and posts a review anchored to real lines.
See it work: pragent-demo PR #1 — a PR with
planted defects, and the review it drew: 9 findings, 3 critical, all anchored
inline.
**[CRITICAL]** search_notes builds its SQL by string concatenation: owner and
term come from request['query'] and are spliced directly, so a term like
`' OR 1=1 --` reads every row in the table.
Fix: Parameterise owner and term with placeholders and a real LIKE pattern.
🪙 ~362 tok (11% · attributed output)
Why this exists
CodeRabbit, Greptile and Qodo are good products with fixed review dimensions and
per-seat pricing. pragent targets the case where a platform team needs to add
its own dimensions — an internal compliance rule, a service-catalog ownership
check, a house performance idiom — without forking a vendor's reviewer. The
analyzers, the data, and the analytics are yours.
It is also self-hosted end to end: the model endpoint is a config value, so the code never has to leave your network.
Status
A pilot is live and reviewing real PRs. The current runtime architecture is
documented in docs/architecture.md; older framework
plans remain in docs/plans/ as historical design material.
What works today:
- a central webhook service, so onboarding a repo is add the bot + commit
.pr-review.json:enabled = true - whole-repo context: the reviewer reads callers and types, not just the hunk
- inline comments with language-highlighted suggested fixes, anchored to post-change lines and validated in Python before posting
- per-commit dedupe, and prior reviews fed back so a re-push synthesises rather than repeats
.pr-review.jsonfor per-repo focus and house rules (also the opt-in flag)- token-usage reporting on every review, measured from opencode
step_finishevents - Langfuse traces, equivalent-cost reporting, evaluation scores, and feedback harvesting
- containment against hostile PR content (see Security)
Not yet: status checks, fail-close, attention tiering enforced in code (it is currently a skill the agent follows), multi-model routing, and a CLI framework.
How a review runs
PR opened on repo with `.pr-review.json:enabled = true`
│ Gitea webhook (HMAC-verified, body-capped, concurrency-bounded)
▼
review_pr()
1. opt-in .pr-review.json:enabled=true on base? if not, skip.
2. dedupe already reviewed this exact sha? stop.
3. fetch diff + .pr-review.json from the BASE branch
4. checkout repo archive at head sha → temp workdir
5. sanitize delete author-controlled agent-instruction files
6. brief .pragent/brief.md, untrusted parts explicitly fenced
7. review opencode agent: read code, run linters, emit findings JSON
8. anchor validate every line against the diff's post-change lines
9. post inline comments + summary, as pragent-bot
Steps 1, 3, 8 and 9 are deterministic Python. The model's only job is step 7 — producing correct findings. It never talks to Gitea, and a finding whose line does not validate becomes a summary bullet rather than a misplaced comment.
Setup
Onboarding a repo, once the service is running for that owner:
- add
pragent-botas a Write collaborator - commit
.pr-review.json: {"enabled": true}to the repo's default branch - open a PR
Standing up the service itself — the webhook, the image, the Gitea SSRF
allow-list, the per-owner webhook registration — is in
pilot/README-webhook.md. A legacy per-repo CI-step
path is in pilot/README.md.
The model endpoint is supplied at runtime via PRAGENT_MODEL_BASE_URL; the
committed opencode.json carries a placeholder.
Per-review token spend, latency, equivalent cost, and evaluation scores are
shipped to a self-hosted Langfuse: pilot/README-langfuse.md.
Emission is a silent no-op unless LANGFUSE_HOST and the key pair are set.
Extending it
The review "factory" is .opencode/ — agent definitions
and skills as plain Markdown. Adding a review dimension is dropping a file in,
not writing code:
| Add | How |
|---|---|
| A review lens | .opencode/agents/<name>.md + one allow-list line in pragent.md |
| Domain knowledge | .opencode/skills/<name>/SKILL.md, referenced from the load table |
| Per-repo rules | .pr-review.json in the repo being reviewed |
Shipped skills: attention-tiering (the cost governor), review-methodology,
findings-schema, linter-playbook, security-lens, malicious-change,
comment-craft.
Security
The reviewer runs an autonomous agent with shell access over a checkout of the PR author's branch, and its bot account holds a Write credential. Anyone who can open a PR can therefore put arbitrary text in front of the model and arbitrary files on its disk — the setup exploited in the April 2026 disclosures against Claude Code Security Review, Gemini CLI Action and Copilot Agent.
Four controls, none of which rely on the model behaving:
- No credentials in the agent's environment. The subprocess environment is built from an allow-list, not inherited. There is nothing to exfiltrate.
- No author-controlled instruction files on disk. Nested
AGENTS.md,CLAUDE.md,.cursorrules, a repoopencode.json— all deleted before the agent starts, so a PR cannot ship its own system prompt. They are still reviewed, as data. - Untrusted-data framing. PR text and diffs are fenced; the agent reports
injection attempts as
criticalfindings instead of following them. - Reviewer config comes from the base branch, so a PR cannot rewrite the rules it is judged by.
Plus: tar-slip guards on the archive, a non-root container, and bounded
concurrency. Full threat model and residual risks: pilot/README-webhook.md.
What it costs
The pilot runs against a self-hosted model and bills nothing per token, but the
token work is real. pilot/cost_model.py prices it against published API
rates, calibrated against runs measured through the usage telemetry
(OBSERVED_RUNS in that file — append to it, don't guess). Tokens are summed
from opencode step_finish events per review.
Each review is governed by hard limits: 20 completed steps, 120,000 total
tokens, 20,000 output tokens, and 480 seconds by default. Limits can be
overridden deployment-wide with PRAGENT_MAX_REVIEW_STEPS,
PRAGENT_MAX_REVIEW_TOKENS, PRAGENT_MAX_REVIEW_OUTPUT_TOKENS, and
PRAGENT_REVIEW_TIMEOUT, or per repository in the trusted base-branch
.pr-review.json; repository values win.
For repositories with broad diffs, the pilot automatically raises headroom to 40/400K, 60/800K, or 80/1.2M steps/tokens as changed lines cross 200, 800, or 2,000. Explicit repository budgets always take precedence, and the global hard ceilings remain in force.
Two measured reviews of a ~1100-line PR in this repo: 28 and 31 agent steps, ~2.1M input tokens each, zero cache reads or writes. The demo repo's PR, same tier: 126K tokens.
The measurements above are historical uncapped runs. A capped run preserves completed output, reports the cap reason in the review, and records it in Langfuse.
| Model | this repo, ~1100-line PR | demo repo PR |
|---|---|---|
| Claude Opus 5 | ~$10.79 | ~$0.71 |
| Claude Sonnet 5 | ~$4.32 | ~$0.28 |
| Claude Haiku 4.5 | ~$2.16 | ~$0.14 |
Three things that estimate wrong if you skip them:
- The loop resends its context every step. Cost is roughly quadratic in step
count, not linear in diff size. This is what
attention-tieringexists to cap. - Repository size dominates diff size. The 16x gap above is the same reviewer on the same tier — the difference is how much repo there was to read.
- Prompt caching is worth about a third of the bill and is not currently
happening on this stack. Check
cache_readbefore budgeting.
python3 pilot/cost_model.py --help # other mixes, volumes, models
Development
python3 -m pytest tests -q # stdlib-only tests, no network
The pilot is stdlib-only Python by design — it runs from a bare python:slim
image with the scripts mounted, and has no dependency resolution to go wrong at
review time.
License
Not yet chosen. Until one is added, no reuse rights are granted.
pilot eval judges test 1788201461