# pragent An AI pull-request reviewer for Gitea that posts **inline comments with suggested fixes**, not a wall of prose โ€” and reports what each review cost. A webhook wakes for any PR on a repo whose default branch carries a `.pr-review.json` with `"enabled": true`. The service checks the repo out at the PR's head commit, reads the changed files *and the code around them*, runs the repo's own linters, and posts a review anchored to real lines. **See it work:** [`pragent-demo` PR #1](../pragent-demo/pulls/1) โ€” a PR with planted defects, and the review it drew: 9 findings, 3 critical, all anchored inline. ``` **[CRITICAL]** search_notes builds its SQL by string concatenation: owner and term come from request['query'] and are spliced directly, so a term like `' OR 1=1 --` reads every row in the table. Fix: Parameterise owner and term with placeholders and a real LIKE pattern. ๐Ÿช™ ~362 tok (11% ยท attributed output) ``` ## Why this exists CodeRabbit, Greptile and Qodo are good products with fixed review dimensions and per-seat pricing. `pragent` targets the case where a platform team needs to **add its own dimensions** โ€” an internal compliance rule, a service-catalog ownership check, a house performance idiom โ€” without forking a vendor's reviewer. The analyzers, the data, and the analytics are yours. It is also self-hosted end to end: the model endpoint is a config value, so the code never has to leave your network. ## Status A **pilot** is live and reviewing real PRs. The current runtime architecture is documented in [`docs/architecture.md`](docs/architecture.md); older framework plans remain in [`docs/plans/`](docs/plans/) as historical design material. What works today: - a central webhook service, so onboarding a repo is *add the bot + commit `.pr-review.json:enabled = true`* - whole-repo context: the reviewer reads callers and types, not just the hunk - inline comments with language-highlighted suggested fixes, anchored to post-change lines and validated in Python before posting - per-commit dedupe, and prior reviews fed back so a re-push synthesises rather than repeats - `.pr-review.json` for per-repo focus and house rules (also the opt-in flag) - token-usage reporting on every review, measured from opencode `step_finish` events - Langfuse traces, equivalent-cost reporting, evaluation scores, and feedback harvesting - containment against hostile PR content (see [Security](#security)) Not yet: status checks, fail-close, attention tiering enforced in code (it is currently a skill the agent follows), multi-model routing, and a CLI framework. ## How a review runs ``` PR opened on repo with `.pr-review.json:enabled = true` โ”‚ Gitea webhook (HMAC-verified, body-capped, concurrency-bounded) โ–ผ review_pr() 1. opt-in .pr-review.json:enabled=true on base? if not, skip. 2. dedupe already reviewed this exact sha? stop. 3. fetch diff + .pr-review.json from the BASE branch 4. checkout repo archive at head sha โ†’ temp workdir 5. sanitize delete author-controlled agent-instruction files 6. brief .pragent/brief.md, untrusted parts explicitly fenced 7. review opencode agent: read code, run linters, emit findings JSON 8. anchor validate every line against the diff's post-change lines 9. post inline comments + summary, as pragent-bot ``` Steps 1, 3, 8 and 9 are deterministic Python. The model's only job is step 7 โ€” producing correct findings. It never talks to Gitea, and a finding whose line does not validate becomes a summary bullet rather than a misplaced comment. ## Setup Onboarding a repo, once the service is running for that owner: 1. add `pragent-bot` as a **Write** collaborator 2. commit `.pr-review.json: {"enabled": true}` to the repo's default branch 3. open a PR Standing up the service itself โ€” the webhook, the image, the Gitea SSRF allow-list, the per-owner webhook registration โ€” is in [`pilot/README-webhook.md`](pilot/README-webhook.md). A legacy per-repo CI-step path is in [`pilot/README.md`](pilot/README.md). The model endpoint is supplied at runtime via `PRAGENT_MODEL_BASE_URL`; the committed `opencode.json` carries a placeholder. Per-review token spend, latency, equivalent cost, and evaluation scores are shipped to a self-hosted Langfuse: [`pilot/README-langfuse.md`](pilot/README-langfuse.md). Emission is a silent no-op unless `LANGFUSE_HOST` and the key pair are set. ## Extending it The review "factory" is [`.opencode/`](.opencode/README.md) โ€” agent definitions and skills as plain Markdown. Adding a review dimension is dropping a file in, not writing code: | Add | How | |---|---| | A review lens | `.opencode/agents/.md` + one allow-list line in `pragent.md` | | Domain knowledge | `.opencode/skills//SKILL.md`, referenced from the load table | | Per-repo rules | `.pr-review.json` in the repo being reviewed | Shipped skills: `attention-tiering` (the cost governor), `review-methodology`, `findings-schema`, `linter-playbook`, `security-lens`, `malicious-change`, `comment-craft`. ## Security The reviewer runs an autonomous agent with shell access over a checkout of **the PR author's branch**, and its bot account holds a Write credential. Anyone who can open a PR can therefore put arbitrary text in front of the model and arbitrary files on its disk โ€” the setup exploited in the [April 2026 disclosures against Claude Code Security Review, Gemini CLI Action and Copilot Agent][csa]. Four controls, none of which rely on the model behaving: 1. **No credentials in the agent's environment.** The subprocess environment is built from an allow-list, not inherited. There is nothing to exfiltrate. 2. **No author-controlled instruction files on disk.** Nested `AGENTS.md`, `CLAUDE.md`, `.cursorrules`, a repo `opencode.json` โ€” all deleted before the agent starts, so a PR cannot ship its own system prompt. They are still *reviewed*, as data. 3. **Untrusted-data framing.** PR text and diffs are fenced; the agent reports injection attempts as `critical` findings instead of following them. 4. **Reviewer config comes from the base branch**, so a PR cannot rewrite the rules it is judged by. Plus: tar-slip guards on the archive, a non-root container, and bounded concurrency. Full threat model and residual risks: `pilot/README-webhook.md`. [csa]: https://labs.cloudsecurityalliance.org/research/csa-research-note-comment-control-github-prompt-injection-20/ ## What it costs The pilot runs against a self-hosted model and bills nothing per token, but the token *work* is real. `pilot/cost_model.py` prices it against published API rates, calibrated against runs measured through the usage telemetry (`OBSERVED_RUNS` in that file โ€” append to it, don't guess). Tokens are summed from opencode `step_finish` events per review. Each review is governed by hard limits: 20 completed steps, 120,000 total tokens, 20,000 output tokens, and 480 seconds by default. Limits can be overridden deployment-wide with `PRAGENT_MAX_REVIEW_STEPS`, `PRAGENT_MAX_REVIEW_TOKENS`, `PRAGENT_MAX_REVIEW_OUTPUT_TOKENS`, and `PRAGENT_REVIEW_TIMEOUT`, or per repository in the trusted base-branch `.pr-review.json`; repository values win. Two measured reviews of a ~1100-line PR in this repo: 28 and 31 agent steps, ~2.1M input tokens each, **zero cache reads or writes**. The demo repo's PR, same tier: 126K tokens. The measurements above are historical uncapped runs. A capped run preserves completed output, reports the cap reason in the review, and records it in Langfuse. | Model | this repo, ~1100-line PR | demo repo PR | |---|---:|---:| | Claude Opus 5 | ~$10.79 | ~$0.71 | | Claude Sonnet 5 | ~$4.32 | ~$0.28 | | Claude Haiku 4.5 | ~$2.16 | ~$0.14 | Three things that estimate wrong if you skip them: 1. **The loop resends its context every step.** Cost is roughly quadratic in step count, not linear in diff size. This is what `attention-tiering` exists to cap. 2. **Repository size dominates diff size.** The 16x gap above is the same reviewer on the same tier โ€” the difference is how much repo there was to read. 3. **Prompt caching is worth about a third of the bill** and is not currently happening on this stack. Check `cache_read` before budgeting. ```bash python3 pilot/cost_model.py --help # other mixes, volumes, models ``` ## Development ```bash python3 -m pytest tests -q # stdlib-only tests, no network ``` The pilot is stdlib-only Python by design โ€” it runs from a bare `python:slim` image with the scripts mounted, and has no dependency resolution to go wrong at review time. ## License Not yet chosen. Until one is added, no reuse rights are granted. _pilot eval judges test 1788201461_