92283c44e8
Ship token spend, latency and equivalent cost for every review to the self-hosted Langfuse so per-model behaviour is queryable as a trend rather than one PR comment at a time. langfuse_trace.py is stdlib-only and emits via the public ingestion API. Traces split into `ollama` and `claude` environments keyed off the bare model name, not the provider: both paths go through the same headroom proxy, so the provider prefix says nothing about which spend story a review belongs to. The pilot's own path bills $0, so the reported cost is the equivalent price from cost_model.PRICES. ai_review.py calls _emit_langfuse on both token-spending exit paths (the normal post and the salvage path). Import and emission are wrapped in a blanket except: with no LANGFUSE_HOST or key pair the whole thing is a silent no-op, and a telemetry failure must never fail a review. These files were previously deployed only by way of the image build's `COPY . /app`, so a clean checkout would have silently dropped tracing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
188 lines
7.9 KiB
Markdown
188 lines
7.9 KiB
Markdown
# pragent
|
|
|
|
An AI pull-request reviewer for Gitea that posts **inline comments with suggested
|
|
fixes**, not a wall of prose — and reports what each review cost.
|
|
|
|
A webhook wakes for any PR on a repo whose default branch carries a
|
|
`.pr-review.json` with `"enabled": true`. The service checks the repo out at the
|
|
PR's head commit, reads the changed files *and the code around them*, runs the
|
|
repo's own linters, and posts a review anchored to real lines.
|
|
|
|
**See it work:** [`pragent-demo` PR #1](../pragent-demo/pulls/1) — a PR with
|
|
planted defects, and the review it drew: 9 findings, 3 critical, all anchored
|
|
inline.
|
|
|
|
```
|
|
**[CRITICAL]** search_notes builds its SQL by string concatenation: owner and
|
|
term come from request['query'] and are spliced directly, so a term like
|
|
`' OR 1=1 --` reads every row in the table.
|
|
|
|
Fix: Parameterise owner and term with placeholders and a real LIKE pattern.
|
|
|
|
🪙 ~362 tok (11% · attributed output)
|
|
```
|
|
|
|
## Why this exists
|
|
|
|
CodeRabbit, Greptile and Qodo are good products with fixed review dimensions and
|
|
per-seat pricing. `pragent` targets the case where a platform team needs to **add
|
|
its own dimensions** — an internal compliance rule, a service-catalog ownership
|
|
check, a house performance idiom — without forking a vendor's reviewer. The
|
|
analyzers, the data, and the analytics are yours.
|
|
|
|
It is also self-hosted end to end: the model endpoint is a config value, so the
|
|
code never has to leave your network.
|
|
|
|
## Status
|
|
|
|
A **pilot** is live and reviewing real PRs. The full framework (`pragent init`,
|
|
tiering as code, analyzer fan-out, `explain` / `replay`) is designed but not
|
|
built — see [`docs/plans/`](docs/plans/).
|
|
|
|
What works today:
|
|
|
|
- a central webhook service, so onboarding a repo is *add the bot + commit
|
|
`.pr-review.json:enabled = true`*
|
|
- whole-repo context: the reviewer reads callers and types, not just the hunk
|
|
- inline comments with language-highlighted suggested fixes, anchored to
|
|
post-change lines and validated in Python before posting
|
|
- per-commit dedupe, and prior reviews fed back so a re-push synthesises rather
|
|
than repeats
|
|
- `.pr-review.json` for per-repo focus and house rules (also the opt-in flag)
|
|
- token-usage reporting on every review, measured from opencode `step_finish`
|
|
events
|
|
- containment against hostile PR content (see [Security](#security))
|
|
|
|
Not yet: status checks, fail-close, attention tiering enforced in code (it is
|
|
currently a skill the agent follows), multi-model routing.
|
|
|
|
## How a review runs
|
|
|
|
```
|
|
PR opened on repo with `.pr-review.json:enabled = true`
|
|
│ Gitea webhook (HMAC-verified, body-capped, concurrency-bounded)
|
|
▼
|
|
review_pr()
|
|
1. opt-in .pr-review.json:enabled=true on base? if not, skip.
|
|
2. dedupe already reviewed this exact sha? stop.
|
|
3. fetch diff + .pr-review.json from the BASE branch
|
|
4. checkout repo archive at head sha → temp workdir
|
|
5. sanitize delete author-controlled agent-instruction files
|
|
6. brief .pragent/brief.md, untrusted parts explicitly fenced
|
|
7. review opencode agent: read code, run linters, emit findings JSON
|
|
8. anchor validate every line against the diff's post-change lines
|
|
9. post inline comments + summary, as pragent-bot
|
|
```
|
|
|
|
Steps 1, 3, 8 and 9 are deterministic Python. The model's only job is step 7 —
|
|
producing correct findings. It never talks to Gitea, and a finding whose line
|
|
does not validate becomes a summary bullet rather than a misplaced comment.
|
|
|
|
## Setup
|
|
|
|
Onboarding a repo, once the service is running for that owner:
|
|
|
|
1. add `pragent-bot` as a **Write** collaborator
|
|
2. commit `.pr-review.json: {"enabled": true}` to the repo's default branch
|
|
3. open a PR
|
|
|
|
Standing up the service itself — the webhook, the image, the Gitea SSRF
|
|
allow-list, the per-owner webhook registration — is in
|
|
[`pilot/README-webhook.md`](pilot/README-webhook.md). A legacy per-repo CI-step
|
|
path is in [`pilot/README.md`](pilot/README.md).
|
|
|
|
The model endpoint is supplied at runtime via `PRAGENT_MODEL_BASE_URL`; the
|
|
committed `opencode.json` carries a placeholder.
|
|
|
|
Per-review token spend, latency and equivalent cost are shipped to a
|
|
self-hosted Langfuse, split into `ollama` and `claude` environments so the two
|
|
spend stories stay separate: [`pilot/README-langfuse.md`](pilot/README-langfuse.md).
|
|
Emission is a silent no-op unless `LANGFUSE_HOST` and the key pair are set.
|
|
|
|
## Extending it
|
|
|
|
The review "factory" is [`.opencode/`](.opencode/README.md) — agent definitions
|
|
and skills as plain Markdown. Adding a review dimension is dropping a file in,
|
|
not writing code:
|
|
|
|
| Add | How |
|
|
|---|---|
|
|
| A review lens | `.opencode/agents/<name>.md` + one allow-list line in `pragent.md` |
|
|
| Domain knowledge | `.opencode/skills/<name>/SKILL.md`, referenced from the load table |
|
|
| Per-repo rules | `.pr-review.json` in the repo being reviewed |
|
|
|
|
Shipped skills: `attention-tiering` (the cost governor), `review-methodology`,
|
|
`findings-schema`, `linter-playbook`, `security-lens`, `malicious-change`,
|
|
`comment-craft`.
|
|
|
|
## Security
|
|
|
|
The reviewer runs an autonomous agent with shell access over a checkout of **the
|
|
PR author's branch**, and its bot account holds a Write credential. Anyone who
|
|
can open a PR can therefore put arbitrary text in front of the model and
|
|
arbitrary files on its disk — the setup exploited in the [April 2026 disclosures
|
|
against Claude Code Security Review, Gemini CLI Action and Copilot Agent][csa].
|
|
|
|
Four controls, none of which rely on the model behaving:
|
|
|
|
1. **No credentials in the agent's environment.** The subprocess environment is
|
|
built from an allow-list, not inherited. There is nothing to exfiltrate.
|
|
2. **No author-controlled instruction files on disk.** Nested `AGENTS.md`,
|
|
`CLAUDE.md`, `.cursorrules`, a repo `opencode.json` — all deleted before the
|
|
agent starts, so a PR cannot ship its own system prompt. They are still
|
|
*reviewed*, as data.
|
|
3. **Untrusted-data framing.** PR text and diffs are fenced; the agent reports
|
|
injection attempts as `critical` findings instead of following them.
|
|
4. **Reviewer config comes from the base branch**, so a PR cannot rewrite the
|
|
rules it is judged by.
|
|
|
|
Plus: tar-slip guards on the archive, a non-root container, and bounded
|
|
concurrency. Full threat model and residual risks: `pilot/README-webhook.md`.
|
|
|
|
[csa]: https://labs.cloudsecurityalliance.org/research/csa-research-note-comment-control-github-prompt-injection-20/
|
|
|
|
## What it costs
|
|
|
|
The pilot runs against a self-hosted model and bills nothing per token, but the
|
|
token *work* is real. `pilot/cost_model.py` prices it against published API
|
|
rates, calibrated against runs measured through the usage telemetry
|
|
(`OBSERVED_RUNS` in that file — append to it, don't guess). Tokens are summed
|
|
from opencode `step_finish` events per review.
|
|
|
|
Two measured reviews of a ~1100-line PR in this repo: 28 and 31 agent steps,
|
|
~2.1M input tokens each, **zero cache reads or writes**. The demo repo's PR, same
|
|
tier: 126K tokens.
|
|
|
|
| Model | this repo, ~1100-line PR | demo repo PR |
|
|
|---|---:|---:|
|
|
| Claude Opus 5 | ~$10.79 | ~$0.71 |
|
|
| Claude Sonnet 5 | ~$4.32 | ~$0.28 |
|
|
| Claude Haiku 4.5 | ~$2.16 | ~$0.14 |
|
|
|
|
Three things that estimate wrong if you skip them:
|
|
|
|
1. **The loop resends its context every step.** Cost is roughly quadratic in step
|
|
count, not linear in diff size. This is what `attention-tiering` exists to cap.
|
|
2. **Repository size dominates diff size.** The 16x gap above is the same
|
|
reviewer on the same tier — the difference is how much repo there was to read.
|
|
3. **Prompt caching is worth about a third of the bill** and is not currently
|
|
happening on this stack. Check `cache_read` before budgeting.
|
|
|
|
```bash
|
|
python3 pilot/cost_model.py --help # other mixes, volumes, models
|
|
```
|
|
|
|
## Development
|
|
|
|
```bash
|
|
python3 -m pytest tests -q # 137 tests, stdlib only, no network
|
|
```
|
|
|
|
The pilot is stdlib-only Python by design — it runs from a bare `python:slim`
|
|
image with the scripts mounted, and has no dependency resolution to go wrong at
|
|
review time.
|
|
|
|
## License
|
|
|
|
Not yet chosen. Until one is added, no reuse rights are granted.
|