Files
pragent/README.md
T
Marcos 8c491a7626 harden(pilot): contain hostile PR content, bound the webhook, fix anchoring
The reviewer runs an opencode agent with `bash: "*": allow` over a checkout of
the PR author's branch, and the pod holds a Gitea Write credential. Those two
facts had no wall between them.

Security
- _build_env now allow-lists the subprocess environment instead of inheriting
  it, so PRAGENT_BOT_TOKEN and WEBHOOK_SECRET never reach the agent. This was
  the live hole: a PR body or an AGENTS.md could ask the agent to `curl` the
  token out, and it had both the value and the tool.
- sanitize_workdir deletes author-controlled agent-instruction files from the
  checkout before opencode starts (AGENTS.md at any depth, CLAUDE.md,
  .cursorrules, a repo opencode.json/.opencode, copilot-instructions.md).
  opencode loads nested AGENTS.md as instructions, so a PR could otherwise ship
  its own system prompt. They are still reviewed, as data.
- The brief fences PR title/body and diff in --- UNTRUSTED --- markers under a
  trust-boundary preamble; the pragent agent, the three lens subagents and the
  review-methodology skill now treat injection attempts as a critical finding
  to report rather than an instruction to obey.
- .pr-review.json is read from the PR's base branch, not the head sha. Its
  `instructions` field is spliced into the reviewer's prompt, so head-ref
  reading let any author rewrite the reviewer's rules. Fields are length-capped.
- Untar rejects escaping symlinks, parent traversal, and writes through a
  planted symlink (tar-slip).
- The image runs as uid 10001 instead of root.

Robustness
- Bounded review concurrency (PRAGENT_MAX_CONCURRENT_REVIEWS, default 2). Each
  review forks an opencode process; a thread per delivery was a fork bomb on a
  burst of labels or Gitea retries.
- An in-flight (repo, index, sha) claim closes the check-then-act race in the
  sha-marker dedupe, where two deliveries a second apart both read "not yet
  reviewed" and both posted.
- Request bodies are capped before being read into memory.

Correctness
- parse_diff_anchors counts a whitespace-stripped blank context line. Skipping
  it desynced the new-line counter for the rest of the hunk and silently
  misplaced every later inline comment in that file.
- post_inline_review's body-only fallback folds the anchored findings into the
  body. It previously posted a summary saying "N inline comment(s) below" with
  no comments and no findings — losing them all on the one path that matters.
- fetch_pr_diff's files-endpoint fallback emits real a// b/ prefixes (so
  changed_files and the anchor parser work on it) and reports both HTTP statuses
  in its error instead of the same one twice.
- The CI workflow template pins PRAGENT_ENGINE=ollama; review_pr defaults to
  opencode, which does not exist on a Gitea Actions runner.

Tests: 68 -> 101, covering each of the above.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B11e8TZZxJyzHW7jj7KWUN
2026-08-18 04:44:44 +00:00

105 lines
4.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# pragent
An extensible, forge-agnostic PR review framework. Not a product — a toolkit that teams
extend with their own review dimensions.
**Status:** design approved; framework build deferred. A **pilot** is live on
`glm-5.2:cloud` with two delivery paths:
- **Central webhook service** (preferred, least per-repo setup): a Gitea
user-level webhook posts PR events to an always-on in-cluster service that
gates on the `AI-REVIEW` label. Onboarding a repo = add `pragent-bot`
collaborator + create the label + label a PR. See
[`pilot/README-webhook.md`](pilot/README-webhook.md).
- **CI-step** (legacy): a per-repo Gitea Action fetches the reviewer script at
runtime. See [`pilot/README.md`](pilot/README.md).
The framework design remains at
[`docs/plans/2026-08-04-pragent-design.md`](docs/plans/2026-08-04-pragent-design.md);
the pilot is its bootstrap and will be superseded by `pragent review` when the
framework build resumes.
## What it is
`pragent` runs as a CI step. It reads a pull request, decides how much attention the
change deserves, runs the analyzers that apply, and posts ranked findings back to the
forge.
```
pragent init # one-time repo scan → .pragent/profile.yml (committed, reviewable)
pragent review # the CI step: tier → analyze → aggregate → publish
pragent explain # why did this PR get this tier / these findings?
pragent replay # re-run a past PR against a new prompt or model (the eval loop)
pragent doctor # config, credentials, and adapter health
```
## Why not CodeRabbit / Greptile / Qodo
Those are good products with fixed review dimensions and per-seat pricing. `pragent`
targets the case where a platform team needs to **add its own dimensions** — an internal
compliance rule, a service-catalog ownership check, a house performance idiom — without
forking a vendor's reviewer. Cost lands in the same range (~$25/dev/month at 350 PRs/mo
for 20 devs), but the analyzers, the data, and the analytics are yours.
## Attention tiers
Every PR is classified before any expensive work happens. Deterministic rules decide
first; an ambiguous case gets one cheap model call as tie-breaker.
| Tier | What it means | Cost/PR |
|---|---|---|
| `trivial` | lockfile bumps, generated code, docs typos | ~$0.005 |
| `lite` | small change, no risk paths | ~$0.08 |
| `full` | the default for real changes | ~$0.802.00 |
| `oversized` | too big to review whole; structural summary + deep pass on the hot subset | ~$5 ceiling |
Every tier decision records *why*, so a surprising outcome is explainable rather than
mysterious.
## Extension points
Five, all documented in the design doc. Teams override or add; nobody forks.
1. **Analyzers** — drop a YAML + prompt in `.pragent/analyzers/`, or install from npm
2. **Forge adapters** — Gitea, GitLab, GitHub, local diff
3. **Tier policy** — thresholds and the path risk map, per repo or per org
4. **Profile enrichers** — extend what `pragent init` learns about a repo
5. **Emitter sinks** — JSONL by default, OpenTelemetry, or your own
Org config can lock keys, so a repo cannot quietly disable the security analyzer.
## Design principles
- **Polyglot by construction.** Language knowledge lives in the repo profile, not in the
reviewer. A new language is a profile change, not a core change.
- **One shared prompt prefix.** All analyzers for a PR share a byte-identical cached
prefix. This is what makes fan-out affordable; it is enforced, not hoped for.
- **Everything is traceable.** Tier reasons, token counts, cost, latency, and finding
outcomes are recorded per run. False-positive rate is measurable per analyzer.
- **Fail open.** A budget ceiling or an analyzer crash yields a partial review with a
clear note, never a blocked pipeline with no explanation.
- **The reviewed code is untrusted input.** The reviewer runs an agent over a
branch anyone with PR access can write. So it holds no credentials in its
environment, the checkout is stripped of files an agent runtime would load as
instructions, PR-authored text is fenced as data, and reviewer config is read
from the base branch. See "Threat model" in
[`pilot/README-webhook.md`](pilot/README-webhook.md).
## Stack
TypeScript + Node, built on the [`pi`](https://github.com/badlogic/pi-mono) agent SDK.
Shipped as an npm package and an OCI image, so CI runners need no local Node install.
## Roadmap
1. Walking skeleton — local diff, one analyzer, rules-only tiering
2. Gitea end to end — adapter, Woodpecker step, PR comments, status checks
3. Profile + full tier — `pragent init`, shared-prefix caching, analyzer fan-out
4. Extensibility hardening — plugin loading, config layering, `explain` / `replay`
5. Second forge — GitLab adapter, Jenkins recipe
6. Analytics maturity — OTel export, feedback loop, eval harness
## License
TBD.