Files
pragent/README.md
T
Marcos 8c491a7626 harden(pilot): contain hostile PR content, bound the webhook, fix anchoring
The reviewer runs an opencode agent with `bash: "*": allow` over a checkout of
the PR author's branch, and the pod holds a Gitea Write credential. Those two
facts had no wall between them.

Security
- _build_env now allow-lists the subprocess environment instead of inheriting
  it, so PRAGENT_BOT_TOKEN and WEBHOOK_SECRET never reach the agent. This was
  the live hole: a PR body or an AGENTS.md could ask the agent to `curl` the
  token out, and it had both the value and the tool.
- sanitize_workdir deletes author-controlled agent-instruction files from the
  checkout before opencode starts (AGENTS.md at any depth, CLAUDE.md,
  .cursorrules, a repo opencode.json/.opencode, copilot-instructions.md).
  opencode loads nested AGENTS.md as instructions, so a PR could otherwise ship
  its own system prompt. They are still reviewed, as data.
- The brief fences PR title/body and diff in --- UNTRUSTED --- markers under a
  trust-boundary preamble; the pragent agent, the three lens subagents and the
  review-methodology skill now treat injection attempts as a critical finding
  to report rather than an instruction to obey.
- .pr-review.json is read from the PR's base branch, not the head sha. Its
  `instructions` field is spliced into the reviewer's prompt, so head-ref
  reading let any author rewrite the reviewer's rules. Fields are length-capped.
- Untar rejects escaping symlinks, parent traversal, and writes through a
  planted symlink (tar-slip).
- The image runs as uid 10001 instead of root.

Robustness
- Bounded review concurrency (PRAGENT_MAX_CONCURRENT_REVIEWS, default 2). Each
  review forks an opencode process; a thread per delivery was a fork bomb on a
  burst of labels or Gitea retries.
- An in-flight (repo, index, sha) claim closes the check-then-act race in the
  sha-marker dedupe, where two deliveries a second apart both read "not yet
  reviewed" and both posted.
- Request bodies are capped before being read into memory.

Correctness
- parse_diff_anchors counts a whitespace-stripped blank context line. Skipping
  it desynced the new-line counter for the rest of the hunk and silently
  misplaced every later inline comment in that file.
- post_inline_review's body-only fallback folds the anchored findings into the
  body. It previously posted a summary saying "N inline comment(s) below" with
  no comments and no findings — losing them all on the one path that matters.
- fetch_pr_diff's files-endpoint fallback emits real a// b/ prefixes (so
  changed_files and the anchor parser work on it) and reports both HTTP statuses
  in its error instead of the same one twice.
- The CI workflow template pins PRAGENT_ENGINE=ollama; review_pr defaults to
  opencode, which does not exist on a Gitea Actions runner.

Tests: 68 -> 101, covering each of the above.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B11e8TZZxJyzHW7jj7KWUN
2026-08-18 04:44:44 +00:00

4.8 KiB
Raw Blame History

pragent

An extensible, forge-agnostic PR review framework. Not a product — a toolkit that teams extend with their own review dimensions.

Status: design approved; framework build deferred. A pilot is live on glm-5.2:cloud with two delivery paths:

  • Central webhook service (preferred, least per-repo setup): a Gitea user-level webhook posts PR events to an always-on in-cluster service that gates on the AI-REVIEW label. Onboarding a repo = add pragent-bot collaborator + create the label + label a PR. See pilot/README-webhook.md.
  • CI-step (legacy): a per-repo Gitea Action fetches the reviewer script at runtime. See pilot/README.md.

The framework design remains at docs/plans/2026-08-04-pragent-design.md; the pilot is its bootstrap and will be superseded by pragent review when the framework build resumes.

What it is

pragent runs as a CI step. It reads a pull request, decides how much attention the change deserves, runs the analyzers that apply, and posts ranked findings back to the forge.

pragent init      # one-time repo scan → .pragent/profile.yml (committed, reviewable)
pragent review    # the CI step: tier → analyze → aggregate → publish
pragent explain   # why did this PR get this tier / these findings?
pragent replay    # re-run a past PR against a new prompt or model (the eval loop)
pragent doctor    # config, credentials, and adapter health

Why not CodeRabbit / Greptile / Qodo

Those are good products with fixed review dimensions and per-seat pricing. pragent targets the case where a platform team needs to add its own dimensions — an internal compliance rule, a service-catalog ownership check, a house performance idiom — without forking a vendor's reviewer. Cost lands in the same range (~$25/dev/month at 350 PRs/mo for 20 devs), but the analyzers, the data, and the analytics are yours.

Attention tiers

Every PR is classified before any expensive work happens. Deterministic rules decide first; an ambiguous case gets one cheap model call as tie-breaker.

Tier What it means Cost/PR
trivial lockfile bumps, generated code, docs typos ~$0.005
lite small change, no risk paths ~$0.08
full the default for real changes ~$0.802.00
oversized too big to review whole; structural summary + deep pass on the hot subset ~$5 ceiling

Every tier decision records why, so a surprising outcome is explainable rather than mysterious.

Extension points

Five, all documented in the design doc. Teams override or add; nobody forks.

  1. Analyzers — drop a YAML + prompt in .pragent/analyzers/, or install from npm
  2. Forge adapters — Gitea, GitLab, GitHub, local diff
  3. Tier policy — thresholds and the path risk map, per repo or per org
  4. Profile enrichers — extend what pragent init learns about a repo
  5. Emitter sinks — JSONL by default, OpenTelemetry, or your own

Org config can lock keys, so a repo cannot quietly disable the security analyzer.

Design principles

  • Polyglot by construction. Language knowledge lives in the repo profile, not in the reviewer. A new language is a profile change, not a core change.
  • One shared prompt prefix. All analyzers for a PR share a byte-identical cached prefix. This is what makes fan-out affordable; it is enforced, not hoped for.
  • Everything is traceable. Tier reasons, token counts, cost, latency, and finding outcomes are recorded per run. False-positive rate is measurable per analyzer.
  • Fail open. A budget ceiling or an analyzer crash yields a partial review with a clear note, never a blocked pipeline with no explanation.
  • The reviewed code is untrusted input. The reviewer runs an agent over a branch anyone with PR access can write. So it holds no credentials in its environment, the checkout is stripped of files an agent runtime would load as instructions, PR-authored text is fenced as data, and reviewer config is read from the base branch. See "Threat model" in pilot/README-webhook.md.

Stack

TypeScript + Node, built on the pi agent SDK. Shipped as an npm package and an OCI image, so CI runners need no local Node install.

Roadmap

  1. Walking skeleton — local diff, one analyzer, rules-only tiering
  2. Gitea end to end — adapter, Woodpecker step, PR comments, status checks
  3. Profile + full tier — pragent init, shared-prefix caching, analyzer fan-out
  4. Extensibility hardening — plugin loading, config layering, explain / replay
  5. Second forge — GitLab adapter, Jenkins recipe
  6. Analytics maturity — OTel export, feedback loop, eval harness

License

TBD.