30d2a3d7da
Skills — the primary now loads conditionally (each one is input tokens), per a load table in pragent.md: - attention-tiering: classify every PR trivial/lite/full/oversized BEFORE reading anything, and cap file reads, linter runs and subagent fan-out per tier. This is the cost governor; the other skills defer to its budget. - linter-playbook: per-ecosystem detect-and-run commands scoped to changed files, the never-install rule, and how to turn a diagnostic into a finding instead of pasting tool output. - security-lens: the inline security checklist for when @security isn't worth delegating, built around a source -> sink test each finding must pass. - malicious-change: hostile-PR detection — injection aimed at the reviewer, install/CI-time hooks, obfuscated payloads, dependency confusion, logic backdoors. Complements the runtime containment added in the previous commit: that stops the agent being hijacked, this makes it report the attempt. - comment-craft: how to write problem/fix/suggestion so a maintainer can act in one read, and what to cut. pilot/cost_model.py — prices a review against published Claude and OpenAI rates (fetched 2026-08-18). Prompt sizes are measured from the factory files rather than guessed; per-tier workloads come from the tiering budgets. The model is explicit about the thing that actually dominates an agent loop: the whole conversation is resent every step, so caching moves ~2.3x of the bill. Blended over a 5/35/55/5 mix with caching on: ~$0.61/PR on Opus 5 or GPT-5.6 Sol, ~$0.24 on Sonnet 5 or Terra, ~$0.12 on Haiku 4.5, ~$0.02 on Luna. At 350 PRs/month that's ~$212 / ~$85 / ~$43 / ~$8.50. Tests: 101 -> 122. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B11e8TZZxJyzHW7jj7KWUN
133 lines
6.3 KiB
Markdown
133 lines
6.3 KiB
Markdown
# pragent `.opencode/` — the review factory
|
|
|
|
pragent's review runs on **opencode** (the AI-coding-agent CLI). This directory is
|
|
a portable **factory**: the `opencode.json` + `.opencode/` are dropped into a
|
|
checked-out copy of the target repo at the PR head sha, then `opencode run` is
|
|
launched there. The `pragent` primary agent reviews the diff with real tools
|
|
(subagents, LSP/linters via bash, webfetch references) and emits a structured
|
|
findings JSON. A thin Python shell posts that JSON back to Gitea as inline
|
|
comments + language-highlighted suggested-fix blocks + a summary (dedupe + anchor validation stay
|
|
deterministic in Python).
|
|
|
|
## Layout
|
|
|
|
```
|
|
opencode.json provider (headroom → glm-5.2:cloud), model, lsp, permission, default_agent
|
|
.opencode/
|
|
agents/
|
|
pragent.md PRIMARY reviewer — reads .pragent/brief.md, runs linters, emits findings JSON
|
|
security.md subagent — injection/auth/secrets/supply-chain lens (dormant)
|
|
tests.md subagent — missing/weak test coverage lens (dormant)
|
|
perf.md subagent — N+1 / O(n²) / hot-path lens (dormant)
|
|
skills/
|
|
review-methodology/SKILL.md severity rubric, what to report, anchoring, trust boundary
|
|
findings-schema/SKILL.md the exact output JSON shape
|
|
attention-tiering/SKILL.md trivial/lite/full/oversized + the budget each tier gets
|
|
linter-playbook/SKILL.md per-ecosystem check commands + turning diagnostics into findings
|
|
security-lens/SKILL.md inline security checklist + the source→sink test
|
|
malicious-change/SKILL.md hostile-PR detection: injection at the reviewer, install hooks, backdoors
|
|
comment-craft/SKILL.md how to write problem/fix/suggestion so a maintainer can act
|
|
commands/
|
|
review.md /review slash command (local interactive use)
|
|
README.md this file
|
|
```
|
|
|
|
## How a review runs
|
|
|
|
```mermaid
|
|
flowchart TD
|
|
WH["webhook_server.py<br/>HMAC + AI-REVIEW gate + dedupe"] --> RP["ai_review.review_pr"]
|
|
RP --> ARCH["fetch repo archive @ head sha<br/>→ /tmp/pragent-work/<repo>-<sha>"]
|
|
ARCH --> BRIEF["write .pragent/brief.md<br/>(title, body, diff, config, prior, sha)"]
|
|
BRIEF --> DROP["drop opencode.json + .opencode/ into workdir"]
|
|
DROP --> OC["opencode run --pure --agent pragent --dir <workdir><br/>--model headroom/glm-5.2:cloud"]
|
|
OC --> PR["pragent primary<br/>load skills · run linters · read code · delegate lenses"]
|
|
PR --> JSON["final message: summary + ```json findings```"]
|
|
JSON --> PARSE["ai_review.parse_review_output<br/>{summary, findings}"]
|
|
PARSE --> ANCHOR["parse_diff_anchors → split_findings"]
|
|
ANCHOR --> POST["post_inline_review<br/>summary + inline lang-tagged fix block + ref links + sha marker"]
|
|
```
|
|
|
|
## Lean by default
|
|
|
|
`attention-tiering` is the cost governor: it classifies every PR as `trivial` /
|
|
`lite` / `full` / `oversized` before any file is read, and each tier caps file
|
|
reads, linter runs, and subagent fan-out. The `pragent` primary does the whole
|
|
review in one pass for small/medium diffs (no subagent calls) and delegates to
|
|
`@security` / `@tests` / `@perf` ONLY at `full`/`oversized` when the lens has
|
|
real surface. Skills are loaded conditionally for the same reason — each one is
|
|
input tokens. Subagent recursion is capped by the primary's `steps` budget.
|
|
|
|
`pilot/cost_model.py` turns those tier assumptions into a per-PR and per-month
|
|
cost figure for any provider — run it after changing the factory to see what the
|
|
change costs.
|
|
|
|
`--pure` is passed at runtime so the reviewer doesn't load the host user's heavy
|
|
global opencode plugins (supermemory/dcp/morph/pty) which hang cold-start. In the
|
|
deploy pod there's no global config, so `--pure` is a no-op there — but it keeps
|
|
host-local runs deterministic.
|
|
|
|
## Extending the factory
|
|
|
|
### Add a review lens (subagent)
|
|
|
|
1. Create `.opencode/agents/<name>.md` with `mode: subagent`, `hidden: true`, a
|
|
`description`, and a read-only `permission` (deny edit/write, allow bash/webfetch,
|
|
`task: deny` so it can't recurse). The body is its system prompt; end it by
|
|
requiring the same findings-JSON shape.
|
|
2. Allow it in the primary's `permission.task` list in `pragent.md`:
|
|
```yaml
|
|
task:
|
|
"*": "deny"
|
|
"security": "allow"
|
|
"tests": "allow"
|
|
"<name>": "allow" # add this
|
|
```
|
|
3. Mention in `pragent.md`'s "Delegate on heavy diffs" step when to invoke it.
|
|
|
|
That's it — the primary can now `@<name>` it via the Task tool. It stays dormant
|
|
(the primary decides when), so adding it costs nothing for small PRs.
|
|
|
|
### Add a skill
|
|
|
|
1. `mkdir .opencode/skills/<name> && touch .opencode/skills/<name>/SKILL.md`
|
|
2. Frontmatter: `name: <name>` (kebab-case, matches dir), `description:` (specific
|
|
enough for the agent to pick it). Body = the knowledge.
|
|
3. Refer to it from `pragent.md` ("Call the `skill` tool for `<name>`").
|
|
|
|
Per-language expertise is free: the host user already has 29 global skills
|
|
(golang-*, react-*, k8s, terraform, testing, typescript, …). opencode
|
|
auto-discovers them via the `skill` tool — the pragent primary loads a matching
|
|
one when the repo's language fits. To ship a pragent-specific one, just drop it
|
|
here.
|
|
|
|
### Change the output shape
|
|
|
|
Edit `.opencode/skills/findings-schema/SKILL.md` (the schema doc) AND the Python
|
|
parser in `pilot/ai_review.py` (`parse_findings`) + the renderers
|
|
(`inline_comment_body`, `summary_bullets`, `format_review_body`). Keep them in
|
|
sync — the parser is tolerant but the agent and parser must agree on field names.
|
|
|
|
### Switch model / provider
|
|
|
|
Edit `opencode.json` `provider` + `model`. The provider points at the on-network
|
|
headroom proxy (`http://100.74.17.70:8789/v1`, Anthropic `/v1/messages` format,
|
|
`apiKey: ollama`) → `glm-5.2:cloud`. To use a different model, add a provider and
|
|
reference it as `<provider>/<model-id>`.
|
|
|
|
## Local one-shot review (no webhook)
|
|
|
|
```bash
|
|
cd ~/Projects/pragent
|
|
opencode run --pure --agent pragent --dir . \
|
|
--model headroom/glm-5.2:cloud \
|
|
"Read .pragent/brief.md if present, else review \`git diff HEAD\`, and output findings."
|
|
```
|
|
|
|
Or in the TUI: `/review` (uses `.opencode/commands/review.md`).
|
|
|
|
## Engine flag
|
|
|
|
`PRAGENT_ENGINE=opencode` (default once wired) uses this factory. `=ollama`
|
|
falls back to the legacy direct model call in `pilot/ai_review.py`. The two
|
|
share all Gitea I/O, dedupe, and posting logic. |