Files
pragent/.opencode
Marcos 30d2a3d7da feat(factory): five review skills + a per-review cost model
Skills — the primary now loads conditionally (each one is input tokens), per a
load table in pragent.md:

- attention-tiering: classify every PR trivial/lite/full/oversized BEFORE
  reading anything, and cap file reads, linter runs and subagent fan-out per
  tier. This is the cost governor; the other skills defer to its budget.
- linter-playbook: per-ecosystem detect-and-run commands scoped to changed
  files, the never-install rule, and how to turn a diagnostic into a finding
  instead of pasting tool output.
- security-lens: the inline security checklist for when @security isn't worth
  delegating, built around a source -> sink test each finding must pass.
- malicious-change: hostile-PR detection — injection aimed at the reviewer,
  install/CI-time hooks, obfuscated payloads, dependency confusion, logic
  backdoors. Complements the runtime containment added in the previous commit:
  that stops the agent being hijacked, this makes it report the attempt.
- comment-craft: how to write problem/fix/suggestion so a maintainer can act in
  one read, and what to cut.

pilot/cost_model.py — prices a review against published Claude and OpenAI rates
(fetched 2026-08-18). Prompt sizes are measured from the factory files rather
than guessed; per-tier workloads come from the tiering budgets. The model is
explicit about the thing that actually dominates an agent loop: the whole
conversation is resent every step, so caching moves ~2.3x of the bill.

Blended over a 5/35/55/5 mix with caching on: ~$0.61/PR on Opus 5 or GPT-5.6
Sol, ~$0.24 on Sonnet 5 or Terra, ~$0.12 on Haiku 4.5, ~$0.02 on Luna. At 350
PRs/month that's ~$212 / ~$85 / ~$43 / ~$8.50.

Tests: 101 -> 122.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B11e8TZZxJyzHW7jj7KWUN
2026-08-18 04:53:49 +00:00
..

pragent .opencode/ — the review factory

pragent's review runs on opencode (the AI-coding-agent CLI). This directory is a portable factory: the opencode.json + .opencode/ are dropped into a checked-out copy of the target repo at the PR head sha, then opencode run is launched there. The pragent primary agent reviews the diff with real tools (subagents, LSP/linters via bash, webfetch references) and emits a structured findings JSON. A thin Python shell posts that JSON back to Gitea as inline comments + language-highlighted suggested-fix blocks + a summary (dedupe + anchor validation stay deterministic in Python).

Layout

opencode.json              provider (headroom → glm-5.2:cloud), model, lsp, permission, default_agent
.opencode/
  agents/
    pragent.md             PRIMARY reviewer — reads .pragent/brief.md, runs linters, emits findings JSON
    security.md            subagent — injection/auth/secrets/supply-chain lens (dormant)
    tests.md               subagent — missing/weak test coverage lens (dormant)
    perf.md                subagent — N+1 / O(n²) / hot-path lens (dormant)
  skills/
    review-methodology/SKILL.md   severity rubric, what to report, anchoring, trust boundary
    findings-schema/SKILL.md      the exact output JSON shape
    attention-tiering/SKILL.md    trivial/lite/full/oversized + the budget each tier gets
    linter-playbook/SKILL.md      per-ecosystem check commands + turning diagnostics into findings
    security-lens/SKILL.md        inline security checklist + the source→sink test
    malicious-change/SKILL.md     hostile-PR detection: injection at the reviewer, install hooks, backdoors
    comment-craft/SKILL.md        how to write problem/fix/suggestion so a maintainer can act
  commands/
    review.md              /review slash command (local interactive use)
  README.md                this file

How a review runs

flowchart TD
  WH["webhook_server.py<br/>HMAC + AI-REVIEW gate + dedupe"] --> RP["ai_review.review_pr"]
  RP --> ARCH["fetch repo archive @ head sha<br/>→ /tmp/pragent-work/<repo>-<sha>"]
  ARCH --> BRIEF["write .pragent/brief.md<br/>(title, body, diff, config, prior, sha)"]
  BRIEF --> DROP["drop opencode.json + .opencode/ into workdir"]
  DROP --> OC["opencode run --pure --agent pragent --dir <workdir><br/>--model headroom/glm-5.2:cloud"]
  OC --> PR["pragent primary<br/>load skills · run linters · read code · delegate lenses"]
  PR --> JSON["final message: summary + ```json findings```"]
  JSON --> PARSE["ai_review.parse_review_output<br/>{summary, findings}"]
  PARSE --> ANCHOR["parse_diff_anchors → split_findings"]
  ANCHOR --> POST["post_inline_review<br/>summary + inline lang-tagged fix block + ref links + sha marker"]

Lean by default

attention-tiering is the cost governor: it classifies every PR as trivial / lite / full / oversized before any file is read, and each tier caps file reads, linter runs, and subagent fan-out. The pragent primary does the whole review in one pass for small/medium diffs (no subagent calls) and delegates to @security / @tests / @perf ONLY at full/oversized when the lens has real surface. Skills are loaded conditionally for the same reason — each one is input tokens. Subagent recursion is capped by the primary's steps budget.

pilot/cost_model.py turns those tier assumptions into a per-PR and per-month cost figure for any provider — run it after changing the factory to see what the change costs.

--pure is passed at runtime so the reviewer doesn't load the host user's heavy global opencode plugins (supermemory/dcp/morph/pty) which hang cold-start. In the deploy pod there's no global config, so --pure is a no-op there — but it keeps host-local runs deterministic.

Extending the factory

Add a review lens (subagent)

  1. Create .opencode/agents/<name>.md with mode: subagent, hidden: true, a description, and a read-only permission (deny edit/write, allow bash/webfetch, task: deny so it can't recurse). The body is its system prompt; end it by requiring the same findings-JSON shape.
  2. Allow it in the primary's permission.task list in pragent.md:
    task:
      "*": "deny"
      "security": "allow"
      "tests": "allow"
      "<name>": "allow"      # add this
    
  3. Mention in pragent.md's "Delegate on heavy diffs" step when to invoke it.

That's it — the primary can now @<name> it via the Task tool. It stays dormant (the primary decides when), so adding it costs nothing for small PRs.

Add a skill

  1. mkdir .opencode/skills/<name> && touch .opencode/skills/<name>/SKILL.md
  2. Frontmatter: name: <name> (kebab-case, matches dir), description: (specific enough for the agent to pick it). Body = the knowledge.
  3. Refer to it from pragent.md ("Call the skill tool for <name>").

Per-language expertise is free: the host user already has 29 global skills (golang-, react-, k8s, terraform, testing, typescript, …). opencode auto-discovers them via the skill tool — the pragent primary loads a matching one when the repo's language fits. To ship a pragent-specific one, just drop it here.

Change the output shape

Edit .opencode/skills/findings-schema/SKILL.md (the schema doc) AND the Python parser in pilot/ai_review.py (parse_findings) + the renderers (inline_comment_body, summary_bullets, format_review_body). Keep them in sync — the parser is tolerant but the agent and parser must agree on field names.

Switch model / provider

Edit opencode.json provider + model. The provider points at the on-network headroom proxy (http://100.74.17.70:8789/v1, Anthropic /v1/messages format, apiKey: ollama) → glm-5.2:cloud. To use a different model, add a provider and reference it as <provider>/<model-id>.

Local one-shot review (no webhook)

cd ~/Projects/pragent
opencode run --pure --agent pragent --dir . \
  --model headroom/glm-5.2:cloud \
  "Read .pragent/brief.md if present, else review \`git diff HEAD\`, and output findings."

Or in the TUI: /review (uses .opencode/commands/review.md).

Engine flag

PRAGENT_ENGINE=opencode (default once wired) uses this factory. =ollama falls back to the legacy direct model call in pilot/ai_review.py. The two share all Gitea I/O, dedupe, and posting logic.