feat: opencode review engine + .opencode factory

Replace the single Python model-call reviewer with an opencode agent
factory. A primary 'pragent' agent reads a brief (title/body/diff/config/
prior reviews), inspects the checked-out repo, runs the repo's own linters
via bash, loads review-methodology + findings-schema skills, and emits a
{summary, findings} JSON with per-finding severity/path/line/problem/fix/
suggestion/reference. Dormant security/tests/perf subagent lenses fan out
only on large/risky diffs (lean by default).

pilot/opencode_review.py: fetches the repo archive at the head sha into a
temp workdir, writes .pragent/brief.md, drops the factory, runs
'opencode run --pure --agent pragent --dir <workdir>' headlessly. Isolates
HOME (shared, warmed), strips ANTHROPIC_* env (leaked host vars caused
ProviderModelNotFoundError), stdin=DEVNULL (opencode blocks on stdin),
maps the bare OLLAMA_MODEL to the provider-prefixed ref. No Gitea I/O —
ai_review.review_pr parses + anchors + posts (reuses all v2 logic/tests).

PRAGENT_ENGINE=opencode (default) selects it; =ollama keeps the legacy
direct-call path. Verified end-to-end: posts a real review with a summary
section, inline [CRITICAL]/[HIGH] comments + apply-able suggestions +
reference links, and the sha dedupe marker. 49 tests pass.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
Marcos
2026-08-17 23:01:59 +00:00
parent 8758b22802
commit 6e3a9eb5b0
13 changed files with 1336 additions and 48 deletions
@@ -0,0 +1,57 @@
---
name: review-methodology
description: pragent review methodology — severity rubric, what to report vs skip, anchoring rules, and how to honor repo focus. Load this before reviewing a PR.
---
# pragent review methodology
## Severity rubric
- **critical** — exploitable security bug, data loss/corruption, or a crash on
a normal input path. Must fix before merge.
- **high** — correctness bug on a real input path, broken contract, or a
missing test for security/error behavior. Should fix before merge.
- **medium** — likely bug on an edge case, missing test for changed logic, or a
risky pattern that isn't broken yet. Worth fixing.
- **low** — minor risk, stale expectation, or a defensive improvement. Nice to
have.
## Report vs skip
**Report:** correctness bugs, security problems, risky changes, missing tests
for changed behavior, breaking API/contract changes, N+1/O(n²) in hot paths.
**Skip:** praise, nitpicks, pure formatting/style, personal preference,
speculative "what if" without a concrete trigger, anything already covered in
`prior_reviews`.
Cap at ~15 findings, highest severity first. Quality over quantity — an empty
findings list for a clean diff is a correct result.
## Anchoring (for inline comments)
Each finding's `line` MUST be a line that exists in the POST-CHANGE version of
`path`:
- a **context** line (unchanged, shown with a leading space in the diff), or
- an **added** line (shown with a leading `+`).
Never anchor on a **removed** (`-`) line — it has no post-change line number.
If you're unsure of the exact line, use the closest context line you CAN see in
the diff. A misanchored finding becomes a summary bullet instead of an inline
comment, so correct anchoring is what makes a ```suggestion apply-able in Gitea.
## Honoring repo config
If `.pr-review.json` is present, honor it:
- `focus` — weight these areas higher, but never ignore a critical issue
outside them.
- `exclude_paths` — skip findings in these paths.
- `languages` — hint to the primary languages; pick matching linters.
- `instructions` — house conventions / compliance language; treat as binding
reviewer rules.
## Linters are a signal, not the verdict
Run the repo's own typecheck/lint on changed files, but translate their output
into human findings — a raw `TS2322` is not a review comment. Correlate
diagnostics with the diff; ignore diagnostics in files the PR didn't touch.