docs: rework design after prior-art review

Red Hat's MIT ai-code-review already implements phases 1-2 (four forge clients,
six providers, CI integration, repo context file). Adds a research writeup,
inserts Phase 0 (evaluate it before building), and folds in seven requirements
the original design missed — chiefly prior-comment synthesis, without which our
own 1.7-runs-per-PR assumption means every push re-posts dismissed findings.

Amends implementation tasks 4, 5, 6, 9, 10 and gates the subagent briefs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011Ye1KNFMkkUtmzTypHXkoK
This commit is contained in:
Marcos
2026-08-04 17:27:15 +00:00
parent 8d690f5717
commit 8a4239a9d9
6 changed files with 501 additions and 1 deletions
+99 -1
View File
@@ -1,7 +1,9 @@
# pragent — Design
**Date:** 2026-08-04
**Status:** Approved (brainstorming complete, ready for implementation planning)
**Status:** Approved, then revised the same day after a prior-art review — see
`docs/research/2026-08-04-prior-art-ai-code-review.md`. **Phase 0 (evaluate the closest
existing tool) now precedes implementation.** Sections 6 and 7 below are the revisions.
## Problem
@@ -182,6 +184,19 @@ which is what makes the confidence gates tunable rather than guessed.
## 5. Rollout phases
**Phase 0 — Evaluate `ai-code-review` (do this first).** Red Hat's MIT-licensed
`ai-code-review` already implements Phases 12 of this design: four forge clients
(GitLab, GitHub, Forgejo, local git), six AI providers, CI integration, a committed
repo-context file. Run it on real repos in the Gitea setup for a week before writing
pragent code. *Done when:* we can name, from use rather than speculation, which of our
differentiators (tiering, analyzer bus, analytics) are worth building, and whether to
build them standalone, on top of that CLI, or as upstream contributions. See
`docs/research/2026-08-04-prior-art-ai-code-review.md` for the full assessment.
Phases 16 below stand as written **if** Phase 0 concludes we build standalone. If it
concludes we extend an existing base, Phases 12 mostly disappear and Phases 36 become
the whole project.
1. **Walking skeleton**`local` forge adapter, one analyzer, rules-only tiering, JSONL
emitter. Runs on a local diff, prints findings. No network beyond the model API.
2. **Gitea end-to-end** — Gitea adapter, Woodpecker step, PR comments, status checks.
@@ -192,6 +207,89 @@ which is what makes the confidence gates tunable rather than guessed.
5. **Second forge** — GitLab adapter, Jenkins runner recipe. Proves the abstraction.
6. **Analytics maturity** — OTel export, feedback loop, per-analyzer eval harness.
## 6. Revisions from the prior-art review
Seven requirements the original design missed. They apply on every path — standalone,
fork, or contribution — so they are part of the design now rather than a backlog.
### 6.1 Prior-comment context and synthesis (highest priority)
Before the main review, fetch **all** existing comments and reviews on the PR, including
resolved ones, and compress them with a **cheap model** (Haiku-class) into a short summary
of what has already been said, fixed, or explicitly rejected. Feed that summary to the
analyzers.
This closes a hole the original design created. The cost model assumes ~1.7 review runs
per PR — every push re-reviews. Without prior-comment context, the second run repeats the
first run's findings and re-argues points a human already dismissed. That is the fastest
way for an AI reviewer to get muted, and it was designed in by accident.
Cost impact is favourable: one Haiku call (~$0.01) to compress a thread, against re-posting
findings that get ignored. Config: `reviewContext.enabled`, `reviewContext.synthesisModel`,
`reviewContext.maxTokens`.
### 6.2 Team context document
`teamContextFile` — a local path **or a URL** — carrying organization-wide review guidance:
security requirements, house conventions, compliance language. It outranks the repo
profile, so one document steers every repo without copying. This is the missing half of
"roll it out across many projects"; org config layering handles thresholds, this handles
judgment.
### 6.3 Deterministic skip conditions
The tier engine gains free, pre-model skips beyond paths and sizes: **draft/WIP MRs**,
`WIP` in the commit message, `wip/` branch prefixes, bot authors, tagged MRs. Each records
its `tier_reason` like any other rule. Reviewing a draft PR at full tier is a pure waste
the original rules would not have caught.
### 6.4 Self-hosted forge configuration
`forgeUrl` per adapter, `sslVerify`, `sslCertPath`. A self-hosted GitLab behind a corporate
CA is the normal case for the company deployment this targets, and the original design had
no way to express it.
### 6.5 Provider matrix as a requirement
The `ModelClient` port must be exercised by more than one provider before we claim it is a
port. Required: Anthropic, plus **Ollama or another local runtime** — "the diff never
leaves our network" is a procurement requirement for regulated repos, not a preference.
Vertex/Bedrock variants follow the same shape.
### 6.6 Per-provider input clamp
Alongside the file and line caps, clamp total diff characters with a provider-aware
default (roughly: 150k Anthropic, 200k Gemini, 100k OpenAI, 50k Ollama). The `oversized`
tier decides *strategy*; the clamp is the backstop that keeps a pathological diff from
blowing the context window regardless of tier.
### 6.7 Smaller additions
- **PR summary** alongside findings — a short description of what the change does
- **`--dry-run`** with mock responses, so a team can wire the pipeline before buying keys
- **Gitea adapter targets Forgejo too** — Forgejo is a Gitea fork with a compatible API;
this is close to free and doubles the addressable forges
- **Library-docs enrichment** (à la Context7) as a profile-enricher plugin, not core
## 7. What remains genuinely ours
After the revisions above, the differentiation is narrower and clearer than the original
design implied. Existing tools — Red Hat's CLI, CodeRabbit, Greptile, Qodo — do adapters,
providers, prompts, and skip logic. None of them do:
1. **Graded attention with recorded provenance.** Four tiers, each decision naming the rule
that fired. Everyone else has binary skip-or-review.
2. **An analyzer plugin bus.** Per-analyzer model, effort, and tool budget, fanned out over
a shared cached prefix — so a team adds a review dimension without forking anything.
3. **Measurement.** Run records, `explain`, `replay`, and finding-outcome feedback, which
together make false-positive rate per analyzer a number you can query and a prompt change
something you can A/B. No tool in this category can answer "did last week's prompt change
help?"
4. **Org-locked policy.** Config a repo cannot downgrade.
If Phase 0 shows an existing base covers everything else well, these four are the project —
and they may be worth contributing upstream rather than shipping standalone.
## Non-goals (for now)
- Central webhook service (CLI-only until a team actually needs zero-setup onboarding)