docs: rework design after prior-art review

Red Hat's MIT ai-code-review already implements phases 1-2 (four forge clients,
six providers, CI integration, repo context file). Adds a research writeup,
inserts Phase 0 (evaluate it before building), and folds in seven requirements
the original design missed — chiefly prior-comment synthesis, without which our
own 1.7-runs-per-PR assumption means every push re-posts dismissed findings.

Amends implementation tasks 4, 5, 6, 9, 10 and gates the subagent briefs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011Ye1KNFMkkUtmzTypHXkoK
This commit is contained in:
Marcos
2026-08-04 17:27:15 +00:00
parent 8d690f5717
commit 8a4239a9d9
6 changed files with 501 additions and 1 deletions
+2
View File
@@ -0,0 +1,2 @@
/cache
/project.local.yml
+166
View File
@@ -0,0 +1,166 @@
# the name by which the project can be referenced within Serena/when chatting with the LLM.
project_name: "pragent"
# list of language servers to start when using the LSP backend; choose from:
# ada al angular ansible bash
# bsl clojure cpp cpp_ccls crystal
# csharp csharp_omnisharp cue dart elixir
# elm erlang fortran fsharp gdscript
# go groovy haskell haxe hlsl
# html java json julia kotlin
# latex lean4 lua luau markdown
# matlab msl nix ocaml pascal
# perl php php_phpactor php_phpantom powershell
# python python_basedpyright python_jedi python_pyrefly python_ty
# qml r rego ruby ruby_solargraph
# rust scala scss solidity svelte
# swift systemverilog terraform toml typescript
# typescript_vts vue yaml zig
# (This list may be outdated; generated with scripts/print_language_list.py;
# For the current list, see values of the LanguageServerId enum here:
# https://github.com/oraios/serena/blob/main/src/solidlsp/ls_config.py)
# For some languages, there are several alternative language servers, e.g. csharp_omnisharp, ruby_solargraph.)
# Note:
# - For C, use cpp
# - For JavaScript, use typescript
# - For Angular projects, use angular (subsumes typescript+html; requires `npm install` in the project root)
# - For Svelte projects, use svelte (subsumes typescript/javascript for .svelte projects; requires npm)
# - For SCSS / Sass / plain CSS, use scss (some-sass-language-server handles all three)
# - For Free Pascal/Lazarus, use pascal
# Special requirements:
# Some language servers require additional setup/installations.
# See here for details: https://oraios.github.io/serena/01-about/020_programming-languages.html#language-servers
# When using multiple language servers, the first language server that supports a given file will be used for that file.
# The first language server is the default language and the respective language server will be used as a fallback.
# Note that when using the JetBrains backend, language servers are not used and this list is correspondingly ignored.
language_servers: []
# the encoding used by text files in the project
# For a list of possible encodings, see https://docs.python.org/3.11/library/codecs.html#standard-encodings
encoding: "utf-8"
# optional shell command to run before the language backend (LSP or JetBrains) is initialised.
# the command runs in the project root directory and is only executed if the project is trusted
# (see trusted_project_path_patterns in the global configuration).
# serena waits for the command to exit: a non-zero exit code is logged as an error but does not
# abort activation. a per-project timeout (activation_command_timeout, default 180s) is the safety
# backstop for non-terminating commands; on expiry the process is killed and activation continues.
# example: activation_command: "npx nx run-many -t build"
activation_command:
# maximum time in seconds to wait for activation_command to complete before killing it (default 180s).
# must be a positive number.
activation_command_timeout: 180.0
# line ending convention to use when writing source files.
# Possible values: unset (use global setting), "lf", "crlf", or "native" (platform default)
# This does not affect Serena's own files (e.g. memories and configuration files), which always use native line endings.
line_ending:
# The language backend to use for this project.
# If not set, the global setting from serena_config.yml is used.
# Valid values: LSP, JetBrains
# Note: the backend is fixed at startup. If a project with a different backend
# is activated post-init, an error will be returned.
language_backend:
# whether to use project's .gitignore files to ignore files
ignore_all_files_in_gitignore: true
# advanced configuration option allowing to configure language server-specific options.
# Maps the language key to the options.
# The settings are considered only if the project is trusted (see global configuration to define trusted projects).
# See https://oraios.github.io/serena/02-usage/050_configuration.html#language-server-specific-settings
ls_specific_settings: {}
# list of workspace folder paths (LSP backend only).
# These folders will be used to build up Serena's symbol index.
# Paths must be within the project root and should thus be relative to the project root.
# Furthermore, the paths should not be filtered by ignore settings.
# Default setting: The entire project root folder (".") is considered.
# In (large) monorepos, this can be used to index only subfolders of the project root, e.g.
# ls_workspace_folders:
# - "./subproject1"
# - "./subproject2"
ls_workspace_folders:
- "."
# list of additional workspace folder paths for cross-package reference support.
# Paths can be absolute or relative to the project root.
# Each folder is registered as an LSP workspace folder, enabling language servers to discover
# symbols and references across package boundaries, but these folders are not indexed by Serena,
# i.e. the respective symbols will not be found using Serena's symbol search tools.
# Example:
# additional_workspace_folders:
# - ../sibling-package
# - ../shared-lib
ls_additional_workspace_folders: []
# list of additional paths to ignore in this project.
# Same syntax as gitignore, so you can use * and **.
# Important: quote patterns that start with `*`, otherwise YAML treats them as aliases.
# Example:
# ignored_paths:
# - "examples/**"
# - ".worktrees/**"
# - "**/bin/**"
# - "**/obj/**"
# Note: global ignored_paths from serena_config.yml are also applied additively.
ignored_paths: []
# whether the project is in read-only mode
# If set to true, all editing tools will be disabled and attempts to use them will result in an error
# Added on 2025-04-18
read_only: false
# list of tool names to exclude.
# This extends the existing exclusions (e.g. from the global configuration)
# Find the list of tools here: https://oraios.github.io/serena/01-about/035_tools.html
excluded_tools: []
# list of tools to include that would otherwise be disabled (particularly optional tools that are disabled by default).
# This extends the existing inclusions (e.g. from the global configuration).
# Find the list of tools here: https://oraios.github.io/serena/01-about/035_tools.html
included_optional_tools: []
# fixed set of tools to use as the base tool set (if non-empty), replacing Serena's default set of tools.
# This cannot be combined with non-empty excluded_tools or included_optional_tools.
# Find the list of tools here: https://oraios.github.io/serena/01-about/035_tools.html
fixed_tools: []
# list of mode names that are to be activated by default, overriding the setting in the global configuration.
# The full set of modes to be activated is base_modes (from global config) + default_modes + added_modes.
# If the setting is undefined/empty, the default_modes from the global configuration (serena_config.yml) apply.
# Otherwise, this overrides the setting from the global configuration (serena_config.yml).
# Therefore, you can set this to [] if you do not want the default modes defined in the global config to apply
# for this project.
# This setting can, in turn, be overridden by CLI parameters (--mode).
# See https://oraios.github.io/serena/02-usage/050_configuration.html#modes
default_modes:
# list of mode names to be activated additionally for this project, e.g. ["query-projects"]
# The full set of modes to be activated is base_modes (from global config) + default_modes + added_modes.
# See https://oraios.github.io/serena/02-usage/050_configuration.html#modes
added_modes:
# initial prompt for the project. It will always be given to the LLM upon activating the project
# (contrary to the memories, which are loaded on demand).
initial_prompt: ""
# time budget (seconds) per tool call for the retrieval of additional symbol information
# such as docstrings or parameter information.
# This overrides the corresponding setting in the global configuration; see the documentation there.
# If null or missing, use the setting from the global configuration.
symbol_info_budget:
# list of regex patterns which, when matched, mark a memory entry as readonly.
# Extends the list from the global configuration, merging the two lists.
read_only_memory_patterns: []
# list of regex patterns for memories to completely ignore.
# Matching memories will not appear in list_memories or activate_project output
# and cannot be accessed via read_memory or write_memory.
# To access ignored memory files, use the read_file tool on the raw file path.
# Extends the list from the global configuration, merging the two lists.
# Example: ["_archive/.*", "_episodes/.*"]
ignored_memory_patterns: []
+99 -1
View File
@@ -1,7 +1,9 @@
# pragent — Design # pragent — Design
**Date:** 2026-08-04 **Date:** 2026-08-04
**Status:** Approved (brainstorming complete, ready for implementation planning) **Status:** Approved, then revised the same day after a prior-art review — see
`docs/research/2026-08-04-prior-art-ai-code-review.md`. **Phase 0 (evaluate the closest
existing tool) now precedes implementation.** Sections 6 and 7 below are the revisions.
## Problem ## Problem
@@ -182,6 +184,19 @@ which is what makes the confidence gates tunable rather than guessed.
## 5. Rollout phases ## 5. Rollout phases
**Phase 0 — Evaluate `ai-code-review` (do this first).** Red Hat's MIT-licensed
`ai-code-review` already implements Phases 12 of this design: four forge clients
(GitLab, GitHub, Forgejo, local git), six AI providers, CI integration, a committed
repo-context file. Run it on real repos in the Gitea setup for a week before writing
pragent code. *Done when:* we can name, from use rather than speculation, which of our
differentiators (tiering, analyzer bus, analytics) are worth building, and whether to
build them standalone, on top of that CLI, or as upstream contributions. See
`docs/research/2026-08-04-prior-art-ai-code-review.md` for the full assessment.
Phases 16 below stand as written **if** Phase 0 concludes we build standalone. If it
concludes we extend an existing base, Phases 12 mostly disappear and Phases 36 become
the whole project.
1. **Walking skeleton**`local` forge adapter, one analyzer, rules-only tiering, JSONL 1. **Walking skeleton**`local` forge adapter, one analyzer, rules-only tiering, JSONL
emitter. Runs on a local diff, prints findings. No network beyond the model API. emitter. Runs on a local diff, prints findings. No network beyond the model API.
2. **Gitea end-to-end** — Gitea adapter, Woodpecker step, PR comments, status checks. 2. **Gitea end-to-end** — Gitea adapter, Woodpecker step, PR comments, status checks.
@@ -192,6 +207,89 @@ which is what makes the confidence gates tunable rather than guessed.
5. **Second forge** — GitLab adapter, Jenkins runner recipe. Proves the abstraction. 5. **Second forge** — GitLab adapter, Jenkins runner recipe. Proves the abstraction.
6. **Analytics maturity** — OTel export, feedback loop, per-analyzer eval harness. 6. **Analytics maturity** — OTel export, feedback loop, per-analyzer eval harness.
## 6. Revisions from the prior-art review
Seven requirements the original design missed. They apply on every path — standalone,
fork, or contribution — so they are part of the design now rather than a backlog.
### 6.1 Prior-comment context and synthesis (highest priority)
Before the main review, fetch **all** existing comments and reviews on the PR, including
resolved ones, and compress them with a **cheap model** (Haiku-class) into a short summary
of what has already been said, fixed, or explicitly rejected. Feed that summary to the
analyzers.
This closes a hole the original design created. The cost model assumes ~1.7 review runs
per PR — every push re-reviews. Without prior-comment context, the second run repeats the
first run's findings and re-argues points a human already dismissed. That is the fastest
way for an AI reviewer to get muted, and it was designed in by accident.
Cost impact is favourable: one Haiku call (~$0.01) to compress a thread, against re-posting
findings that get ignored. Config: `reviewContext.enabled`, `reviewContext.synthesisModel`,
`reviewContext.maxTokens`.
### 6.2 Team context document
`teamContextFile` — a local path **or a URL** — carrying organization-wide review guidance:
security requirements, house conventions, compliance language. It outranks the repo
profile, so one document steers every repo without copying. This is the missing half of
"roll it out across many projects"; org config layering handles thresholds, this handles
judgment.
### 6.3 Deterministic skip conditions
The tier engine gains free, pre-model skips beyond paths and sizes: **draft/WIP MRs**,
`WIP` in the commit message, `wip/` branch prefixes, bot authors, tagged MRs. Each records
its `tier_reason` like any other rule. Reviewing a draft PR at full tier is a pure waste
the original rules would not have caught.
### 6.4 Self-hosted forge configuration
`forgeUrl` per adapter, `sslVerify`, `sslCertPath`. A self-hosted GitLab behind a corporate
CA is the normal case for the company deployment this targets, and the original design had
no way to express it.
### 6.5 Provider matrix as a requirement
The `ModelClient` port must be exercised by more than one provider before we claim it is a
port. Required: Anthropic, plus **Ollama or another local runtime** — "the diff never
leaves our network" is a procurement requirement for regulated repos, not a preference.
Vertex/Bedrock variants follow the same shape.
### 6.6 Per-provider input clamp
Alongside the file and line caps, clamp total diff characters with a provider-aware
default (roughly: 150k Anthropic, 200k Gemini, 100k OpenAI, 50k Ollama). The `oversized`
tier decides *strategy*; the clamp is the backstop that keeps a pathological diff from
blowing the context window regardless of tier.
### 6.7 Smaller additions
- **PR summary** alongside findings — a short description of what the change does
- **`--dry-run`** with mock responses, so a team can wire the pipeline before buying keys
- **Gitea adapter targets Forgejo too** — Forgejo is a Gitea fork with a compatible API;
this is close to free and doubles the addressable forges
- **Library-docs enrichment** (à la Context7) as a profile-enricher plugin, not core
## 7. What remains genuinely ours
After the revisions above, the differentiation is narrower and clearer than the original
design implied. Existing tools — Red Hat's CLI, CodeRabbit, Greptile, Qodo — do adapters,
providers, prompts, and skip logic. None of them do:
1. **Graded attention with recorded provenance.** Four tiers, each decision naming the rule
that fired. Everyone else has binary skip-or-review.
2. **An analyzer plugin bus.** Per-analyzer model, effort, and tool budget, fanned out over
a shared cached prefix — so a team adds a review dimension without forking anything.
3. **Measurement.** Run records, `explain`, `replay`, and finding-outcome feedback, which
together make false-positive rate per analyzer a number you can query and a prompt change
something you can A/B. No tool in this category can answer "did last week's prompt change
help?"
4. **Org-locked policy.** Config a repo cannot downgrade.
If Phase 0 shows an existing base covers everything else well, these four are the project —
and they may be worth contributing upstream rather than shipping standalone.
## Non-goals (for now) ## Non-goals (for now)
- Central webhook service (CLI-only until a team actually needs zero-setup onboarding) - Central webhook service (CLI-only until a team actually needs zero-setup onboarding)
@@ -12,6 +12,73 @@
--- ---
## ⚠️ Read this before dispatching Task 1
A prior-art review on 2026-08-04 found Red Hat's MIT-licensed `ai-code-review`
(https://gitlab.com/redhat/edge/ci-cd/ai-code-review) already implements Phases 12 of
this design: four forge clients, six AI providers, CI integration, a committed
repo-context file. **Phase 0 of the design is now "evaluate that tool for a week."**
See `docs/research/2026-08-04-prior-art-ai-code-review.md`.
Do not start Task 1 until Phase 0 concludes we are building standalone. If it concludes we
extend an existing base instead, most of the tasks below become unnecessary and the
project starts at the tier engine, analyzer bus, and analytics layer.
The amendments in the next section apply **whichever** path we take.
## Amendments (2026-08-04, post prior-art review)
Apply these on top of the task steps below. Where an amendment conflicts with the original
task text, the amendment wins.
**Task 4 (tier engine)** — add deterministic skip conditions before the path and size
rules, since they are free and catch cases the original rules miss:
- `ReviewContext` gains `isDraft: boolean`, `authorIsBot: boolean`, `commitSubject: string`,
`branch: string`, `labels: string[]`.
- `classify()` takes the context, not just the file list. New rules, highest precedence:
draft MR → `trivial` / `rule:draft`; `/^\s*(wip|draft)\b/i` on the commit subject →
`rule:wip_commit`; `wip/` branch prefix → `rule:wip_branch`; bot author →
`rule:bot_author`; a configured skip label → `rule:skip_label(<name>)`.
- Tests: one per rule, plus one asserting a **risk-path change in a draft MR still skips**
decide that deliberately and encode it (a draft is explicitly not ready for review; the
risk path will be caught when it opens).
**Task 5 (config)** — add keys: `skipLabels: string[]` (default `["skip-review"]`),
`maxChars: number` (provider-aware default, 150000 for Anthropic), `maxFiles: number`
(default 100), `excludePatterns: string[]` (lockfiles, minified, `dist/**`, `node_modules/**`,
`__pycache__/**`), `teamContextFile: string | null`, `dryRun: boolean`. Environment
variables become a config layer between the file and CLI flags: **CLI > env > file >
defaults** — CI systems supply secrets and URLs by environment, and the original design
had no way to receive them.
**Task 6 (analyzer runner)** — the prompt builder takes an optional `priorContext: string`
and, when present, includes it under a heading instructing the model not to repeat points
already made or explicitly rejected. Phase 1 always passes `undefined`; Phase 2 fills it.
Adding the parameter now costs one line and avoids reshaping the prompt module later.
**Task 6 (input clamp)** — before building the prompt, truncate the joined diff at
`config.maxChars` and drop files matching `excludePatterns`. When truncation happens, say
so in the prompt (`[diff truncated at N characters]`) rather than silently sending a
partial diff, and record it on the run record.
**Task 9 (run record)** — add `skipped: boolean`, `truncated: boolean`, and
`synthesis_cost_usd: number` (0 in Phase 1) to `RunRecord` and `toWireFormat`. Adding
fields later means old JSONL lines lack them, which breaks the analytics queries that are
the point of the format.
**Task 10 (CLI)** — add `--dry-run`, which uses a mock `ModelClient` returning a fixed
finding. This lets a team wire up the pipeline and verify plumbing before an API key
exists, and it makes the end-to-end test runnable in CI without network access.
**New Phase 2 task (not in this plan) — review-context synthesis.** Fetch all prior
comments including resolved ones, compress with a Haiku-class model, pass the result to
analyzers as `priorContext`. This is the highest-value item in the whole backlog: without
it, our own 1.7-runs-per-PR cost assumption means every push re-posts findings a human
already dismissed. Plan it before Phase 2's Gitea adapter work is considered done.
---
## Ground rules for the implementer ## Ground rules for the implementer
- **TDD, strictly.** Write the failing test, watch it fail, write the minimum code, watch it pass, commit. A step that says "run it and see it fail" is not decoration — a test that passes before the implementation exists is a broken test. - **TDD, strictly.** Write the failing test, watch it fail, write the minimum code, watch it pass, commit. A step that says "run it and see it fail" is not decoration — a test that passes before the implementation exists is a broken test.
@@ -6,6 +6,15 @@
- `docs/plans/2026-08-04-pragent-design.md` — why the system is shaped this way - `docs/plans/2026-08-04-pragent-design.md` — why the system is shaped this way
- `docs/plans/2026-08-04-pragent-implementation.md` — the authoritative task list, with full code - `docs/plans/2026-08-04-pragent-implementation.md` — the authoritative task list, with full code
> **Hold before dispatching.** A prior-art review (`docs/research/2026-08-04-prior-art-ai-code-review.md`)
> found an MIT-licensed tool that already covers Phases 12. The design now opens with
> Phase 0: evaluate it first. Do not dispatch Task 1 until that decision is made.
>
> When you do dispatch, the implementation plan's **Amendments** section changes tasks 4, 5,
> 6, 9, and 10. Add this line to the shared context block: *"Read the Amendments section at
> the top of the implementation plan before your task's steps. Where it conflicts with your
> task text, the amendment wins."*
This file does not restate the code. It supplies what a *cold* subagent needs that the implementation plan assumes: repo conventions, the exact slice of work, the boundaries it must not cross, and how the orchestrator decides whether the work is done. This file does not restate the code. It supplies what a *cold* subagent needs that the implementation plan assumes: repo conventions, the exact slice of work, the boundaries it must not cross, and how the orchestrator decides whether the work is done.
--- ---
@@ -0,0 +1,158 @@
# Prior art review — Red Hat `ai-code-review`
**Date:** 2026-08-04
**Subject:** https://gitlab.com/redhat/edge/ci-cd/ai-code-review (MIT, Python, 249 commits, 40 tags, created 2025-08-29)
**Also published as:** `ai-code-review-cli` on PyPI
**Why it matters:** it is the same shape as pragent's Phases 12, already shipped and maintained.
## What it is
An MIT-licensed CLI that reviews local changes, GitLab MRs, GitHub PRs, and Forgejo PRs.
Runs as a CI job or locally in a container. Python, LangChain for provider abstraction,
Pydantic for config and structured output, Jinja2 for rendering.
Repository structure (via GitLab API):
```
src/ai_code_review/
cli.py
core/
base_platform_client.py ← our ForgeAdapter port, same idea
gitlab_client.py github_client.py forgejo_client.py local_git_client.py
review_engine.py ← one engine, not a plugin bus
providers/
anthropic.py anthropic_vertex.py gemini.py gemini_vertex.py
ollama.py openai.py ← our ModelClient port, already six impls
models/ config.py, platform.py, review.py, settings_sources.py
utils/ prompts.py, review_templates/*.md.j2
.ai_review/
config.yml.example
project.md ← our repo profile, same idea, hand/agent-written
```
The convergence is not a coincidence — CLI-in-CI, an adapter per forge, a provider port,
and a committed repo-context file are what this problem shape pushes you toward. That is
mild evidence our architecture is right, and strong evidence we should not spend weeks
rebuilding the parts they have already debugged.
## What they have that our design missed
These are real gaps, ordered by how much they matter.
### 1. Review-context synthesis (the big one)
`enable_review_context` fetches **all** prior comments and reviews on the MR — including
resolved ones — and `enable_review_synthesis` runs a **cheap model first** (Haiku / Flash /
`gpt-4o-mini`) to compress them before the main review, so the reviewer does not repeat
suggestions that were already made, addressed, or explicitly rejected by a human.
Our design has nothing here, and the omission is worse than it looks. Our own cost model
assumes **1.7 review runs per PR** — every push re-reviews. Without prior-comment context,
run 2 repeats run 1's findings and argues with the human who dismissed them. That is the
single fastest way for an AI reviewer to get muted, and we designed it in by accident.
Their two-phase structure is also the cheap fix: a small model compresses the comment
thread, the expensive model sees the summary.
### 2. Team/org context file, loadable from a URL
`team_context_file` accepts a local path **or a remote URL**, and outranks the project
context. One company-standards document, fetched by every repo, no copying.
We have org config layering (thresholds, locked keys) but no shared *review guidance*
document. For the "roll out across many projects" goal, this is the missing half.
### 3. Skip conditions we did not consider
Draft/WIP MRs, "WIP" in the commit message, `wip/` branch prefixes, bot commits, tagged
MRs. All deterministic, all free. Our tier engine only looks at paths and sizes — it would
happily spend $2 reviewing a draft.
### 4. Enterprise self-hosting details
`gitlab_url` / `github_url` / `forgejo_url`, `ssl_verify`, `ssl_cert_path`. Obvious in
hindsight and completely absent from our design. A self-hosted GitLab behind a corporate
CA is the *normal* case for the company deployment we are targeting.
### 5. Provider breadth as a hard requirement
Six providers including Ollama (local) and both Vertex variants. For regulated repos
"the diff never leaves our network" is a procurement requirement, not a preference. Our
`ModelClient` port allows this, but our plan pins Anthropic and never states the matrix.
### 6. Adaptive input clamping
`max_chars` defaults per provider (Gemini 200k, Anthropic 150k, Ollama 50k, OpenAI 100k),
plus `max_files: 100` and `exclude_patterns`. Our `oversized` tier caps files and lines but
never clamps characters, and our exclusion list lives only in the tier rules.
### 7. Smaller things worth stealing
- **MR summary generation** alongside findings (`include_mr_summary`) — users like it
- **`dry_run`** with mock responses — lets a team wire the pipeline before buying keys
- **Context7 integration** — pulls official library docs into the review; a good argument
for our profile-enricher extension point
- **Forgejo support** — Forgejo is a Gitea fork with a compatible API, so our Gitea adapter
should target both and say so
## What we have that they do not
This is the honest differentiation list. It is shorter than the gap list, but it is real.
| Capability | Them | pragent |
|---|---|---|
| Attention control | Binary skip / review | Four tiers with a recorded `tier_reason` per decision |
| Review dimensions | One engine, one prompt template | Analyzer plugin bus: per-analyzer model, effort, tool budget |
| Cost engineering | Char clamps per provider | Shared cached prompt prefix across analyzers, per-PR spend ceiling |
| Measurement | None | JSONL/OTel run records, `explain`, `replay`, finding-outcome feedback |
| Governance | Priority order (repo can override anything) | Org-locked config keys a repo cannot downgrade |
The measurement column is the one that matters. They cannot answer "what is the
false-positive rate of our security review, and did last week's prompt change improve it?"
Neither can CodeRabbit or Greptile. That is a real gap in the category, not just in this tool.
## Verdict
**Do not start Phase 1 as written.** Insert an evaluation phase first.
The plan currently spends 11 tasks rebuilding a local git adapter, a provider client, a
CLI, and config loading — all of which this project already has, tested, in six provider
variants, across four forges. Building that from scratch to *then* discover it behaves like
theirs is the expensive way to learn something a week of use would tell us.
Three paths, in order of my preference:
**A. Evaluate first, then decide (recommended).** Run their tool on real repos in the Gitea
setup for a week. Two outcomes, both useful:
- It covers ~80% of the need → pragent shrinks to what is genuinely missing (tiering,
analyzer bus, analytics), possibly built *on top of* or *contributed to* their CLI.
- It falls short in ways we can name → we build, with requirements informed by a working
baseline instead of by speculation.
**B. Fork and extend.** Take their platform clients and provider layer, add the tier engine,
analyzer bus, and analytics. Saves most of Phases 12. Costs: Python instead of TypeScript
(fine — the reviewed repos are polyglot either way), a plugin layer retrofitted into
someone else's architecture, and ongoing divergence from an actively developed upstream
(249 commits since August 2025).
**C. Build as planned, steal the ideas.** Keep our architecture, fold in the seven gaps
above. Cleanest design, most work, and it means maintaining a forge and provider matrix
that someone else maintains for free.
I recommend **A**, because it is cheap and it makes the choice between B and C on evidence
rather than on taste. The design and plan in this repo are not wasted either way — the
tiering model, analytics schema, and analyzer contract are what we would add to any base.
## Requirement changes regardless of path
Fold these into the design now, since they apply to all three options:
1. Prior-comment context + cheap-model synthesis before the main review
2. Team/org context document, local path or URL, outranking the repo profile
3. Skip conditions: draft, WIP commit/branch, bot author, tagged MR
4. Self-hosted forge URLs, `ssl_verify`, custom CA path
5. Provider matrix as an explicit requirement, Ollama included, with the diff-never-leaves
the-network case called out
6. Per-provider character clamp alongside the existing file/line caps
7. `--dry-run` with mock responses
8. Gitea adapter targets Forgejo too