docs: initial pragent design and README

Extensible PR review framework: CLI-in-CI delivery, pi SDK agent loop,
deterministic-plus-LLM tiering, committed repo profile, JSONL/OTel analytics.
Design doc records the decisions, cost model, and rollout phases.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011Ye1KNFMkkUtmzTypHXkoK
This commit is contained in:
Marcos
2026-08-04 16:35:35 +00:00
commit 1c0c576e40
3 changed files with 292 additions and 0 deletions
+85
View File
@@ -0,0 +1,85 @@
# pragent
An extensible, forge-agnostic PR review framework. Not a product — a toolkit that teams
extend with their own review dimensions.
**Status:** design approved, implementation not started. See
[`docs/plans/2026-08-04-pragent-design.md`](docs/plans/2026-08-04-pragent-design.md).
## What it is
`pragent` runs as a CI step. It reads a pull request, decides how much attention the
change deserves, runs the analyzers that apply, and posts ranked findings back to the
forge.
```
pragent init # one-time repo scan → .pragent/profile.yml (committed, reviewable)
pragent review # the CI step: tier → analyze → aggregate → publish
pragent explain # why did this PR get this tier / these findings?
pragent replay # re-run a past PR against a new prompt or model (the eval loop)
pragent doctor # config, credentials, and adapter health
```
## Why not CodeRabbit / Greptile / Qodo
Those are good products with fixed review dimensions and per-seat pricing. `pragent`
targets the case where a platform team needs to **add its own dimensions** — an internal
compliance rule, a service-catalog ownership check, a house performance idiom — without
forking a vendor's reviewer. Cost lands in the same range (~$25/dev/month at 350 PRs/mo
for 20 devs), but the analyzers, the data, and the analytics are yours.
## Attention tiers
Every PR is classified before any expensive work happens. Deterministic rules decide
first; an ambiguous case gets one cheap model call as tie-breaker.
| Tier | What it means | Cost/PR |
|---|---|---|
| `trivial` | lockfile bumps, generated code, docs typos | ~$0.005 |
| `lite` | small change, no risk paths | ~$0.08 |
| `full` | the default for real changes | ~$0.802.00 |
| `oversized` | too big to review whole; structural summary + deep pass on the hot subset | ~$5 ceiling |
Every tier decision records *why*, so a surprising outcome is explainable rather than
mysterious.
## Extension points
Five, all documented in the design doc. Teams override or add; nobody forks.
1. **Analyzers** — drop a YAML + prompt in `.pragent/analyzers/`, or install from npm
2. **Forge adapters** — Gitea, GitLab, GitHub, local diff
3. **Tier policy** — thresholds and the path risk map, per repo or per org
4. **Profile enrichers** — extend what `pragent init` learns about a repo
5. **Emitter sinks** — JSONL by default, OpenTelemetry, or your own
Org config can lock keys, so a repo cannot quietly disable the security analyzer.
## Design principles
- **Polyglot by construction.** Language knowledge lives in the repo profile, not in the
reviewer. A new language is a profile change, not a core change.
- **One shared prompt prefix.** All analyzers for a PR share a byte-identical cached
prefix. This is what makes fan-out affordable; it is enforced, not hoped for.
- **Everything is traceable.** Tier reasons, token counts, cost, latency, and finding
outcomes are recorded per run. False-positive rate is measurable per analyzer.
- **Fail open.** A budget ceiling or an analyzer crash yields a partial review with a
clear note, never a blocked pipeline with no explanation.
## Stack
TypeScript + Node, built on the [`pi`](https://github.com/badlogic/pi-mono) agent SDK.
Shipped as an npm package and an OCI image, so CI runners need no local Node install.
## Roadmap
1. Walking skeleton — local diff, one analyzer, rules-only tiering
2. Gitea end to end — adapter, Woodpecker step, PR comments, status checks
3. Profile + full tier — `pragent init`, shared-prefix caching, analyzer fan-out
4. Extensibility hardening — plugin loading, config layering, `explain` / `replay`
5. Second forge — GitLab adapter, Jenkins recipe
6. Analytics maturity — OTel export, feedback loop, eval harness
## License
TBD.