Files
pragent/pilot
Marcos 2613b3e3af fix(cost-model): calibrate against the first measured review
PR #7 ran under the AI-USAGE label and reported real numbers: 28 agent steps,
348s, 2,071,025 input / 17,303 output tokens, and zero cache reads or writes.
The model predicted ~$0.73 on Opus 5 for that tier. The measurement prices it
at $10.79 — the model was ~15x low.

Two wrong assumptions:

- Step count and per-step growth. `full` assumed 12 steps and 1,200 tokens per
  tool result; the run did 28 steps averaging ~3,300. Cost is roughly quadratic
  in steps, so this compounds. Tier defaults are re-derived from the measured
  per-step growth rather than from guesses.
- Caching. The model defaulted to prompt caching on. The headroom/glm-5.2 path
  reports 0 read / 0 write, so the stable prefix is paid at full input price on
  every step. Budget with caching off until that column is nonzero.

Adds OBSERVED_RUNS as an append-only calibration anchor, an observed-runs
section in the report, and a regression test asserting the model stays within
2.5x of the measurement — so the next drift is caught by the suite rather than
by a surprising invoice.

Corrected blended figures at 350 PRs/month: ~$1,740 Opus 5, ~$1,755 GPT-5.6
Sol, ~$696 Sonnet 5, ~$348 Haiku 4.5, ~$70 GPT-5.6 Luna.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01B11e8TZZxJyzHW7jj7KWUN
2026-08-18 05:08:22 +00:00
..

pragent pilot — AI Review bot

A minimal AI code-review bot for Gitea, running as a CI step on the existing act-runner. This is the pilot — a small, self-contained reviewer that predates the full pragent framework (whose design lives in docs/plans/2026-08-04-pragent-design.md). The framework will later absorb this; until then, this is what runs.

How it works

  1. You add pragent-bot to a repo and commit .gitea/workflows/ai-review.yml.
  2. On a PR, you add the AI-REVIEW label.
  3. Gitea Actions runs the workflow on the act-runner; it fetches the PR diff, asks glm-5.2:cloud (on-network via the headroom proxy) to review it, and posts the findings back as a PR review authored by pragent-bot.
  4. Remove the label to stop re-reviews on further pushes.

Fail-open: the job always exits 0 and never blocks CI. Errors become a short "review failed" comment.

Onboard a repo (3 steps)

1. Add pragent-bot as collaborator

Repo → Settings → Collaborators → Add → pragent-bot → permission Write. (Write is required to post reviews/comments.)

Or via API (with an admin/owner token):

curl -X PUT -H "Authorization: token $OWNER_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"permission":"write"}' \
  "http://100.74.17.70:30000/api/v1/repos/OWNER/REPO/collaborators/pragent-bot"

2. Add the PRAGENT_BOT_TOKEN secret

Repo → Settings → Actions → Secrets → New secret → name PRAGENT_BOT_TOKEN, value = the bot's access token (ask the platform admin; stored mode-600 at ~/.claude/.pragent-bot-token on the admin host).

3. Commit the workflow

Copy pilot/workflow-template.yml into the target repo as .gitea/workflows/ai-review.yml and commit it. That's it.

Use it

Open a PR (or push to an open one), add the AI-REVIEW label. The review appears within ~3090s depending on diff size and model latency.

What's intentionally NOT in the pilot

Deferred to the full framework (by design, see the design doc):

  • Attention tiering (trivial/lite/full/oversized) and per-tier cost control.
  • Multiple analyzer fan-out over a shared cached prompt prefix.
  • Prior-comment synthesis (so each push re-posts; the latest review is tagged with the head SHA so it's easy to spot).
  • Inline line comments and status checks.
  • pragent explain / replay / analytics JSONL.
  • A second forge (GitLab) and the provider matrix.

Pieces

File Role
pilot/ai_review.py The reviewer script (stdlib only). Single source of truth — fetched at runtime by each repo's workflow.
pilot/workflow-template.yml The Gitea Action consumers copy into .gitea/workflows/ai-review.yml.
tests/pilot/test_ai_review.py Unit tests for the pure helpers (no network).

Run the tests

cd ~/Projects/pragent
PYTHONPATH=pilot python3 -m pytest tests/pilot/    # if pytest available
# or, without pytest:
python3 - <<'PY'
import os, sys, importlib.util
sys.path.insert(0, os.path.abspath("pilot"))
import ai_review  # noqa: F401
spec = importlib.util.spec_from_file_location("t", "tests/pilot/test_ai_review.py")
m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m)
fails = 0
for n in sorted(x for x in dir(m) if x.startswith("test_")):
    try: getattr(m, n)(); print("PASS", n)
    except Exception as e: fails += 1; print("FAIL", n, e)
print("failed:", fails)
PY

Configuration knobs (env in the workflow)

Env Default Purpose
OLLAMA_MODEL glm-5.2:cloud Model id passed to the headroom proxy.
OLLAMA_MAX_TOKENS 6000 Output token cap.
DIFF_MAX_CHARS 150000 Diff truncation cap (with a noted truncation marker).
OLLAMA_URL http://100.74.17.70:8789 headroom proxy (tailnet). If the act-runner can't reach the tailnet IP, expose 8789 as an in-cluster Service+Endpoints and set this to the cluster DNS name.