Commit Graph

90 Commits

Author SHA1 Message Date
Claude 0d73a075cd feat(opencode): bump qwen3.8-27b context limit to 57344 2026-08-22 19:21:53 +00:00
Claude e5d4e252de feat(opencode): multi-provider support with local provider for qwen3.8-27b 2026-08-22 19:13:06 +00:00
Claude 92020d4d46 feat(ai_review): route per-repo model override through PRICES provider field 2026-08-22 19:13:04 +00:00
Claude e344831b05 feat(cost_model): add provider field on Price + register qwen3.8-27b (local) 2026-08-22 19:13:04 +00:00
Claude 2ba3ccbd95 docs(pilot): dashboard now behind Logto SSO 2026-08-22 18:05:02 +00:00
Claude 99255bb167 feat(dashboard): drop token auth, trust oauth2-proxy X-Forwarded-User 2026-08-22 18:03:33 +00:00
Claude c9809bce9b Merge deploy agent: k8s manifest + README 2026-08-22 17:01:34 +00:00
Claude 5fe83b0b2a k8s(pilot): dashboard deploy manifest + README; pin NodePort 31540 2026-08-22 17:01:34 +00:00
claude 218a8dc271 docs(pilot): dashboard README 2026-08-22 17:00:12 +00:00
Claude e688ea61c5 Merge dashboard agent: data + HTTP server + mutation endpoints 2026-08-22 15:17:58 +00:00
Claude dd49d83933 feat(dashboard): model select dropdown in edit form
The repo edit form already renders a <select id="model"> with one
<option value="…"> per cost_model.PRICES key (sorted) plus a
"— keep current —" placeholder. This commit adds the test that pins
that behaviour: every key must appear as an <option>, no others.

1 test under tests/pilot/test_dashboard_select.py.
2026-08-22 15:15:29 +00:00
Claude 6f7957010e feat(dashboard): mutate .pr-review.json via Gitea contents API
POST /r/<owner>/<name>/edit reads the current .pr-review.json via the
Gitea contents API (404 → start from {}), updates static_message and
model, then PUTs the file back as the bot identity. Static message is
stripped and capped at 400 chars; model is validated against
cost_model.PRICES (unknown → flash + redirect, no save). CSRF token is
the dashboard auth token (same value, simplest binding).

Helpers (_http) moved to urllib.request at module top so tests can patch
the entry point; the helper returns the Gitea status code so 404 (file
missing) cleanly omits sha on the create-PUT path.

6 tests under tests/pilot/test_dashboard_edit.py covering the happy
path, the 400-char cap, the unknown-model rejection, auth + CSRF gates,
and the file-missing creation path.
2026-08-22 15:15:04 +00:00
Claude be24ef245c feat(dashboard): stdlib HTTP server with overview/repo/PR pages
Mirrors webhook_server.py's BaseHTTPRequestHandler + ThreadingHTTPServer
shape. Pure stdlib, no pip deps. Routes:

  GET  /                            overview (totals + 7-day sparkline + top repos)
  GET  /r/<owner>/<name>            repo summary (severity histogram + top findings)
  GET  /r/<owner>/<name>/<index>    one PR's findings
  GET  /r/<owner>/<name>/<index>/raw raw Markdown body (via Gitea contents API)
  GET  /static/style.css            dark-mode stylesheet
  GET  /login                       login form
  POST /login                       compare token, set HttpOnly+SameSite cookie
  POST /r/<owner>/<name>/edit       (Tasks C+D)

Auth: when PRAGENT_DASHBOARD_TOKEN is set, every route except /login and
/static/* requires Cookie: pragent_dash=<token>. Unset → tailnet-only.

All HTML rendered via string.Template; every dynamic value is escaped
with html.escape(..., quote=True). No .format, no f-string templates.

12 tests under tests/pilot/test_dashboard.py.
2026-08-22 15:11:36 +00:00
Claude bb9b6aa12d feat(dashboard): read-only query module over feedback SQLite
Three pure functions — overview / repo_summary / pr_summary — that open
the feedback SQLite via feedback.init, run their queries, and return
plain dicts/lists. All three tolerate a missing or empty DB by returning
a zero-shaped dict.

Cost is hardcoded 0.0: per-review usage:cost isn't stored, only the raw
review rows are. Surfacing a rolled-up dollar figure without telemetry
would be guessing, so we don't.

17 new tests under tests/pilot/test_dashboard_data.py.
2026-08-22 15:07:57 +00:00
Marcos fe5bebb4cf Merge agent-B: cost-line fix + static_message + per-repo model override
# Conflicts:
#	pilot/ai_review.py
#	tests/pilot/test_ai_review.py
2026-08-22 15:01:38 +00:00
Marcos d0f99b9763 feat(ai_review): parse .pr-review.json:model as per-repo override
Repos that need to pin the review engine (e.g. 'this project requires
claude-sonnet-5 for the budget line' or 'route everything through
gpt-5.6-luna for now') can declare a top-level 'model' string in
.pr-review.json. The parser validates the value against cost_model.PRICES
(lazy import — ollama path stays dep-free) and silently drops unknown
values with a stderr pointer to the valid key set so a typo in the
config file surfaces in the logs instead of silently falling back.

The orchestrator gains a small _resolve_display_model(base, config)
helper that implements a 3-way precedence:
  1. OPENCODE_MODEL env (operator override, used verbatim)
  2. config['model'] (per-repo override)
  3. f'headroom/{base}' (default)

review_pr resolves once (lazy fallback before config is loaded) and
re-resolves after .pr-review.json is fetched, then threads the result
into the opencode subprocess, REVIEW_HEADER, and the cost-line
parenthetical. Same single value everywhere — no more mix of bare
OLLAMA_MODEL id in the header and a stale free-tier literal in the cost
line.

Tests cover parsing acceptance, parsing rejection (capsys stderr),
type validation, precedence in all 4 (env×config) combinations, and an
end-to-end sanity check that format_review_body shows the override and
not the base id.
2026-08-22 14:46:47 +00:00
Marcos f3a09b9397 Merge agent-A: move feedback*.py from WIP into tree 2026-08-22 14:46:46 +00:00
Marcos 4e5f43ada7 feat(feedback): move feedback poster from WIP into pilot/ 2026-08-22 14:46:07 +00:00
Marcos 8472f35a58 feat(feedback): move feedback analyzer from WIP into pilot/ 2026-08-22 14:46:07 +00:00
Marcos 69e1fc06a2 feat(feedback): move feedback harvester from WIP into pilot/ 2026-08-22 14:46:07 +00:00
Marcos 6cdccb48ad feat(feedback): move feedback storage layer from WIP into pilot/ 2026-08-22 14:46:04 +00:00
Marcos 7f37a36722 feat(ai_review): parse .pr-review.json:static_message + render as banner
Repos can pin a free-text notice (e.g. 'this repo is in maintenance mode',
or 'reviewers: focus on the public API only for this quarter') in
.pr-review.json:static_message. The string is stripped and capped at 400
chars (mirror of the existing instructions cap), then rendered as a
Markdown blockquote (> {msg}) directly under the REVIEW_HEADER so it
surfaces on every review without scrolling.

Plumbing: parse_repo_config exposes 'static_message'; format_review_body
accepts a static_message kwarg and inserts the blockquote before the
'### Summary of Changes' section; review_pr threads
config.get('static_message') to both call sites (salvage path + happy
path). Empty / non-string values are silently dropped, mirroring the
parser's 'ignore blank' handling for every other text field.
2026-08-22 14:43:59 +00:00
Marcos b6b8173ccb fix(ai_review): show actual model in cost line (was hardcoded glm-5.2:cloud)
The webhook pod now routes through headroom's MiniMax-M2.7 endpoint, but the
cost line in the AI-usage collapsible still read 'headroom glm-5.2:cloud'.
Resolve a single display_model at the top of review_pr (OPENCODE_MODEL env
wins, else headroom/{OLLAMA_MODEL}) and pass it to:
  * the opencode subprocess (was already doing this on the same line, now
    sharing the value)
  * format_review_body so REVIEW_HEADER also reflects the actual run
  * _render_collapsible_usage so the parenthetical reads
    '({display_model} — free tier)' or '({display_model} — billed)'.

Test additions in tests/pilot/test_ai_review.py cover:
  * the parenthetical picks up the passed-in model verbatim
  * the full provider prefix survives (headroom/<id>) for the opencode path
  * nonzero cost flips the inner clause from 'free tier' to 'billed'
  * the existing nonzero-cost assertion flips to assert 'billed' instead
    of dropping the parenthetical entirely
2026-08-22 14:41:56 +00:00
Marcos a9e1b7ddfc fix(opencode): patch apiKey from PRAGENT_MODEL_API_KEY at install_config
opencode's @ai-sdk/anthropic provider sends the config's apiKey as the
x-api-key header. The committed opencode.json carries apiKey="ollama"
(placeholder) so the repo can be public. When the headroom upstream
switches to an auth-gated provider (e.g. MiniMax), the placeholder
returns 'No credentials found'. install_config now also patches
options.apiKey from PRAGENT_MODEL_API_KEY when set, mirroring the
existing baseURL patching.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-22 13:12:46 +00:00
Marcos 4c06a9ab3c chore(opencode): swap glm-5.2:cloud → MiniMax-M2.7
headroom-hub 8789 was repointed to api.minimax.io/anthropic.
glm-5.2:cloud is no longer a registered model there.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-22 03:45:09 +00:00
claude d38e1c8693 test(opencode_review): skip posthash golden when feedback module absent
The golden-vector test references the unmerged pilot/feedback*.py module
(WIP, untracked). Without an explicit skip, main branch fails 1 test.
2026-08-22 01:43:05 +00:00
Marcos 8e8ae54669 Merge feature/pragent-update: 16-task pragent update
Removes AI-REVIEW/AI-USAGE label gates (replaced with .pr-review.json:enabled=true on base ref).
Adds multi-provider cost table, fmt_tokens, trivial/info severities, walkthrough/risk_verdict/test_coverage
sections, per-PR merge confidence 1-5, always-on usage rendering, README rewrite.

362 tests pass. Caveat: pilot/feedback*.py modules referenced by webhook_server.py remain untracked.
2026-08-22 01:38:49 +00:00
claude 661199dad2 chore(cost_model): drop stale AI-USAGE label references 2026-08-22 01:37:23 +00:00
claude 3e03fb80a7 docs: onboarding uses .pr-review.json:enabled, not labels 2026-08-22 01:33:42 +00:00
claude 2432228d68 feat(webhook): gate on .pr-review.json:enabled, drop labels 2026-08-22 01:30:07 +00:00
claude 2c7d4803f1 feat(webhook): is_repo_enabled reads .pr-review.json:enabled 2026-08-22 01:24:10 +00:00
claude 86352a3771 refactor(review): always render usage; drop report_usage flag 2026-08-22 01:18:34 +00:00
claude e1b74d982d refactor(ai_review): drop AI_REVIEW/AI_USAGE label plumbing 2026-08-22 01:13:59 +00:00
claude 33e4c16782 feat(usage): multi-provider equivalent cost table 2026-08-22 01:06:02 +00:00
claude 979c93bdbb feat(review): render walkthrough + risk_verdict + test_coverage 2026-08-22 00:58:51 +00:00
claude 66628aae8d fix(opencode_review): harden _synthesize_summary_fields (none-safe, dedupe emoji map) 2026-08-22 00:52:35 +00:00
claude e5a6e8923d feat(opencode_review): python fallback for summary fields 2026-08-22 00:46:08 +00:00
claude 2e982846d9 feat(review): walkthrough/risk_verdict/test_coverage schema 2026-08-22 00:36:50 +00:00
claude 72b77a96e0 fix(confidence): pass multi_lens_observed kwarg to merge_confidence 2026-08-22 00:29:28 +00:00
claude f3a125666b feat(ai_review): per-PR merge confidence 1-5 in header 2026-08-22 00:23:24 +00:00
claude 9dae850887 test(config): fix cap-at-12 test to actually exercise the cap 2026-08-22 00:19:25 +00:00
claude 0f7903377a feat(config): parse enabled + compare_against 2026-08-22 00:12:06 +00:00
claude 0b295e2443 fix(severity): extend badge label set + update SYSTEM_PROMPT 2026-08-22 00:06:22 +00:00
claude 23ec1bf74a feat(severity): add trivial + info levels 2026-08-21 23:59:48 +00:00
claude 2cf4bdbfe8 feat(ai_review): render humanized token counts in usage + inline 2026-08-21 23:54:10 +00:00
claude 3bcf825104 feat(ai_review): fmt_tokens() humanizes token counts 2026-08-21 23:46:24 +00:00
claude 67339da8d0 feat(cost-model): add GPT, Gemini, Grok prices 2026-08-21 23:42:07 +00:00
claude 99014a4cd3 chore(gitignore): exclude .worktrees/ 2026-08-21 23:39:44 +00:00
claude f52b8d7803 docs(plans): 2026-08-21 pragent update implementation plan 2026-08-21 23:39:11 +00:00
claude f9f6ab4bf0 docs(plans): 2026-08-21 pragent update design
- Repo opt-in (.pr-review.json:enabled) replaces AI-REVIEW/AI-USAGE labels.
- Token humanization (fmt_tokens), multi-provider equivalent cost (Claude,
  GPT, Gemini, Grok), richer review summary (walkthrough + risk verdict +
  test coverage), trivial/info severity levels, per-PR merge confidence 1-5.
- Defers reachability severity demotion, rules-mining from feedback, and
  heavy infra (sequence diagrams, T-rex, cross-repo) to future work.
2026-08-21 23:32:49 +00:00