11 KiB
pragent pilot — central dashboard service
A read-only overview + per-repo / per-PR drilldown over the same SQLite
feedback DB the webhook writes, plus a small form to mutate .pr-review.json
on a covered repo via the Gitea contents API. Companion to the
webhook service; reuses the webhook image
(pragent-webhook:optin) — the pilot modules are baked into /app/pilot/,
and the dashboard is just python3 -m pilot.dashboard.
Architecture
Browser (tailnet)
│
▼
NodePort :31540 on 100.74.17.70 (kubernets)
│
▼
Service pragent-dashboard.pragent.svc.cluster.local (NodePort, ns pragent)
│
▼
pragent-dashboard pod (uid 10001, /data RO, no subprocess fan-out)
│
├── read /data/feedback.db (PVC pragent-feedback-data, RO)
├── GET .../repos/{o}/{r}/... (Gitea contents API, bot token)
└── PUT .../repos/{o}/{r}/contents/.pr-review.json
(edit form submit; Gitea commits a new sha)
Fail-soft. Nothing is ever written to local disk by the dashboard — the
SQLite file is read-only and .pr-review.json mutations go through Gitea's
contents API so the commit history records who changed what.
What it does
- Overview (
GET /): summary stats across all onboarded repos — total reviews, distinct PRs, finding counts by severity, false-positive / accepted-pattern scores (see "Feedback loop" in README-webhook.md), and a sparkline of review activity. - Repo drilldown (
GET /r/<owner>/<name>): per-repo PRs with their last-review status, finding counts, and links to PR-level drilldowns. - PR drilldown (
GET /r/<owner>/<name>/<index>): the bot's review(s) on that PR, inline findings, and reaction / resolved status harvested byfeedback_harvest.py. - Raw review (
GET /r/<owner>/<name>/<index>/raw): the markdown body of the most recent review, for copy-paste / diff-with-prose workflows. - Edit form (
POST /r/<owner>/<name>/edit): a small HTML page that loads the current.pr-review.jsonfrom the repo's default branch and lets the operator edit the JSON (validated, then PUT to Gitea contents API). This is how repo-localfocus/instructions/reviewers/severity_thresholdget tuned per-repo after seeing the feedback roll-up. - Auth (
GET /login,POST /login): single-user cookiepragent_dash=<token>whenPRAGENT_DASHBOARD_TOKENis set in the Secret. Unset = no auth (tailnet-only mode; logged on startup).
All routes return HTML (or plain text for /raw) with the same stylesheet
(/static/style.css).
Routes
| method | path | auth | description |
|---|---|---|---|
| GET | / |
yes | Overview |
| GET | /login |
no | Login form |
| POST | /login |
no | Sets the pragent_dash cookie |
| GET | /static/style.css |
no | Stylesheet |
| GET | /r/<owner>/<name> |
yes | Repo drilldown |
| GET | /r/<owner>/<name>/<index> |
yes | PR drilldown |
| GET | /r/<owner>/<name>/<index>/raw |
yes | Most recent review body as markdown |
| POST | /r/<owner>/<name>/edit |
yes | Edit .pr-review.json on the default branch |
There is no /health route — don't add one to the k8s probes without
updating pilot/dashboard.py (the handler returns 404 on unknown paths,
so a probe would loop forever).
Mutations flow through Gitea, not local fs
The edit endpoint reads the current .pr-review.json from
GET /repos/{o}/{r}/contents/.pr-review.json?ref=<default-branch>, lets
the operator edit it in a form (validated as JSON, length-capped per
field, no schema migration), and PUTs the new content back via the
contents API with a commit message like
pragent dashboard: update .pr-review.json. Every edit is a real Gitea
commit on the default branch, attributable to pragent-bot, and the
next webhook fire picks up the new config — no Pod restart, no image
rebuild, no pod-level state.
The /data mount is read-only (see the readOnly: true on the
volumeMount in ~/k8s/pragent-dashboard.yaml): the dashboard never
writes the SQLite file, only the webhook + the daily cronjob do, and
keeping it RO means a buggy deploy can't corrupt the harvested feedback.
Auth (single-user cookie)
PRAGENT_DASHBOARD_TOKEN in the pragent-webhook Secret is the value
the dashboard expects on the pragent_dash cookie. The login form
accepts it on POST / login and sets the cookie. Empty / unset env var
disables auth (logged on startup: auth OFF (tailnet-only mode)).
Generate a token once:
openssl rand -hex 16
Add to the Secret (the dashboard reads it as PRAGENT_DASHBOARD_TOKEN):
# base64 the token first
printf '<token>' | base64
# then JSON-patch the Secret
microk8s kubectl patch secret pragent-webhook -n pragent --type=json \
-p='[{"op":"add","path":"/data/PRAGENT_DASHBOARD_TOKEN","value":"<b64>"}]'
The token also lives in ~/.config/pragent/dashboard-token (mode 600) so
you can paste it into the login form without re-reading the Secret.
Deploy
The dashboard shares the webhook image, so there's nothing to rebuild.
After editing pilot/dashboard.py or pilot/dashboard_data.py, redo the
webhook image rebuild + containerd import (see README-webhook.md § "K8s
deployment") and roll both deployments.
K="microk8s kubectl"
# 1. (one-time) generate + persist the auth token
TOKEN=$(openssl rand -hex 16)
printf "%s" "$TOKEN" > ~/.config/pragent/dashboard-token
chmod 600 ~/.config/pragent/dashboard-token
printf "%s" "$TOKEN" | base64 \
| xargs -I{} $K -n pragent patch secret pragent-webhook --type=json \
-p='[{"op":"add","path":"/data/PRAGENT_DASHBOARD_TOKEN","value":"{}"}]'
# 2. apply the manifest
$K apply -f ~/k8s/pragent-dashboard.yaml
# 3. roll on image / code changes
$K -n pragent rollout restart deploy/pragent-dashboard
$K -n pragent rollout status deploy/pragent-dashboard --timeout=120s
$K -n pragent logs -f deploy/pragent-dashboard
K8s manifest at ~/k8s/pragent-dashboard.yaml. Key fields:
image: pragent-webhook:optin+imagePullPolicy: Never— local containerd only, same image as the webhook.nodeSelector: kubernetes.io/hostname: kubernets— pinned to the node holding the/dataPVC.securityContext: runAsNonRoot: true, runAsUser: 10001, runAsGroup: 10001, fsGroup: 10001— matches the image's USER directive; fsGroup makes the RO hostpath volume readable.volumeMounts.feedback-data.readOnly: true— dashboard is read-only over/data; mutations go through Gitea, not local fs.- No
readinessProbe/livenessProbe— the dashboard has no/healthroute. If you add one topilot/dashboard.py, add a probe here too. resources.requests: {cpu: 100m, memory: 256Mi}/limits: {cpu: 500m, memory: 512Mi}— read-heavy + tiny writes, no opencode subprocess fan-out, much smaller than the webhook.Service.type: NodePort,nodePort: 31540— between pages-proxy (30081) and browserless (30100), outside the 30096..30969 media range.
Smoke test
# 1. port-forward (in-cluster)
microk8s kubectl port-forward -n pragent svc/pragent-dashboard 8181:80 &
sleep 2
curl -s -o /tmp/dash-overview.html -w "HTTP %{http_code}\n" \
http://localhost:8181/
grep -o "<title>[^<]*</title>" /tmp/dash-overview.html
head -50 /tmp/dash-overview.html
kill %1
# 2. NodePort (host-side, Tailscale IP — only reachable on 100.74.17.70
# or 192.168.1.80; no public DNS yet)
curl -s -o /tmp/dash-nodeport.html -w "HTTP %{http_code}\n" \
http://100.74.17.70:31540/
grep -o "<title>[^<]*</title>" /tmp/dash-nodeport.html
# 3. auth: POST the login cookie value, capture into a jar
curl -sS -c /tmp/dash.jar -X POST \
-d "token=$(cat ~/.config/pragent/dashboard-token)" \
http://100.74.17.70:31540/login
# 4. authenticated overview
curl -sS -b /tmp/dash.jar -o /tmp/dash-auth.html -w "HTTP %{http_code}\n" \
http://100.74.17.70:31540/
grep -o "<title>[^<]*</title>" /tmp/dash-auth.html
# 5. pod logs
microk8s kubectl logs -n pragent -l app=pragent-dashboard --tail=50
The HTML should contain a <title> (whatever the dashboard renders) and
never Traceback or any Python exception output. A 401 / redirect
to /login on the unauthenticated GETs is expected when
PRAGENT_DASHBOARD_TOKEN is set.
Threat model / security notes
- Tailnet-only by default. NodePort 31540 is exposed on the Tailscale / LAN interfaces of kubernets (100.74.17.70, 192.168.1.80). No public DNS, no Caddy route yet — keep it that way until the auth story is solid. When you do add a public route, terminate TLS at Caddy (wildcard cert via Cloudflare DNS-01) and rely on the cookie auth + Logto SSO gating pattern from the code-server / minecraft-sso / livecodes setups.
PRAGENT_BOT_TOKENis Gitea Write scoped to onboarded repos, so a successful auth bypass on the dashboard is Gitea repo write access, not just read. The cookie is the only auth factor — there is no CSRF token in the current implementation (single-operator trust model). If the dashboard goes public, add a CSRF token to the edit form (hidden input + double-submit cookie) before opening it up.- Read-only
/datamount. The dashboard can't corrupt the harvested SQLite file even if it's compromised. The webhook and the daily cronjob are the only writers. uid 10001+runAsNonRoot: true. No host-level escalation if the dashboard is popped — it has no caps, no/procmounts, no NetworkPolicy exemption (cluster default deny applies until you grant egress togitea-http.gitea.svc.cluster.local:3000and100.74.17.70:8787if you ever need the headroom proxy — currently the dashboard doesn't).- No author-controlled input is
eval-ed. The edit form parses the JSON, validates types / lengths, and re-serialises before the Gitea PUT. The review-side hostile-input concerns fromREADME-webhook.md§ "Threat model" do not apply to the dashboard — the dashboard is a read-mostly viewer over already-harvested, already-posted data.
Known limitations (pilot)
- Single-user cookie auth (no per-user sessions, no Logto SSO yet).
- No CSRF protection on the edit form (intentional — single-operator trust model; add before going public).
- Read-only
/datameans the dashboard can't backfill the DB if the cronjob is paused; if you turn off feedback harvesting (PRAGENT_FEEDBACK_DB=empty in the webhook), the overview is empty. - No
/healthroute — if the dashboard process wedges on a Gitea hang, k8s won't restart it. Add a/healthroute topilot/dashboard.py- a probe here before relying on this in production.
- The overview is a single-process render over a SQLite file that the daily cronjob also writes. A long Gitea hang during a page render can stall the dashboard until the client request times out (30 s). The underlying SQLite reader is read-only and concurrent-safe, so no data corruption — just a slow page.