Files
pragent/pilot/README-langfuse.md

5.9 KiB

pragent → Langfuse

Every review the pilot runs ships one trace to a self-hosted Langfuse. The review body already prints a usage table, but that table lives and dies inside one Gitea PR. Langfuse is where the same numbers become a trend: tokens per review, latency per model, equivalent cost per repo, and how those move when the model or the tiering changes.

The ollama / claude split

Both paths route through the same headroom proxy, so the provider prefix does not distinguish them — headroom/claude-sonnet-5 is Claude spend, headroom/glm-5.2:cloud is not. The split is keyed off the bare model name and lands on the trace's environment:

resolved model environment
headroom/claude-sonnet-5 claude
claude-opus-5 claude
headroom/glm-5.2:cloud ollama
headroom/MiniMax-M2.7 ollama
vllm-qwen38/qwen3.8-27b ollama

Langfuse takes an environment selector on every view, filter and cost breakdown, so the two spend stories stay separate inside one project — one key pair to rotate instead of two. Tags carry the finer cut: provider:headroom, model:<bare>, engine:opencode, repo:<owner/name>, lens:<id> per fan-out lens.

Each trace metadata record also includes the completed iteration count, tool calls, per-iteration token details, and whether a configured budget stopped the run. A capped run is tagged in the review body and can be filtered in Langfuse with cap_hit / cap_reason metadata.

To split into two projects later, point LANGFUSE_PUBLIC_KEY / LANGFUSE_SECRET_KEY at the second project on whichever deployment runs the Claude path. Nothing in the code needs to change.

What a trace carries

  • trace pr-reviewsessionId = owner/repo#index, so every push to one PR groups together. Input is the PR identity; output is the summary + finding count; metadata carries steps, duration, severity counts and the provider's own reported cost.
  • generation opencode-reviewmodel, usageDetails, costDetails.

usageDetails.input is the uncached input. opencode reports cache_read inside input, and Langfuse sums the keys it is given, so passing both verbatim would bill the resent prefix twice.

How cost is priced

Langfuse has no price table of its own here — we compute the number and ship it as costDetails.total, so what Langfuse charts is exactly what cost_model.PRICES says.

A model that genuinely bills (claude-*, gpt-*, gemini-*, grok-*) is priced as itself: basis actual.

A model that costs nothing through the headroom proxy is priced against a comparison target instead: basis equivalent:<target>. That covers the models absent from PRICES (MiniMax-M2.7 — which is what the webhook actually runs — and glm-5.2:cloud) as well as entries priced at all zeros (the self-hosted vLLM qwen3.8-27b). Without this Langfuse would show a flat $0.00 line, since the pilot's own path is free.

The target follows the same precedence as the review body, so the PR and Langfuse never disagree:

.pr-review.json:cost_target  >  PRAGENT_PRICE_TARGET  >  claude-sonnet-5

An equivalent cost is a hypothetical, not money spent, so every trace is tagged cost:actual or cost:equivalent:<target> and the generation metadata carries cost_basis. Filter on it before reading any cost chart as spend.

If the comparison target itself is unknown, the trace ships usage with no cost block — better no number than a wrong one.

Anthropic prices in cost_model.PRICES were fetched 2026-08-18; re-check them before quoting anything externally.

Configuration

env meaning
LANGFUSE_HOST http://langfuse-web.langfuse.svc.cluster.local:3000
LANGFUSE_PUBLIC_KEY pk-lf-…
LANGFUSE_SECRET_KEY sk-lf-…
LANGFUSE_TIMEOUT seconds, default 5
LANGFUSE_DEBUG 1 to log ingestion failures to stderr

Unset host or either key ⇒ emission is a silent no-op. That is the default, so a checkout without Langfuse behaves exactly as before.

Fail-open

langfuse_trace is stdlib-only (urllib) and every entry point swallows its own exceptions; _emit_langfuse in ai_review.py wraps even the import. A Langfuse outage cannot fail, delay past LANGFUSE_TIMEOUT, or alter a review.

Both token-spending exit paths emit — the normal post and the salvage path where the agent produced unparseable output. That run cost the same as a clean one, and is precisely the failure worth trending.

Deployment

Cluster side lives outside this repo: ~/k8s/langfuse.yaml (ClickHouse + web + worker, reusing the gitea postgres, gitea valkey and minio), ~/k8s/oauth2-proxy-langfuse.yaml (the Logto gate), and ~/k8s/langfuse-setup.sh, which provisions the database, the bucket, the secrets, and wires pragent-webhook with the three env vars above.

The UI is at https://langfuse.marcospaulo.dev.br:

browser -> Caddy (VPS, TLS, DNS-01) -> tailscale
        -> 100.74.17.70:30361 -> oauth2-proxy (Logto, email allowlist)
        -> langfuse-web (ClusterIP)

Logto sits at both layers off one app (langfuse, two redirect URIs): the proxy gates the domain, and Langfuse's own NextAuth uses the same Logto as a custom OIDC provider, so the inner login is a silent redirect rather than a second password.

pragent does not go through any of that. It posts to langfuse-web.langfuse.svc.cluster.local:3000 from inside the cluster, on API-key auth — putting ingestion behind an interactive SSO gate would break it on the first review.