5d44121b28
Two llm_as_judge evaluators score the review generation directly: a NUMERIC 0-1 on finding actionability, a BOOLEAN on whether the summary agrees with the findings. Both run on every observation whose trace name is pr-review or opencode-review. The judge is kimi-k2.7-code through the headroom hub. Local Ollama returns Anthropic-format responses but the thinking blocks lack the signature field Langfuse Zod schema requires; the evaluator preflight fails as Invalid JSON response. A small judge-proxy pod on 8802 forwards to the hub and patches every thinking block with a synthetic signature before returning. Trace + generation output now includes the findings themselves (capped at 25) rather than just the count, so a judge has something to grade. generation input/output mirrors the trace so an observation-level evaluator can read them. Idempotent: existing evaluators and rules are skipped on re-run, not duplicated. The connection is upserted on provider.