6 Commits

Author SHA1 Message Date
masi 9a9923d034 Merge pull request 'test: eval judges e2e' (#14) from test/eval-judges-trigger into main 2026-09-01 00:26:59 +00:00
Claude 3f30bdd450 merge: resolve eval judge rule conflict 2026-09-01 00:26:25 +00:00
Claude 5b8be61e2f test: trigger review after DB rule insert 2026-08-31 19:19:59 +00:00
Claude 3109bd7c2d pilot(eval): switch rule target from observation to trace
The standard /api/public/ingestion path feeds only the trace-upsert
queue; evalService.createEvalJobs only dispatches targetObject in
{TRACE, DATASET}. Observation rules fire exclusively from the OTel
pipeline, which this pilot does not use. The trace body already carries
review input/output via langfuse_trace, so a trace rule sees the same
material an observation rule would.
2026-08-31 19:18:59 +00:00
gitea_admin d746b1fdc2 test: trigger review for eval judges 2026-08-31 18:37:41 +00:00
gitea_admin 193a90e63e Merge pull request 'feat(eval): LLM-as-judge evaluators, dataset item fixes, and Experiments runs' (#13) from fix/dataset-item-url-safe-ids into main 2026-08-31 18:34:02 +00:00
4 changed files with 25 additions and 10 deletions
+1
View File
@@ -0,0 +1 @@
# judge trigger 1788203999
+2
View File
@@ -185,3 +185,5 @@ review time.
## License
Not yet chosen. Until one is added, no reuse rights are granted.
_pilot eval judges test 1788201461_
+13 -7
View File
@@ -208,18 +208,24 @@ def ensure_evaluators() -> dict:
# ---------------------------------------------------------------------------
def rule_body(name: str, judge_name: str, sampling: float) -> dict:
"""POST /evaluation-rules shape for an LLM-as-judge observation rule.
"""POST /evaluation-rules shape for an LLM-as-judge trace rule.
The judge is referenced by `name`+`scope`, not by id — ids name specific
versions, names name the evaluator across versions. Mapping is required at
both the rule root (the server validates it there) and inside `evaluator`
(the API echoes it back). Filter is on `traceName` because that is the only
stringOptions column the observation-rule schema exposes.
Target is `trace` rather than `observation` on purpose: the standard
`/api/public/ingestion` path that ships review traces here feeds only
the trace-upsert queue, and `evalService.createEvalJobs` only creates
jobs for `targetObject ∈ {TRACE, DATASET}`. Observation rules are
triggered exclusively from the OTel ingestion pipeline, which this
pilot does not use. A trace rule reads the trace's own input/output —
`langfuse_trace` already writes `_review_input`/`_review_output` onto
the trace body for exactly this reason.
Mapping is required at both the rule root (server validates it there)
and inside `evaluator` (the API echoes it back).
"""
return {
"name": name,
"enabled": True,
"target": "observation",
"target": "trace",
"sampling": sampling,
"filter": [
{"column": "traceName", "operator": "any of",
+9 -3
View File
@@ -11,10 +11,16 @@ import eval_judges as ej # noqa: E402
# --- rule_body ------------------------------------------------------------
def test_rule_body_targets_observations():
"""Trace-level rules wouldn't see observation input/output."""
def test_rule_body_targets_traces():
"""Trace target matches the path `/api/public/ingestion` triggers.
Observation rules only fire from the OTel ingestion pipeline; this
pilot uses standard ingestion, so its jobs only come from
`evalService.createEvalJobs` and that dispatcher handles
`targetObject ∈ {TRACE, DATASET}`.
"""
body = ej.rule_body("rule-x", "finding_actionability", 1.0)
assert body["target"] == "observation"
assert body["target"] == "trace"
assert body["enabled"] is True