Combine Scores from Other Evaluators
A code evaluator can read the scores that other evaluators wrote on the same span. Use this to make one score from an LLM judge and a code check, to gate an expensive result on a cheap one, or to write an overall quality score.
def evaluate(ctx):
# The judge's 0 to 1 score, or 0.0 when there is none.
helpful = ctx.score("Helpfulness", default=0.0)
# The value of the "pii_free" score of the "PII leak" evaluator.
pii_free = ctx.score("PII leak", "pii_free", default=True)
overall = Score(
name="support_quality",
value=float(helpful) if pii_free else 0.0,
higher_is_better=True,
comment="Helpfulness score" if pii_free else "PII in the reply",
)
return EvaluationResult(scores=[overall])
How to read a score
| Expression | Gives |
|---|---|
ctx.score("Helpfulness") | The value of the first score that the evaluator wrote, or None |
ctx.score("PII leak", "pii_free") | The value of the score called pii_free. When the evaluator wrote no score with that name, the value of its first score |
ctx.score("Helpfulness", default=0.0) | The value, or 0.0 when the evaluator wrote no score on this span |
ctx.scores.get("PII leak") | The list of scores that the evaluator wrote, or None |
ctx.scores["PII leak"] | The same list. Raises KeyError when there is none, so prefer .get |
The list finds a score by its name:
.get("pii_free") gives the Score,
.value("pii_free", default=None) gives its value,
and .names() gives the names in order.
The key is the evaluator name as you see it on the Evaluators tab. For an LLM-as-Judge evaluator, the list holds one score with the evaluator's name: its value is the judge's 0 to 1 score and its comment is the judge's reasoning.
Dependencies are automatic
You do not list the evaluators that your code reads. When you save the evaluator, or the template that it runs, Oodle finds each name in the code and records a dependency on that evaluator. The evaluator form shows them as Reads scores from, and the dependency graph shows them as a separate kind of edge.
Oodle finds a name only when:
- The code reads it as
ctx.scores["Name"],ctx.scores.get("Name")orctx.score("Name"). - The name is a string literal in single or double quotes. A name that the code builds at run time gets no dependency and reads nothing.
- The line is not a comment. A read in a commented-out line does not make a dependency.
Oodle refuses to save an evaluator when:
- A name in the code is not the name of an evaluator in your instance. The error names it.
- The code reads the evaluator's own name.
- The dependencies make a cycle, for example A reads B
and B reads A. The error shows the cycle, such as
A → B → A. Oodle checks cycles over both kinds of dependency: scores that code reads, and the gate dependencies that you choose.
While other evaluators read an evaluator's scores, you cannot rename or delete it. The error names the evaluators that read it. Change their code first.
When it runs
An evaluator that reads scores runs after every evaluator that it reads, in the same evaluation cycle. It runs only on the spans that every one of those evaluators scored. A span that one of them has not scored yet waits for a later cycle.
This is different from a gate dependency:
Reads scores (ctx.scores) | Gate dependency | |
|---|---|---|
| You set it | From the code, automatically | On the evaluator form, by hand |
| Runs on | The spans that every input evaluator scored | Only the spans that the parent evaluator flagged (scored above 0) |
| Gets the parent's scores | Yes, in ctx.scores | No |
Use a gate to save cost: an LLM judge that runs only on the spans that a cheap code check flagged. Use a score read to combine results.
The same spans
An evaluator that reads scores and the evaluators that it reads must have the same filter rules, and every evaluator that it reads must be enabled. Otherwise the reader never gets their scores, so a save that breaks this is refused.
To change the filter rules of such a group, disable the evaluator that reads the scores first. Then change the filter rules of every evaluator in the group, and enable the reader again. The check runs when it is enabled.
In experiments
In an experiment, Oodle orders the evaluators of the run so that an evaluator that reads scores runs after the evaluators that it names, on each item. A run whose evaluators make a cycle fails with an error that names them.
Backfills
A backfill of an evaluator that reads other evaluators' scores also runs those evaluators, and the evaluators that they read in turn. In each time slice of the backfill, the evaluators run in dependency order, so each input runs before the evaluator that reads it. The results of the added evaluators are recorded like the results of the evaluators that you chose.
- The evaluator that reads scores runs only on the spans that every one of its inputs scored in that slice.
- Gate dependencies do not apply in a backfill: each evaluator runs on every span that matches its filters.
- The notes of the run list each added evaluator and
the reason, for example
also ran Helpfulness: Quality reads its scores. They also list each evaluator that did not run, because its evaluators form a cycle or because one of its inputs could not run.
An added evaluator costs what it costs in any run. An LLM-as-Judge input makes one judge call for each span that it scores. Before you start a backfill, the form shows the added evaluators as Also runs.
Test with simulated scores
A test run has no other evaluators, so give the scores
that the code reads. Send scores with the test run: a
map of evaluator name to a list of scores. With MCP
(genai_evaluator_workbench):
{"action": "try_code", "sourceCode": "def evaluate(ctx): ...",
"scores": {
"Helpfulness": [{"name": "Helpfulness", "value": 0.8, "data_type": "NUMERIC"}],
"PII leak": [{"name": "pii_free", "value": true, "data_type": "BOOLEAN"}]
}}
Without scores, ctx.scores is empty and
ctx.score(...) gives its default, so test that your
code handles a missing score. Always return at least
one score: an empty list fails with INVALID_RESULT.
See the cookbook for a complete example.
Support
If you need assistance or have any questions, please reach out to us through:
- Email at [email protected]