Skip to main content

Code Evaluator Cookbook

Each recipe is a complete code evaluator. Paste it into a new code template, click Run on a sample span, and save it. The functions come from the oodle_eval library.

Built-in check plus your own score​

Run a built-in check and add a score of your own. Each library check returns a list of scores, so add your score to the list.

This is the same evaluator as an Opik custom metric that runs ConversationDegenerationMetric and adds its own score:

from oodle_eval.v1 import metrics, text


def evaluate(ctx):
# degenerated (BOOLEAN) and degeneration (0 to 1)
scores = metrics.conversation_degeneration(ctx, threshold=0.6)

reply = text.reply(ctx)
mine = Score(
name="mentions_refund",
value="refund" in reply.lower(),
data_type="BOOLEAN",
higher_is_better=True,
)
return EvaluationResult(scores=scores + [mine])

In Oodle, the evaluator runs on each chat span as it arrives. ctx holds the conversation: text.turns(ctx) gives the turns as {"role", "content"} dicts, and text.reply(ctx) gives the newest reply.

Weighted quality score​

Run several checks and make one 0 to 1 score from them. Keep the separate scores too, so a drop in the combined score shows its cause.

from oodle_eval.v1 import combine, metrics


def evaluate(ctx):
readable = metrics.readability(ctx, max_grade=10)
tone = metrics.tone(ctx)
length = metrics.length_budget(ctx, max_words=250)

# The primary score of each check: all BOOLEAN, so all 0 or 1.
primary = [readable["readability_ok"], tone["tone_ok"], length["within_length"]]
quality = combine.weighted_mean(
primary,
weights={"readability_ok": 1, "tone_ok": 2, "within_length": 1},
name="quality",
)
return EvaluationResult(scores=readable + tone + length + [quality])

weighted_mean reads a BOOLEAN score as 0 or 1, and skips categorical scores and scores with no value. It uses each value as it is, so give it scores on the same scale that read the same way. Here that excludes reading_ease, which is 0 to 100. A name that weights does not list has the weight 1.

Gate on a check​

Score the content of a JSON reply only when the reply is valid JSON. When it is not, return the JSON check alone: a content score for a reply that did not parse means nothing.

from oodle_eval.v1 import combine, metrics, text


def evaluate(ctx):
shape = metrics.json_validity(ctx, required_keys=["answer", "sources"])
if not combine.passed(shape["is_json"]):
return EvaluationResult(scores=shape)

data = text.extract_json(text.reply(ctx))
cited = Score(
name="has_sources",
value=isinstance(data, dict) and bool(data.get("sources")),
data_type="BOOLEAN",
higher_is_better=True,
)
return EvaluationResult(scores=shape + [cited])

To gate a different evaluator on this one, for example an LLM judge that must run only on valid replies, use an evaluator dependency.

All checks pass​

Return one pass or fail score for a set of rules, plus the separate scores:

from oodle_eval.v1 import combine, metrics


def evaluate(ctx):
checks = (
metrics.pii_leak(ctx)
+ metrics.refusal(ctx)
+ metrics.keyword_check(ctx, forbidden=["guarantee", "promise"])
)
return EvaluationResult(scores=checks + [combine.all_pass(checks, name="safe_reply")])

all_pass reads each score with combine.passed, in the direction that the score declares. The refused score is lower-is-better, so a reply that refuses fails the set. A numeric score passes at 0.5 or more (below 0.5 for a lower-is-better score), unless you give at=.

Tool trajectory​

Check that an agent called the tools that you expect. In an experiment, ctx.trace.spans holds the agent's steps. On a live span, the check reads the tool calls in the span's output.

from oodle_eval.v1 import metrics


def evaluate(ctx):
return EvaluationResult(scores=metrics.tool_trajectory(
ctx,
expected=["search_orders", "issue_refund"],
mode="superset",
))
modePasses when
strictThe same calls in the same order
unorderedThe same calls in any order
subsetEvery call is an expected one: no extra tool
supersetEvery expected call happened. Extra calls are allowed

Make expected a setting to use one template for many agents:

from oodle_eval.v1 import metrics


def evaluate(ctx):
return EvaluationResult(scores=metrics.tool_trajectory(ctx, **ctx.params))

Use a shared library​

Put your team's rules in a shared library and import them from each evaluator. This library holds a banned-phrase list and a check that uses it:

# shared library "acme_policy"
from oodle_eval.v1 import metrics

BANNED = ["guarantee", "risk-free", "cannot lose"]


def policy_check(ctx):
return metrics.keyword_check(ctx, forbidden=BANNED, name="policy_ok")
# code evaluator
from shared.acme_policy import policy_check
from oodle_eval.v1 import metrics


def evaluate(ctx):
return EvaluationResult(scores=policy_check(ctx) + metrics.pii_leak(ctx))

When you change BANNED, every evaluator that imports acme_policy uses the new list, unless its template pins an earlier version.

Combine an LLM judge with a code check​

Read the score of an LLM-as-Judge evaluator and a code check that run on the same span, and write one overall score. The names are evaluator names from the Evaluators tab. Oodle runs this evaluator after both of them. See Combine Scores from Other Evaluators.

from oodle_eval.v1 import combine


def evaluate(ctx):
relevance = ctx.score("Answer relevance", default=0.0)
pii_free = ctx.score("PII leak", "pii_free", default=True)

overall = combine.weighted_mean(
[
Score(name="relevance", value=relevance),
Score(name="pii_free", value=pii_free, data_type="BOOLEAN"),
],
weights={"relevance": 3, "pii_free": 1},
name="overall",
)
return EvaluationResult(scores=[overall])

Compare with a reference value​

In an experiment, the checks that compare with a reference use the dataset item's expected output (ctx.expected_output) when you do not pass reference. On a live span there is no expected output: give the reference as a setting, or read it from a span attribute.

from oodle_eval.v1 import metrics, text


def evaluate(ctx):
# The expected output in an experiment, else a span attribute.
expected = ctx.expected_output
if expected is None:
expected = text.metadata_value(ctx, "app.expected_answer", default="")
return EvaluationResult(scores=(
metrics.exact_match(ctx, reference=expected, normalize=True)
+ metrics.levenshtein_ratio(ctx, reference=expected)
))

A RAG answer is grounded in its context​

Check that a retrieval-augmented answer stays inside what was retrieved: its sentences use the context's words, its numbers and names are in the context, and its citations point at real sources. Then make one pass or fail score.

from oodle_eval.v1 import combine, metrics


def evaluate(ctx):
grounded = metrics.grounding(ctx, threshold=0.8)
numbers = metrics.unsupported_numbers(ctx)
citations = metrics.citation_check(ctx)

checks = grounded + numbers + citations
primary = [
grounded["grounded"],
numbers["numbers_supported"],
citations["citations_ok"],
]
return EvaluationResult(scores=checks + [combine.all_pass(primary, name="rag_ok")])

The three checks find the retrieved context in this order:

  1. The context argument, when you pass one.
  2. The span attribute that context_key names (context by default). Pass context_key= when your application writes the context to another attribute.
  3. The text of the input messages: a RAG call usually puts what it retrieved in its prompt.

Support

If you need assistance or have any questions, please reach out to us through: