Code Evaluator Cookbook
Each recipe is a complete code evaluator. Paste it into
a new code template, click Run on a sample span, and
save it. The functions come from the
oodle_eval library.
Built-in check plus your own score
Run a built-in check and add a score of your own. Each library check returns a list of scores, so add your score to the list.
This is the same evaluator as an Opik custom metric that
runs ConversationDegenerationMetric and adds its own
score:
- Oodle
- Opik
from oodle_eval.v1 import metrics, text
def evaluate(ctx):
# degenerated (BOOLEAN) and degeneration (0 to 1)
scores = metrics.conversation_degeneration(ctx, threshold=0.6)
reply = text.reply(ctx)
mine = Score(
name="mentions_refund",
value="refund" in reply.lower(),
data_type="BOOLEAN",
higher_is_better=True,
)
return EvaluationResult(scores=scores + [mine])
from opik.evaluation.metrics import ConversationDegenerationMetric
from opik.evaluation.metrics import base_metric, score_result
class SupportReply(base_metric.BaseMetric):
def __init__(self, name="support_reply"):
super().__init__(name=name)
self.degeneration = ConversationDegenerationMetric()
def score(self, conversation, **ignored_kwargs):
degeneration = self.degeneration.score(conversation=conversation)
reply = conversation[-1]["content"]
mine = score_result.ScoreResult(
name="mentions_refund",
value=float("refund" in reply.lower()),
)
return [degeneration, mine]
In Oodle, the evaluator runs on each chat span as it
arrives. ctx holds the conversation: text.turns(ctx)
gives the turns as {"role", "content"} dicts, and
text.reply(ctx) gives the newest reply.
Weighted quality score
Run several checks and make one 0 to 1 score from them. Keep the separate scores too, so a drop in the combined score shows its cause.
from oodle_eval.v1 import combine, metrics
def evaluate(ctx):
readable = metrics.readability(ctx, max_grade=10)
tone = metrics.tone(ctx)
length = metrics.length_budget(ctx, max_words=250)
# The primary score of each check: all BOOLEAN, so all 0 or 1.
primary = [readable["readability_ok"], tone["tone_ok"], length["within_length"]]
quality = combine.weighted_mean(
primary,
weights={"readability_ok": 1, "tone_ok": 2, "within_length": 1},
name="quality",
)
return EvaluationResult(scores=readable + tone + length + [quality])
weighted_mean reads a BOOLEAN score as 0 or 1, and
skips categorical scores and scores with no value. It
uses each value as it is, so give it scores on the same
scale that read the same way. Here that excludes
reading_ease, which is 0 to 100. A name that
weights does not list has the weight 1.
Gate on a check
Score the content of a JSON reply only when the reply is valid JSON. When it is not, return the JSON check alone: a content score for a reply that did not parse means nothing.
from oodle_eval.v1 import combine, metrics, text
def evaluate(ctx):
shape = metrics.json_validity(ctx, required_keys=["answer", "sources"])
if not combine.passed(shape["is_json"]):
return EvaluationResult(scores=shape)
data = text.extract_json(text.reply(ctx))
cited = Score(
name="has_sources",
value=isinstance(data, dict) and bool(data.get("sources")),
data_type="BOOLEAN",
higher_is_better=True,
)
return EvaluationResult(scores=shape + [cited])
To gate a different evaluator on this one, for example an LLM judge that must run only on valid replies, use an evaluator dependency.
All checks pass
Return one pass or fail score for a set of rules, plus the separate scores:
from oodle_eval.v1 import combine, metrics
def evaluate(ctx):
checks = (
metrics.pii_leak(ctx)
+ metrics.refusal(ctx)
+ metrics.keyword_check(ctx, forbidden=["guarantee", "promise"])
)
return EvaluationResult(scores=checks + [combine.all_pass(checks, name="safe_reply")])
all_pass reads each score with combine.passed, in
the direction that the score declares. The refused
score is lower-is-better, so a reply that refuses fails
the set. A numeric score passes at 0.5 or more (below
0.5 for a lower-is-better score), unless you give at=.
Tool trajectory
Check that an agent called the tools that you expect.
In an experiment, ctx.trace.spans holds the agent's
steps. On a live span, the check reads the tool calls in
the span's output.
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.tool_trajectory(
ctx,
expected=["search_orders", "issue_refund"],
mode="superset",
))
mode | Passes when |
|---|---|
strict | The same calls in the same order |
unordered | The same calls in any order |
subset | Every call is an expected one: no extra tool |
superset | Every expected call happened. Extra calls are allowed |
Make expected a setting
to use one template for many agents:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.tool_trajectory(ctx, **ctx.params))
Use a shared library
Put your team's rules in a shared library and import them from each evaluator. This library holds a banned-phrase list and a check that uses it:
# shared library "acme_policy"
from oodle_eval.v1 import metrics
BANNED = ["guarantee", "risk-free", "cannot lose"]
def policy_check(ctx):
return metrics.keyword_check(ctx, forbidden=BANNED, name="policy_ok")
# code evaluator
from shared.acme_policy import policy_check
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=policy_check(ctx) + metrics.pii_leak(ctx))
When you change BANNED, every evaluator that imports
acme_policy uses the new list, unless its template
pins an earlier version.
Combine an LLM judge with a code check
Read the score of an LLM-as-Judge evaluator and a code check that run on the same span, and write one overall score. The names are evaluator names from the Evaluators tab. Oodle runs this evaluator after both of them. See Combine Scores from Other Evaluators.
from oodle_eval.v1 import combine
def evaluate(ctx):
relevance = ctx.score("Answer relevance", default=0.0)
pii_free = ctx.score("PII leak", "pii_free", default=True)
overall = combine.weighted_mean(
[
Score(name="relevance", value=relevance),
Score(name="pii_free", value=pii_free, data_type="BOOLEAN"),
],
weights={"relevance": 3, "pii_free": 1},
name="overall",
)
return EvaluationResult(scores=[overall])
Compare with a reference value
In an experiment, the checks that compare with a
reference use the dataset item's expected output
(ctx.expected_output) when you do not pass
reference. On a live span there is no expected
output: give the reference as a setting, or read it
from a span attribute.
from oodle_eval.v1 import metrics, text
def evaluate(ctx):
# The expected output in an experiment, else a span attribute.
expected = ctx.expected_output
if expected is None:
expected = text.metadata_value(ctx, "app.expected_answer", default="")
return EvaluationResult(scores=(
metrics.exact_match(ctx, reference=expected, normalize=True)
+ metrics.levenshtein_ratio(ctx, reference=expected)
))
A RAG answer is grounded in its context
Check that a retrieval-augmented answer stays inside what was retrieved: its sentences use the context's words, its numbers and names are in the context, and its citations point at real sources. Then make one pass or fail score.
from oodle_eval.v1 import combine, metrics
def evaluate(ctx):
grounded = metrics.grounding(ctx, threshold=0.8)
numbers = metrics.unsupported_numbers(ctx)
citations = metrics.citation_check(ctx)
checks = grounded + numbers + citations
primary = [
grounded["grounded"],
numbers["numbers_supported"],
citations["citations_ok"],
]
return EvaluationResult(scores=checks + [combine.all_pass(primary, name="rag_ok")])
The three checks find the retrieved context in this order:
- The
contextargument, when you pass one. - The span attribute that
context_keynames (contextby default). Passcontext_key=when your application writes the context to another attribute. - The text of the input messages: a RAG call usually puts what it retrieved in its prompt.
Support
If you need assistance or have any questions, please reach out to us through:
- Email at [email protected]