Skip to main content

oodle_eval Library Reference

Code evaluators can import the oodle_eval library. It holds the built-in checks as functions, helpers that read text and tool calls out of a span, and functions that combine scores. See Code Evaluators for how an evaluator runs, and the Cookbook for complete examples.

Version: 1.0.0

from oodle_eval.v1 import metrics, text, combine

Everything under oodle_eval.v1 keeps its behaviour: a later release changes it only to fix a bug. A change that would move your scores goes into a new version module, so an evaluator that imports v1 gives the same scores after an upgrade.

Each function in metrics takes the evaluation context ctx (or a plain string) as its first argument, takes its settings as keyword arguments, and returns Scores: a list of Score objects with the primary score first. Add the lists together to return scores from several checks.

The evaluation context​

PathValue
ctx.observation.inputThe span's input: a string, a message list (OpenAI, Anthropic or OpenTelemetry gen_ai shape) or any JSON value.
ctx.observation.outputThe span's output, in the same shapes as the input. A structured output or tool result arrives parsed.
ctx.observation.metadataThe span's attributes as a dict. A numeric attribute can carry a type suffix, such as duration_ms_int.
ctx.trace.trace_idThe trace id. Empty for a live span.
ctx.trace.spansThe agent's steps on an experiment item: dicts with name, kind (llm, tool, agent or span), tool_name, model, input, output, duration_ms and error. Empty for a live span.
ctx.trace.tools_called()The tool name of each tool step in ctx.trace.spans, in order.
ctx.paramsThe evaluator's settings as a dict: the value set on the evaluator, else the template's default, else None. Every setting the template declares is a key.
ctx.experimentThe dataset item on an experiment run: item_id, expected_output, item_metadata, run_id and dataset_id. Every field is None (metadata {}) on a live span.
ctx.expected_outputShort for ctx.experiment.expected_output: the item's expected output, or None on a live span. Checks that compare against a reference use it when their reference setting is empty.
ctx.scoresThe scores other evaluators wrote on the same span, as {evaluator name: Scores}. The evaluator runs after the ones its code names here. Use ctx.scores.get(name): an evaluator with no score on the span is not a key.
ctx.score(evaluator, score_name=None, default=None)The value of one score another evaluator wrote: the score called score_name, else its first score, else default.

Runtime types​

Score and EvaluationResult are already in the globals of your evaluator code. Import the others, and import all of them in a shared library, from oodle_eval.runtime.

Score​

Score(value, data_type='NUMERIC', comment=None, name=None, higher_is_better=None)
One named score.

Args:
value: A number for NUMERIC, True or False for BOOLEAN,
or a string for CATEGORICAL.
data_type: "NUMERIC", "BOOLEAN" or "CATEGORICAL".
comment: A short reason, shown next to the score.
name: The score name. An evaluator that returns more
than one score must name each.
higher_is_better: Which way is good for this score. None
uses the evaluator's own declaration.

EvaluationResult​

EvaluationResult(scores)
What `evaluate(ctx)` returns.

Args:
scores: A list of `Score` (a `Scores` is a list).

Scores​

Scores(iterable=())
A list of `Score` that also finds a score by its name.

Every built-in check returns one. It is a list, so
`scores + [Score(...)]` and `EvaluationResult(scores=scores)`
work as they do for a list. `scores["name"]` gives the score
of that name (KeyError when there is none); an int index
still gives the score at that place.

EvaluationContext​

EvaluationContext(observation, trace=None, params=None, scores=None, experiment=None)
What `evaluate(ctx)` gets.

Args:
observation: The `ObservationContext` under evaluation.
trace: The `TraceContext`. Empty for a live span.
params: The evaluator's settings, resolved: the value set
on the evaluator, else the template's default, else
None. Every setting the template declares is a key.
An empty dict when the template declares none.
scores: The scores other evaluators wrote on the same
span, as {evaluator name: Scores}. Only evaluators
the code names in `ctx.scores[...]`,
`ctx.scores.get(...)` or `ctx.score(...)` are here.
An empty dict when there are none.
experiment: The `ExperimentContext` of an experiment item.
Empty (every field None, metadata {}) on a live span.

ObservationContext​

ObservationContext(input=None, output=None, metadata=None)
The span or dataset item under evaluation.

Args:
input: The input, as the span stored it: a string, a
message list or any JSON value.
output: The output, in the same shapes as `input`.
metadata: The span's attributes, as a dict.

TraceContext​

TraceContext(trace_id='', spans=None)
The agent's trace, as ordered steps.

Args:
trace_id: The trace id. Empty for a live span.
spans: A list of dicts with name, kind (llm, tool, agent
or span), tool_name, model, input, output,
duration_ms and error. Filled for an experiment
item; empty for a live span.

ExperimentContext​

ExperimentContext(item_id=None, expected_output=None, item_metadata=None, run_id=None, dataset_id=None)
The dataset item an experiment runs, empty on a live span.

Args:
item_id: The dataset item's id. None on a live span.
expected_output: The item's expected output, as the
dataset stores it: a string or any JSON value. None
when the item has none.
item_metadata: The item's metadata dict. {} when none.
run_id: The experiment run's id.
dataset_id: The dataset's id.

metrics​

Import path: oodle_eval.v1.metrics

Built-in checks. Each returns `Scores`, primary score first.

Every check takes the evaluation context `ctx` first, or a plain
string to check that text. Its settings are keyword arguments,
so a template can pass its settings straight through:

def evaluate(ctx):
return EvaluationResult(scores=metrics.keyword_check(ctx, **ctx.params))

A setting passed as None takes its default, so an evaluator that
leaves a setting empty gets the check's own behaviour. Where the
summary says so, None on a limit means "do not check this limit".

`name` renames the primary score; the other scores keep their
names. Use `combine.rename` to change them all.

Checks that compare against a reference take it as the
`reference` setting. When that is None, they use the experiment
item's expected output (`ctx.expected_output`), so the same check
works on live spans with a fixed reference and in experiments
with the dataset's ground truth.

The checks live in one module per area, and every one is also
here, so `metrics.keyword_check` and
`from oodle_eval.v1.metrics import keyword_check` both work:

- `metrics.format`: JSON, patterns, keywords, length, Markdown,
HTML, CSV, TOML, SQL
- `metrics.quality`: readability, language, tone, sentiment
- `metrics.similarity`: comparison with a reference, a question
or a context
- `metrics.safety`: PII and secrets, refusals, prompt injection,
URLs
- `metrics.conversation`: repetition and forgotten facts
- `metrics.context`: grounding, unsupported numbers, citations
- `metrics.agent`: tool calls, tool trajectory, token and
latency budget
- `metrics.trace`: steps, failures, loops and duration of the
agent's trace

To write a check of your own the same way, see
`oodle_eval.v1.util`.
FunctionSummary
json_validityCheck that the output is JSON and has the shape you expect.
regex_matchCheck that the text matches every required pattern and no forbidden one.
keyword_checkCheck that the text mentions the required words and none of the forbidden ones.
containsCheck that the text contains the given substrings.
length_budgetCheck that the reply stays inside a word and character range.
markdown_structureCheck the reply's Markdown: code fences, tables, headings and lists.
html_validityCheck that the HTML in the reply has balanced tags.
csv_validityCheck that the reply is CSV with the same number of fields in every row.
toml_validityCheck that the reply parses as TOML and has the keys you require.
sql_shapeCheck that generated SQL is only the kind of statement you allow.
readabilityScore Flesch reading ease and check the Flesch-Kincaid grade.
detect_languageReturn the language of a text and a confidence.
language_adherenceCheck that the reply is written in the language you expect.
toneFlag a reply that shouts, piles on exclamation marks, reads as negative, or uses a banned phrase.
sentimentScore the reply's sentiment from -1 (negative) to 1 (positive).
exact_matchCheck that the reply equals a reference text.
levenshtein_ratioScore how close the reply is to a reference, by edit distance.
jaccard_similarityScore the overlap of the reply's and the reference's word sets.
rougeScore ROUGE-L, ROUGE-1 and ROUGE-2 F1 against a reference.
bleuScore sentence BLEU against a reference.
numeric_diffScore how close the first number in the reply is to a reference number.
json_diffScore how close the output's JSON is to a reference JSON value.
bm25_relevanceScore how relevant the reply is to the question or context, by BM25.
tfidf_similarityScore the cosine similarity of the reply and the question or context, by TF-IDF.
pii_leakCheck the reply for an email, phone, card number, national id, IP address or credential.
refusalCheck whether the model declined to do what it was asked.
url_checkCheck the URLs in the text: well formed, allowed scheme and domain, not private.
prompt_injectionFlag text that tries to override the model's instructions.
conversation_degenerationCheck whether the assistant starts to repeat itself, in the reply or across turns.
knowledge_retentionCheck that the reply does not forget what the user already said.
groundingScore the share of the reply's sentences whose words appear in the context.
unsupported_numbersFind numbers and names in the reply that the context does not have.
citation_checkCheck that every citation in the reply points at a real source.
tool_call_validityCheck that the tool calls are well formed, allowed and not stuck in a loop.
tools_usedReturn the names of the tools the agent called, in order.
tool_trajectoryCheck the agent's tool calls against the tools you expect.
token_latency_budgetCheck that the call stays inside its token, duration and time to first token limits.
span_countCount the agent's steps in ctx.trace.spans.
error_spansCount the agent's steps that failed, in ctx.trace.spans.
repeated_tool_callsDetect the agent calling the same tool with the same input again and again.
total_durationMeasure how long the agent ran, from its trace.

json_validity​

json_validity(ctx, *, extract_from_code_fence=True, required_keys=(), key_types=None, schema=None, name='is_json')

Score names: is_json, json_schema_ok

Check that the output is JSON and has the shape you expect.

A structured output or tool result that arrives already
parsed counts as JSON with no second parse.

Args:
ctx: The evaluation context (reads the output), a string,
or a parsed JSON value.
extract_from_code_fence: Read the JSON inside a ```json
fenced block when the reply wraps it in prose. False
demands bare JSON.
required_keys: Keys the top-level object must carry.
key_types: {key: type} for top-level keys, where type is
"string", "number", "integer", "boolean", "array",
"object" or "null". A missing key is reported by
`required_keys`, not here.
schema: A JSON Schema subset the value must match: type,
required, properties, items, enum, minimum, maximum,
minLength and maxLength. Other keywords are ignored.
name: Name of the primary score.

Returns:
is_json (BOOLEAN): the output parses as JSON.
json_schema_ok (BOOLEAN): it parses and matches
`required_keys`, `key_types` and `schema`.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.json_validity(ctx))

regex_match​

regex_match(ctx, *, must_match=(), must_not_match=(), field='output', name='regex_pass')

Score names: regex_pass

Check that the text matches every required pattern and no forbidden one.

Patterns use Python's re syntax, in multi-line mode. Put
(?i) at the start of a pattern to ignore case. A pattern that
does not compile fails the check and says so.

Args:
ctx: The evaluation context, or the text to check.
must_match: Patterns that must each match.
must_not_match: Patterns that must not match.
field: "output" (the assistant's reply), "input" (the
user's messages) or "both".
name: Name of the primary score.

Returns:
regex_pass (BOOLEAN): every rule held.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.regex_match(ctx))

keyword_check​

keyword_check(ctx, *, required=(), forbidden=(), mode='all', case_sensitive=False, field='output', name='keywords_ok')

Score names: keywords_ok, keyword_coverage

Check that the text mentions the required words and none of the forbidden ones.

Matching is on whole words, so "art" does not match "start".
Use `contains` to match any part of a word.

Args:
ctx: The evaluation context, or the text to check.
required: Keywords or phrases the text must mention.
forbidden: Keywords or phrases the text must not mention.
mode: "all": every required keyword must appear. "any":
one is enough.
case_sensitive: Match case exactly.
field: "output" (the assistant's reply), "input" (the
user's messages) or "both".
name: Name of the primary score.

Returns:
keywords_ok (BOOLEAN): the required rule held and no
forbidden keyword appeared.
keyword_coverage (NUMERIC, 0 to 1): the share of required
keywords found. Only when `required` is set.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.keyword_check(ctx))

contains​

contains(ctx, *, values=(), mode='any', case_sensitive=False, field='output', name='contains')

Score names: contains

Check that the text contains the given substrings.

Unlike `keyword_check`, a value matches anywhere, also inside
a word. Like OpenAI's `like` / `ilike` string check and Opik's
Contains.

Args:
ctx: The evaluation context, or the text to check.
values: The substrings to look for. Empty: the experiment
item's expected output (`ctx.expected_output`), a text
or a list of texts.
mode: "any": one value is enough. "all": every value
must appear. "none": no value may appear.
case_sensitive: Match case exactly.
field: "output" (the assistant's reply), "input" (the
user's messages) or "both".
name: Name of the primary score.

Returns:
contains (BOOLEAN): the text meets `mode`.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.contains(ctx))

length_budget​

length_budget(ctx, *, min_words=1, max_words=300, min_chars=None, max_chars=4000, name='within_length')

Score names: within_length

Check that the reply stays inside a word and character range.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
min_words: Fewest words. None: not checked.
max_words: Most words. None: not checked.
min_chars: Fewest characters. None: not checked.
max_chars: Most characters. None: not checked.
name: Name of the primary score.

Returns:
within_length (BOOLEAN): every bound held.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.length_budget(ctx))

markdown_structure​

markdown_structure(ctx, *, required_headings=(), min_headings=0, require_list=False, require_table=False, name='markdown_ok')

Score names: markdown_ok, heading_count

Check the reply's Markdown: code fences, tables, headings and lists.

Code fences (``` or ~~~) must be closed. A table's rows must
have as many cells as its header. Text inside a code fence is
not read as Markdown.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
required_headings: Headings the reply must have, matched
in any case, with no punctuation, inside the heading
text.
min_headings: Fewest headings.
require_list: The reply must have a bullet or numbered
list.
require_table: The reply must have a table.
name: Name of the primary score.

Returns:
markdown_ok (BOOLEAN): every rule held.
heading_count (NUMERIC): how many headings the reply has.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.markdown_structure(ctx))

html_validity​

html_validity(ctx, *, allowed_tags=(), name='html_valid')

Score names: html_valid

Check that the HTML in the reply has balanced tags.

Reads a ```html fenced block when there is one, else the
whole reply. Void tags (br, img, ...) need no end tag, and
tags whose end tag HTML makes optional (p, li, td, ...) are
not reported when left open.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
allowed_tags: Tag names allowed. Empty: any tag.
name: Name of the primary score.

Returns:
html_valid (BOOLEAN): every tag is closed in order, and
allowed.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.html_validity(ctx))

csv_validity​

csv_validity(ctx, *, delimiter=None, has_header=True, required_columns=(), min_rows=1, name='csv_valid')

Score names: csv_valid, row_count

Check that the reply is CSV with the same number of fields in every row.

Reads a ```csv fenced block when there is one, else the whole
reply. Blank lines are skipped.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
delimiter: The field separator. Empty: the one of , ; tab
and | that the first row uses most.
has_header: The first row names the columns.
required_columns: Column names the header must have.
Needs `has_header`.
min_rows: Fewest data rows, the header not counted.
name: Name of the primary score.

Returns:
csv_valid (BOOLEAN): the rules held.
row_count (NUMERIC): data rows, the header not counted.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.csv_validity(ctx))

toml_validity​

toml_validity(ctx, *, required_keys=(), name='toml_valid')

Score names: toml_valid

Check that the reply parses as TOML and has the keys you require.

Reads a ```toml fenced block when there is one, else the
whole reply.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
required_keys: Keys the document must have. A dotted key
such as "server.port" looks inside tables.
name: Name of the primary score.

Returns:
toml_valid (BOOLEAN): it parses and has every required key.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.toml_validity(ctx))

sql_shape​

sql_shape(ctx, *, allowed_statements=('SELECT',), allowed_keywords=(), max_statements=1, name='sql_ok')

Score names: sql_ok, statement_type

Check that generated SQL is only the kind of statement you allow.

Reads a ```sql fenced block when there is one, else the whole
reply. Comments and string literals are left out before the
check, so a keyword inside a string does not count. This is a
pattern check, not a SQL parser: use it to catch a model that
writes where it should only read.

Args:
ctx: The evaluation context (reads the reply), or the
SQL text.
allowed_statements: Statement types allowed, such as
"SELECT" or "INSERT". A WITH statement counts as the
write it does, else as SELECT.
allowed_keywords: Write keywords (DROP, DELETE, INSERT,
UPDATE, ALTER, CREATE, TRUNCATE, GRANT, ...) allowed
anywhere. An allowed statement type is also allowed
as a keyword.
max_statements: Most statements, split on ";".
name: Name of the primary score.

Returns:
sql_ok (BOOLEAN): every rule held.
statement_type (CATEGORICAL): the first statement's type,
or "none".

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.sql_shape(ctx))

readability​

readability(ctx, *, min_grade=None, max_grade=10, min_words=20, name='readability_ok')

Score names: readability_ok, reading_ease

Score Flesch reading ease and check the Flesch-Kincaid grade.

Syllables are counted with a vowel-group rule, so the grade
is an estimate, close enough to tell a plain answer from a
dense one. Code blocks, URLs and markdown are left out.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
min_grade: Lowest grade allowed. None: not checked.
max_grade: Highest grade allowed. None: not checked.
min_words: Replies shorter than this are too short to
grade and pass.
name: Name of the primary score.

Returns:
readability_ok (BOOLEAN): the grade is in range.
reading_ease (NUMERIC, 0 to 1): Flesch reading ease over
100. Not returned for a reply too short to grade.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.readability(ctx))

detect_language​

detect_language(value)
Return the language of a text and a confidence.

Latin-script languages are told apart by their most common
words; other scripts by their characters. Mixed or very short
text can read as "unknown".

Args:
value: The text.

Returns:
(code, confidence): an ISO 639-1 code (en, es, fr, de,
it, pt, nl, ru, uk, el, ar, he, hi, th, zh, ja, ko) or
"unknown", and a confidence from 0 to 1.

language_adherence​

language_adherence(ctx, *, expected_language='en', min_words=5, name='language_match')

Score names: language_match, detected_language

Check that the reply is written in the language you expect.

No model is needed, so it runs in the sandbox. See
`detect_language` for how the language is found.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
expected_language: ISO 639-1 code: en, es, fr, de, it,
pt, nl, ru, uk, el, ar, he, hi, th, zh, ja or ko.
min_words: Replies shorter than this pass with no verdict.
name: Name of the primary score.

Returns:
language_match (BOOLEAN): the detected language is the
expected one.
detected_language (CATEGORICAL): the detected code, or
"unknown". Not returned for a reply too short to
detect.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.language_adherence(ctx))

tone​

tone(ctx, *, max_uppercase_ratio=0.3, max_exclamations=2, min_sentiment=-0.3, forbidden_phrases=('calm down', 'obviously', 'as i already said', "that's not my problem", 'you should have', 'stupid', 'shut up'), extra_positive_words=(), extra_negative_words=(), name='tone_ok')

Score names: tone_ok, sentiment

Flag a reply that shouts, piles on exclamation marks, reads as negative, or uses a banned phrase.

Sentiment is a small word list with negation, scored the way
VADER normalises its compound score, so it needs no model.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
max_uppercase_ratio: Share of capital letters, over words
of 3 or more letters, above which the reply shouts.
Checked only on 20 or more such letters.
max_exclamations: Most exclamation marks allowed.
min_sentiment: Sentiment runs from -1 to 1. Below this
is negative.
forbidden_phrases: Phrases that fail the check, matched
in lower case anywhere in the reply.
extra_positive_words: Words to add to the positive list.
extra_negative_words: Words to add to the negative list.
name: Name of the primary score.

Returns:
tone_ok (BOOLEAN): no issue found.
sentiment (NUMERIC, -1 to 1): the reply's sentiment.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.tone(ctx))

sentiment​

sentiment(ctx, *, extra_positive_words=(), extra_negative_words=(), name='sentiment')

Score names: sentiment, sentiment_label

Score the reply's sentiment from -1 (negative) to 1 (positive).

The same word-list method as `tone`: no model is needed.
Like Opik Sentiment.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
extra_positive_words: Words to add to the positive list.
extra_negative_words: Words to add to the negative list.
name: Name of the primary score.

Returns:
sentiment (NUMERIC, -1 to 1): the compound sentiment.
sentiment_label (CATEGORICAL): "positive" at 0.05 and
above, "negative" at -0.05 and below, else "neutral".

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.sentiment(ctx))

exact_match​

exact_match(ctx, *, reference=None, normalize=False, case_sensitive=True, name='exact_match')

Score names: exact_match

Check that the reply equals a reference text.

Space at the start and end is ignored. Like autoevals
ExactMatch, Opik Equals and OpenAI's `eq` string check.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
reference: The expected text. None: the
experiment item's expected output (`ctx.expected_output`).
normalize: Compare after `text.normalize`: lower case,
no punctuation, single spaces.
case_sensitive: Match case exactly. Ignored when
`normalize` is set.
name: Name of the primary score.

Returns:
exact_match (BOOLEAN): the texts are equal.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.exact_match(ctx, reference=ctx.params.get("reference")))

levenshtein_ratio​

levenshtein_ratio(ctx, *, reference=None, normalize=False, name='levenshtein_ratio')

Score names: levenshtein_ratio

Score how close the reply is to a reference, by edit distance.

1 minus the Levenshtein distance over the longer length.
Each text is compared on its first 1000 characters. Like
autoevals Levenshtein and Opik LevenshteinRatio.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
reference: The expected text. None: the
experiment item's expected output (`ctx.expected_output`).
normalize: Compare after `text.normalize`.
name: Name of the primary score.

Returns:
levenshtein_ratio (NUMERIC, 0 to 1): 1 is identical.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.levenshtein_ratio(ctx, reference=ctx.params.get("reference")))

jaccard_similarity​

jaccard_similarity(ctx, *, reference=None, name='jaccard_similarity')

Score names: jaccard_similarity

Score the overlap of the reply's and the reference's word sets.

Words are compared in lower case with no punctuation.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
reference: The expected text. None: the
experiment item's expected output (`ctx.expected_output`).
name: Name of the primary score.

Returns:
jaccard_similarity (NUMERIC, 0 to 1): shared words over
all words. 1 when both texts have no words.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.jaccard_similarity(ctx, reference=ctx.params.get("reference")))

rouge​

rouge(ctx, *, reference=None, name='rouge_l')

Score names: rouge_l, rouge_1, rouge_2

Score ROUGE-L, ROUGE-1 and ROUGE-2 F1 against a reference.

Tokens are normalised words (see `text.normalize`). ROUGE-L
uses the first 400 tokens of each text. Like OpenAI's
text_similarity grader and Opik ROUGE.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
reference: The expected text. None: the
experiment item's expected output (`ctx.expected_output`).
name: Name of the primary score.

Returns:
rouge_l (NUMERIC, 0 to 1): longest common subsequence F1.
rouge_1 (NUMERIC, 0 to 1): word overlap F1.
rouge_2 (NUMERIC, 0 to 1): word pair overlap F1.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.rouge(ctx, reference=ctx.params.get("reference")))

bleu​

bleu(ctx, *, reference=None, max_n=4, name='bleu')

Score names: bleu

Score sentence BLEU against a reference.

Geometric mean of the 1- to `max_n`-gram precisions, with
add-one smoothing for n above 1 so a short reply does not
score 0, times the brevity penalty. Tokens are normalised
words. Like Opik's and OpenAI's BLEU.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
reference: The expected text. None: the
experiment item's expected output (`ctx.expected_output`).
max_n: The longest n-gram, 1 to 4.
name: Name of the primary score.

Returns:
bleu (NUMERIC, 0 to 1): 1 is identical.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.bleu(ctx, reference=ctx.params.get("reference")))

numeric_diff​

numeric_diff(ctx, *, reference=None, name='numeric_diff')

Score names: numeric_diff

Score how close the first number in the reply is to a reference number.

1 minus the difference over the sum of the two sizes, as
autoevals NumericDiff does. "1,200" reads as 1200.

Args:
ctx: The evaluation context (reads the reply), the text,
or a number.
reference: The expected number, or a text holding it.
None: the experiment item's expected output
(`ctx.expected_output`).
name: Name of the primary score.

Returns:
numeric_diff (NUMERIC, 0 to 1): 1 is equal. 0 when the
reply holds no number.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.numeric_diff(ctx, reference=ctx.params.get("reference")))

json_diff​

json_diff(ctx, *, reference=None, name='json_diff')

Score names: json_diff

Score how close the output's JSON is to a reference JSON value.

Objects score the mean over the union of their keys (a key
on one side only scores 0), arrays the mean over the longer
length, strings by edit similarity, numbers as in
`numeric_diff`. Like autoevals JSONDiff.

Args:
ctx: The evaluation context (reads the output), a text,
or a parsed value.
reference: The expected value, or its JSON text. A text
that is not JSON is compared as a string. None: the
experiment item's expected output (`ctx.expected_output`).
name: Name of the primary score.

Returns:
json_diff (NUMERIC, 0 to 1): 1 is the same structure and
values. 0 when the output holds no JSON.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.json_diff(ctx, reference=ctx.params.get("reference")))

bm25_relevance​

bm25_relevance(ctx, *, against='input', context=None, context_key='context', k1=1.5, b=0.75, name='bm25_relevance')

Score names: bm25_relevance

Score how relevant the reply is to the question or context, by BM25.

The question's (or context's) content words are the query,
and each sentence of the reply is a document; the best
sentence's BM25 score is divided by the score a sentence of
average length that holds each query word once would get,
and capped at 1. Word frequencies come from the sentences of
both texts. Use it to rank or to flag replies that do not
touch the question; it is not a calibrated probability.

Args:
ctx: The evaluation context, or the reply text (then give
`context`).
against: "input": the user's last message. "context":
the context (see `grounding` for where it comes from).
context: The context text, or a list of passages.
context_key: The span attribute that holds the context
when `context` is not given.
k1: BM25 term saturation.
b: BM25 length normalisation, 0 to 1.
name: Name of the primary score.

Returns:
bm25_relevance (NUMERIC, 0 to 1): higher is more relevant.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.bm25_relevance(ctx))

tfidf_similarity​

tfidf_similarity(ctx, *, against='input', context=None, context_key='context', name='tfidf_similarity')

Score names: tfidf_similarity

Score the cosine similarity of the reply and the question or context, by TF-IDF.

Content words only (no stop words). Word weights come from
the sentences of both texts.

Args:
ctx: The evaluation context, or the reply text (then give
`context`).
against: "input": the user's last message. "context":
the context (see `grounding` for where it comes from).
context: The context text, or a list of passages.
context_key: The span attribute that holds the context
when `context` is not given.
name: Name of the primary score.

Returns:
tfidf_similarity (NUMERIC, 0 to 1): 1 is the same words
in the same proportions.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.tfidf_similarity(ctx))

pii_leak​

pii_leak(ctx, *, checks=('email', 'phone', 'credit_card', 'us_ssn', 'ip_address', 'secret'), allowed_email_domains=('example.com',), ignore_values_from_input=True, name='pii_free')

Score names: pii_free, pii_count

Check the reply for an email, phone, card number, national id, IP address or credential.

Card numbers are checked with the Luhn sum, so a long order
number does not read as a card. The comment names the kinds
found, never the values, so a score does not copy the leak.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
checks: The kinds to look for: "email", "phone",
"credit_card", "us_ssn", "ip_address", "secret".
allowed_email_domains: Email domains that are fine to
show, such as your support address.
ignore_values_from_input: A value the user typed is not a
leak when the model says it back. False flags every
occurrence.
name: Name of the primary score.

Returns:
pii_free (BOOLEAN): nothing found.
pii_count (NUMERIC): how many values were found.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.pii_leak(ctx))

refusal​

refusal(ctx, *, patterns=None, extra_patterns=(), scan_first_chars=400, name='refused')

Score names: refused

Check whether the model declined to do what it was asked.

A refusal is sometimes correct, so this scores the event.
Pair it with a filter on the spans where a refusal is a bug.
Patterns are matched without regard to case.

Args:
ctx: The evaluation context (reads the reply), or the
text to check.
patterns: Patterns that mark a refusal. None uses the
built-in English list.
extra_patterns: Patterns to add to `patterns`.
scan_first_chars: Look only at the start of the reply. A
long answer that says "I can't guarantee this" in
paragraph four is not a refusal.
name: Name of the primary score.

Returns:
refused (BOOLEAN): a pattern matched. Lower is better.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.refusal(ctx))

url_check​

url_check(ctx, *, allowed_schemes=('http', 'https'), allowed_domains=(), blocked_domains=(), allow_private=False, require_url=False, field='output', name='urls_ok')

Score names: urls_ok, url_count

Check the URLs in the text: well formed, allowed scheme and domain, not private.

A URL is anything of the form scheme://... . A domain rule
also covers its subdomains. An IP that is not public
(private, loopback, link-local, reserved) and localhost fail
unless `allow_private` is set.

Args:
ctx: The evaluation context, or the text to check.
allowed_schemes: Schemes allowed, such as "https".
allowed_domains: Domains allowed. Empty: any domain.
blocked_domains: Domains that fail the check.
allow_private: Allow private and local addresses.
require_url: Fail when the text has no URL.
field: "output" (the assistant's reply), "input" (the
user's messages) or "both".
name: Name of the primary score.

Returns:
urls_ok (BOOLEAN): every URL passed.
url_count (NUMERIC): how many URLs the text has.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.url_check(ctx))

prompt_injection​

prompt_injection(ctx, *, field='input', extra_patterns=(), min_base64_chars=120, threshold=0.5, name='injection_detected')

Score names: injection_detected, injection_risk, injection_signals

Flag text that tries to override the model's instructions.

Signals, each with a weight: "ignore previous instructions"
phrasing (0.6), asking for the system prompt (0.4), role
markers such as "system:" or <|im_start|> (0.4), persona
switches such as "developer mode" (0.3), hidden Unicode tag
characters (0.6), three or more invisible format characters
(0.2), a long base64 blob that decodes (0.3), and each of
your `extra_patterns` (0.5). The risk is the sum of the
signals found, capped at 1. A heuristic: it catches the
common attacks, not a determined one.

Args:
ctx: The evaluation context, or the text to check.
field: "input": the user's and the tools' messages, where
direct and indirect injection arrive. "output": the
reply. "both".
extra_patterns: Python regular expressions to add as
signals.
min_base64_chars: Shortest base64 run that counts.
threshold: Risk at which the text counts as an injection.
name: Name of the primary score.

Returns:
injection_detected (BOOLEAN): risk reached `threshold`.
Lower is better.
injection_risk (NUMERIC, 0 to 1): lower is better.
injection_signals (CATEGORICAL): the signals found,
comma separated, or "none".

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.prompt_injection(ctx))

conversation_degeneration​

conversation_degeneration(ctx, *, turns_to_consider=5, min_words=8, threshold=0.6, name='degenerated')

Score names: degenerated, degeneration

Check whether the assistant starts to repeat itself, in the reply or across turns.

Point it at chat spans: their input holds the conversation
so far, and the output is the newest reply. Three signals,
each 0 (healthy) to 1 (degenerate):

- repetition: share of the reply's word 3-grams that occur
more than once in it (a reply that loops on a phrase)
- turn similarity: how close the reply is to the most
similar earlier assistant turn (the same answer again)
- low diversity: 1 minus the word entropy of the reply, over
the most it could be for its length

Args:
ctx: The evaluation context, or the reply text (then only
repetition and low diversity count).
turns_to_consider: Earlier assistant turns to compare
against.
min_words: Replies shorter than this are not compared
across turns: two "Done." replies are not a sign of
degeneration.
threshold: The degeneration score at which the reply
counts as degenerated.
name: Name of the primary score.

Returns:
degenerated (BOOLEAN): degeneration reached `threshold`.
Lower is better.
degeneration (NUMERIC, 0 to 1): the largest of the three
signals. Lower is better. Not returned when the reply
has no text.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.conversation_degeneration(ctx))

knowledge_retention​

knowledge_retention(ctx, *, turns_to_consider=5, contradiction_kinds=('email', 'phone', 'name'), name='knowledge_retained')

Score names: knowledge_retained, knowledge_retention

Check that the reply does not forget what the user already said.

Point it at chat spans: their input holds the conversation
so far. It collects facts the user stated in earlier turns
(email, phone, order or ticket number, name), then fails the
reply when it asks for one of those again, or quotes a
different value of the same kind.

Args:
ctx: The evaluation context.
turns_to_consider: Earlier user turns to read facts from.
contradiction_kinds: Kinds where the reply quoting a
different value is a mistake: "email", "phone",
"order_id", "name". Order numbers are left out by
default: an agent that opens a new ticket quotes a
new number, and that is correct.
name: Name of the primary score.

Returns:
knowledge_retained (BOOLEAN): no fact was forgotten.
knowledge_retention (NUMERIC, 0 to 1): the share of facts
kept. Not returned when the user stated no facts.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.knowledge_retention(ctx))

grounding​

grounding(ctx, *, context=None, context_key='context', min_overlap=0.6, threshold=0.8, name='grounded')

Score names: grounded, grounding

Score the share of the reply's sentences whose words appear in the context.

A sentence counts as grounded when at least `min_overlap` of
its content words (no stop words, 3 letters or more) appear
in the context. Sentences with no content words are not
counted. A word check, not an entailment model: it catches a
reply that talks about things the context never mentions.

The context is, in order: the `context` argument, the span
attribute `context_key`, else the text of every input
message, where a RAG call usually puts what it retrieved.

Args:
ctx: The evaluation context, or the reply text (then give
`context`).
context: The context text, or a list of passages.
context_key: The span attribute that holds the context.
min_overlap: Share of a sentence's words, 0 to 1, that
must be in the context.
threshold: Share of grounded sentences at which the reply
counts as grounded.
name: Name of the primary score.

Returns:
grounded (BOOLEAN): grounding reached `threshold`.
grounding (NUMERIC, 0 to 1): the share of grounded
sentences.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.grounding(ctx))

unsupported_numbers​

unsupported_numbers(ctx, *, context=None, context_key='context', check_entities=True, name='numbers_supported')

Score names: numbers_supported, unsupported_count

Find numbers and names in the reply that the context does not have.

Numbers match by value ("1,200" equals "1200"). One-digit
whole numbers and list positions are ignored. Names are runs
of capitalised words that do not start a sentence, matched in
any case. The context comes from the same places as in
`grounding`.

Args:
ctx: The evaluation context, or the reply text (then give
`context`).
context: The context text, or a list of passages.
context_key: The span attribute that holds the context.
check_entities: Also check names.
name: Name of the primary score.

Returns:
numbers_supported (BOOLEAN): everything found is in the
context.
unsupported_count (NUMERIC): how many are not. Lower is
better.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.unsupported_numbers(ctx))

citation_check​

citation_check(ctx, *, sources=None, context=None, context_key='context', require_citation=False, check_urls=True, name='citations_ok')

Score names: citations_ok, citation_count

Check that every citation in the reply points at a real source.

A marker such as [2], [1, 3] or [2-4] must name a source: a
number from 1 to the length of `sources` when given, else a
number the context marks as [n] or as a numbered line. A URL
in the reply must appear in the sources or the context. The
context comes from the same places as in `grounding`.

Args:
ctx: The evaluation context, or the reply text (then give
`context` or `sources`).
sources: The list of sources the reply may cite, in order.
context: The context text, or a list of passages.
context_key: The span attribute that holds the context.
require_citation: Fail a reply with no citation at all.
check_urls: Check the URLs in the reply too.
name: Name of the primary score.

Returns:
citations_ok (BOOLEAN): every citation resolves.
citation_count (NUMERIC): markers and URLs found.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.citation_check(ctx))

tool_call_validity​

tool_call_validity(ctx, *, allowed_tools=(), required_args=None, max_identical_calls=2, name='tool_calls_valid')

Score names: tool_calls_valid, tool_call_loop

Check that the tool calls are well formed, allowed and not stuck in a loop.

Checks each call's arguments parse as a JSON object and carry
the arguments you require, that the tool is on the allow
list, and that the model does not repeat a call it already
made with the same arguments earlier in the conversation.
On an experiment item it also reads the agent's tool steps
in ctx.trace.spans.

Args:
ctx: The evaluation context.
allowed_tools: Tool names allowed. Empty: any tool.
required_args: {tool name: [argument names]} each call
of that tool must carry.
max_identical_calls: The same call (name and arguments)
seen this many times counts as a loop.
name: Name of the primary score.

Returns:
tool_calls_valid (BOOLEAN): every call is well formed and
allowed.
tool_call_loop (BOOLEAN): a call repeated. Lower is
better. Not returned when there are no calls.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.tool_call_validity(ctx))

tools_used​

tools_used(ctx)
Return the names of the tools the agent called, in order.

Reads the tool steps in ctx.trace.spans. When the trace has
none (a live span), reads the tool calls in the output.

Args:
ctx: The evaluation context.

Returns:
A list of tool names.

tool_trajectory​

tool_trajectory(ctx, *, expected=(), mode='strict', name='trajectory_match')

Score names: trajectory_match, tool_correctness

Check the agent's tool calls against the tools you expect.

Like LangChain agentevals trajectory match and DeepEval
ToolCorrectness. Calls are compared by tool name.

Args:
ctx: The evaluation context (see `tools_used`), or a
list of tool names.
expected: The tool names you expect, in order.
mode: "strict": the same calls in the same order.
"unordered": the same calls in any order.
"subset": every call is an expected one (no extra
tool). "superset": every expected call happened
(extra calls allowed). Repeated names count.
name: Name of the primary score.

Returns:
trajectory_match (BOOLEAN): the calls meet `mode`.
tool_correctness (NUMERIC, 0 to 1): the share of expected
calls that happened. 1 when nothing is expected.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.tool_trajectory(ctx))

token_latency_budget​

token_latency_budget(ctx, *, max_output_tokens=2000, max_input_tokens=None, max_duration_ms=30000, max_time_to_first_token_ms=5000, name='within_budget')

Score names: within_budget, budget_used

Check that the call stays inside its token, duration and time to first token limits.

Reads the span's own attributes, so it needs no text. A span
that reports none of them passes and says so.

Args:
ctx: The evaluation context, or the metadata dict.
max_output_tokens: None: not checked.
max_input_tokens: None: not checked.
max_duration_ms: None: not checked.
max_time_to_first_token_ms: None: not checked.
name: Name of the primary score.

Returns:
within_budget (BOOLEAN): no limit was passed.
budget_used (NUMERIC): the largest share of any limit;
above 1 is over. Lower is better. Not returned when
the span reports none of the attributes.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.token_latency_budget(ctx))

span_count​

span_count(ctx, *, kind=None, max_spans=None, name='span_count')

Score names: span_count, within_span_limit

Count the agent's steps in ctx.trace.spans.

The trace is filled on an experiment item and empty on a live
span, so on a live span the count is 0.

Args:
ctx: The evaluation context, or a list of span dicts.
kind: Count only steps of this kind: "llm", "tool",
"agent" or "span". None: every step.
max_spans: Most steps allowed. None: not checked.
name: Name of the primary score.

Returns:
span_count (NUMERIC): the count. Lower is better.
within_span_limit (BOOLEAN): the count is at most
`max_spans`. Only when `max_spans` is set.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.span_count(ctx))

error_spans​

error_spans(ctx, *, max_errors=0, name='no_errors')

Score names: no_errors, error_count

Count the agent's steps that failed, in ctx.trace.spans.

A step failed when its `error` field is set, or its `status`
is "error".

Args:
ctx: The evaluation context, or a list of span dicts.
max_errors: Most failed steps allowed.
name: Name of the primary score.

Returns:
no_errors (BOOLEAN): at most `max_errors` steps failed.
error_count (NUMERIC): failed steps. Lower is better.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.error_spans(ctx))

repeated_tool_calls​

repeated_tool_calls(ctx, *, max_repeats=2, consecutive_only=False, name='tool_loop')

Score names: tool_loop, most_repeated

Detect the agent calling the same tool with the same input again and again.

Reads the tool steps in ctx.trace.spans; with none (a live
span), the tool calls in the output.

Args:
ctx: The evaluation context, or a list of span dicts.
max_repeats: The same call seen this many times counts as
a loop.
consecutive_only: Count only calls repeated back to back.
name: Name of the primary score.

Returns:
tool_loop (BOOLEAN): a call repeated `max_repeats` times.
Lower is better.
most_repeated (NUMERIC): the most times one call was
made. Lower is better.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.repeated_tool_calls(ctx))

total_duration​

total_duration(ctx, *, max_duration_ms=None, name='total_duration_ms')

Score names: total_duration_ms, within_duration

Measure how long the agent ran, from its trace.

With agent steps in ctx.trace.spans, the longest agent step;
else the sum of every step's duration_ms (steps can overlap,
so the sum is an upper bound); with no steps (a live span),
the span's own duration_ms attribute.

Args:
ctx: The evaluation context, or a list of span dicts.
max_duration_ms: Longest time allowed. None: not checked.
name: Name of the primary score.

Returns:
total_duration_ms (NUMERIC): lower is better.
within_duration (BOOLEAN): at most `max_duration_ms`.
Only when it is set.

Example:

from oodle_eval.v1 import metrics

def evaluate(ctx):
return EvaluationResult(scores=metrics.total_duration(ctx))

text​

Import path: oodle_eval.v1.text

Read text, turns and tool calls out of a span, whatever its shape.

A span's input and output arrive as a plain string, one message
dict, or a list of messages. A message holds its text as a
`content` string, as a list of content blocks (OpenAI and
Anthropic), or as OpenTelemetry gen_ai `parts`. The helpers here
read all of these, so a check does not care which SDK wrote the
span.

Every function takes either the evaluation context `ctx` or the
value itself, where the summary says so.
FunctionSummary
message_textReturn the plain text of a message value.
replyReturn the text of the assistant's reply in the output.
user_textReturn the text of every user message in the input.
turnsReturn the conversation as a list of turns.
last_user_messageReturn the text of the last user turn in the input.
tool_callsReturn every tool call in a value, in order.
wordsReturn the words of a text.
sentencesReturn the sentences of a text.
normalizeReturn a text in lower case, with no punctuation and single spaces.
extract_jsonReturn the first JSON value found in a text.
metadata_valueReturn the first span attribute found among names.
metadata_numberReturn the first span attribute among names that is a number.

message_text​

message_text(value, role=None, default_role=None)
Return the plain text of a message value.

Args:
value: A string, a message dict, or a list of messages
and strings. Any other JSON value is returned as its
JSON text.
role: Keep only messages of this role ("user",
"assistant", "system", "tool"). None keeps every
message.
default_role: The role of a message that has no `role`
key. None means `role`, so a bare message counts as
the role asked for.

Returns:
The text of the kept messages, joined by new lines. A
string value is returned as it is, whatever `role` is.

reply​

reply(ctx)
Return the text of the assistant's reply in the output.

Args:
ctx: The evaluation context, or the output value itself.

Returns:
The assistant messages' text. A message with no role
counts as the assistant's.

user_text​

user_text(ctx)
Return the text of every user message in the input.

Args:
ctx: The evaluation context, or the input value itself.

Returns:
The user messages' text. A message with no role counts
as the user's. A plain string input is returned as it is.

turns​

turns(ctx, include_reply=True)
Return the conversation as a list of turns.

The shape is the thread shape DeepEval, Ragas and Opik use.

Args:
ctx: The evaluation context.
include_reply: Add the output's messages at the end.

Returns:
A list of {"role": str, "content": str}. An input
message with no role is a user turn; an output message
with no role is an assistant turn. A plain string input
is one user turn. Turns with no text are left out.

last_user_message​

last_user_message(ctx)
Return the text of the last user turn in the input.

Args:
ctx: The evaluation context.

Returns:
The text, or "" when the input has no user turn.

tool_calls​

tool_calls(ctx_or_value)
Return every tool call in a value, in order.

Reads OpenAI chat `tool_calls`, OpenAI Responses
`function_call` items, OpenTelemetry gen_ai parts of type
`tool_call`, and Anthropic `tool_use` content blocks.

Args:
ctx_or_value: The evaluation context (reads the output)
or a message value.

Returns:
A list of {"id", "name", "arguments"}. `arguments` is
parsed when it is a JSON string, else the value as it
came (a string that is not JSON stays a string). `name`
is "" when the call has none.

words​

words(text, lower=True)
Return the words of a text.

Args:
text: The text.
lower: Lower-case the words first.

Returns:
A list of runs of letters, digits and underscores.

sentences​

sentences(text)
Return the sentences of a text.

A full stop, question mark or exclamation mark followed by a
space, or a blank line, ends a sentence.

Args:
text: The text.

Returns:
A list of sentences with the space around them removed.

normalize​

normalize(text)
Return a text in lower case, with no punctuation and single spaces.

Use it before an exact comparison, so "Paris." equals "paris".

Args:
text: The text.

Returns:
The normalised text.

extract_json​

extract_json(text)
Return the first JSON value found in a text.

Tries, in order: the whole text, the first fenced code
block (```json or ```), then the first balanced {...} or
[...] that parses.

Args:
text: The text. A dict or list is returned as it is.

Returns:
The parsed value, or None when no JSON is found.

metadata_value​

metadata_value(ctx, *names, default=None)
Return the first span attribute found among `names`.

A numeric attribute can be stored with a type suffix, so
each name is also tried with `_int`, `_float`, `_double`,
`_bool`, `_str` and `_string`.

Args:
ctx: The evaluation context, or the metadata dict.
*names: Attribute names to try, first found wins.
default: Returned when no name is found.

Returns:
The attribute value as stored, or `default`.

metadata_number​

metadata_number(ctx, *names, default=None)
Return the first span attribute among `names` that is a number.

Like `metadata_value`, but skips a value that does not read
as a number, and returns a float.

Args:
ctx: The evaluation context, or the metadata dict.
*names: Attribute names to try, first found wins.
default: Returned when no name holds a number.

Returns:
The number as a float, or `default`.

combine​

Import path: oodle_eval.v1.combine

Combine several scores into one.

Every function takes a list of scores: a `Scores` from a check,
a plain list, or a list of such lists, which is read as one flat
list. A score can also be a dict with "name", "value",
"data_type" and "higher_is_better", as `ctx.scores` holds.

scores = metrics.tone(ctx) + metrics.readability(ctx)
overall = combine.weighted_mean(scores, weights={"tone_ok": 2})
return EvaluationResult(scores=scores + [overall])

The models for these are OpenAI's multi grader, promptfoo's
weights and assert sets, DeepEval's composite metrics, Opik's
AggregatedMetric and LangSmith's composite scores.
FunctionSummary
weighted_meanReturn the weighted mean of the scores' values as one NUMERIC score.
passedReturn whether one score reads as a pass.
all_passReturn True when every score passes, as a BOOLEAN score.
any_failReturn True when a score fails, as a BOOLEAN score.
minimumReturn the smallest value among the scores, as a NUMERIC score.
maximumReturn the largest value among the scores, as a NUMERIC score.
thresholdTurn one score into a BOOLEAN pass at a line.
renameReturn copies of the scores with new names.

weighted_mean​

weighted_mean(scores, weights=None, name='combined')
Return the weighted mean of the scores' values as one NUMERIC score.

A BOOLEAN score counts as 1 or 0. CATEGORICAL scores and
scores with no value are skipped. Values are used as they
are: pick scores that read the same way (turn a
lower-is-better score round with `threshold` or `passed`
first).

Args:
scores: The scores to combine.
weights: {score name: weight}, or a list of weights in
the order of `scores`. A name missing from the dict
weighs 1. None: every score weighs 1.
name: Name of the combined score.

Returns:
A NUMERIC Score, higher is better.

Raises:
ValueError: no score has a number to combine.

passed​

passed(score, at=None)
Return whether one score reads as a pass.

A BOOLEAN score passes when it is True, or when it is False
for a lower-is-better score (such as `refused`). A NUMERIC
score passes when it is at or above `at`, or at or below
`at` for a lower-is-better score. With no `at`, 0.5 is the
line: at or above for higher-is-better, below for
lower-is-better.

Args:
score: A Score or a score dict.
at: The pass line for a NUMERIC score.

Returns:
True or False.

Raises:
ValueError: the score is CATEGORICAL or has no value.

all_pass​

all_pass(scores, name='all_pass', at=None)
Return True when every score passes, as a BOOLEAN score.

Each score is read with `passed`. CATEGORICAL scores and
scores with no value are skipped.

Args:
scores: The scores to check.
at: The pass line for NUMERIC scores, as in `passed`.
name: Name of the result score.

Returns:
A BOOLEAN Score, higher is better. True when there is
nothing to check.

any_fail​

any_fail(scores, name='any_fail', at=None)
Return True when a score fails, as a BOOLEAN score.

The opposite of `all_pass`.

Args:
scores: The scores to check.
at: The pass line for NUMERIC scores, as in `passed`.
name: Name of the result score.

Returns:
A BOOLEAN Score, lower is better.

minimum​

minimum(scores, name='minimum')
Return the smallest value among the scores, as a NUMERIC score.

A BOOLEAN score counts as 1 or 0; CATEGORICAL scores and
scores with no value are skipped.

Args:
scores: The scores to read.
name: Name of the result score.

Returns:
A NUMERIC Score. Its direction is the one the scores
share, else unset.

Raises:
ValueError: no score has a number.

maximum​

maximum(scores, name='maximum')
Return the largest value among the scores, as a NUMERIC score.

A BOOLEAN score counts as 1 or 0; CATEGORICAL scores and
scores with no value are skipped.

Args:
scores: The scores to read.
name: Name of the result score.

Returns:
A NUMERIC Score. Its direction is the one the scores
share, else unset.

Raises:
ValueError: no score has a number.

threshold​

threshold(score, at, name=None, above=True)
Turn one score into a BOOLEAN pass at a line.

Args:
score: A Score or a score dict. A BOOLEAN counts as 1
or 0.
at: The line.
above: True: pass at or above `at`. False: pass below
`at`.
name: Name of the result score. None: the score's name
with "_ok" added.

Returns:
A BOOLEAN Score, higher is better. False when the score
has no number.

rename​

rename(scores, prefix='', names=None)
Return copies of the scores with new names.

Use it when two checks return scores of the same name, or to
tell apart the scores of one check run twice.

Args:
scores: The scores to rename.
prefix: Text to put before every name.
names: {old name: new name}, applied before `prefix`.

Returns:
A new Scores. The scores passed in do not change.

util​

Import path: oodle_eval.v1.util

Building blocks for writing your own checks, the way the built-in ones are written.

from oodle_eval.v1 import text, util

@util.declares("polite", "politeness")
def politeness(ctx, *, words=None, name="polite"):
words = util.opt(words, ["please", "thank"])
reply = util.subject_text(ctx).lower()
found = [w for w in words if w in reply]
return Scores([
util.boolean_score(util.opt(name, "polite"), bool(found), f"found {found}"),
util.numeric_score("politeness", len(found) / len(words), "share of words"),
])

A check built this way takes a context or a plain string, reads
None as "the default", and returns `Scores` whose names the
editor knows before the first run. Message helpers (reply,
turns, tool calls, JSON) are in `oodle_eval.v1.text`.

The sandbox gives each item 5 seconds: cut long text with
`MAX_TEXT_CHARS` and `MAX_SENTENCES` before anything that grows
faster than the text.
FunctionSummary
declaresRecord the score names a check returns, primary first.
optReturn value, or default when value is None.
is_contextReturn whether value is an evaluation context.
subject_textReturn the text a check reads: the reply, the user's turns, or both.
reference_or_expectedReturn the reference setting, else the experiment item's expected output.
require_referenceReturn a reference as text, or say how to set one.
boolean_scoreReturn a BOOLEAN Score.
numeric_scoreReturn a NUMERIC Score.
strip_codeReturn a text with code blocks, inline code and URLs made spaces.
fenced_blockReturn the first fenced code block of one of langs, else the whole text.
join_problemsReturn problems as one comment, the first limit named.
flat_textReturn any value as text: a string as it is, a list line by line, else JSON.
context_textReturn the retrieved context to check a reply against.
content_wordsReturn the words of a text that carry content.
find_numberReturn the first number in a value, as a float.
find_urlsReturn the URLs in a text, with trailing punctuation removed.
is_messagesReturn whether a value is a list of chat messages.

declares​

declares(*names)
Record the score names a check returns, primary first.

The library reference reads them from the function, so the
editor can offer a score name before the check has run once.

Args:
*names: The score names, primary first.

Returns:
A decorator that sets `__scores__` on the function and
returns the function unchanged.

opt​

opt(value, default)
Return `value`, or `default` when `value` is None.

An evaluator's unset setting arrives as None, so a check
reads each setting through this to get its own default.

Args:
value: The setting as passed.
default: The value to use for None.

Returns:
`value`, or `default`.

is_context​

is_context(value)
Return whether `value` is an evaluation context.

Args:
value: Anything.

Returns:
True for an object with `observation` and `trace`.

subject_text​

subject_text(ctx, field='output', limit=None)
Return the text a check reads: the reply, the user's turns, or both.

Args:
ctx: The evaluation context, a plain string (returned as
it is), or a message value (its assistant text).
field: "output" (the assistant's reply), "input" (the
user's messages) or "both". Used for a context only.
limit: The most characters to return. None: all.

Returns:
The text.

Raises:
ValueError: `field` is not one of the three.

reference_or_expected​

reference_or_expected(ctx, reference)
Return the reference setting, else the experiment item's expected output.

Args:
ctx: The evaluation context, or anything else.
reference: The reference setting.

Returns:
`reference` when set; else `ctx.expected_output` for a
context; else None.

require_reference​

require_reference(reference, check)
Return a reference as text, or say how to set one.

Args:
reference: The reference, a string or any JSON value.
check: The check's name, for the error message.

Returns:
The reference itself when it is a string, else its JSON
text.

Raises:
ValueError: `reference` is None.

boolean_score​

boolean_score(name, value, comment, higher_is_better=True)
Return a BOOLEAN Score.

Args:
name: The score name.
value: Read as True or False.
comment: A short reason.
higher_is_better: False for a score that flags a problem
when true, such as `refused`.

Returns:
A Score.

numeric_score​

numeric_score(name, value, comment, higher_is_better=True)
Return a NUMERIC Score.

Args:
name: The score name.
value: A number.
comment: A short reason.
higher_is_better: False for a count of problems or a
duration.

Returns:
A Score.

strip_code​

strip_code(value)
Return a text with code blocks, inline code and URLs made spaces.

Use it before a check of prose, so code does not read as
words.

Args:
value: The text.

Returns:
The text without code and URLs.

fenced_block​

fenced_block(value, *langs)
Return the first fenced code block of one of `langs`, else the whole text.

Args:
value: The text.
*langs: Fence labels to accept, such as "json". A block
with no label is also accepted.

Returns:
The block's content, or `value` when there is none.

join_problems​

join_problems(problems, limit=10)
Return problems as one comment, the first `limit` named.

Args:
problems: A list of short texts.
limit: The most to name; the rest are counted.

Returns:
"a; b; and 3 more", or "" for no problems.

flat_text​

flat_text(value)
Return any value as text: a string as it is, a list line by line, else JSON.

Args:
value: A string, a list, or any JSON value.

Returns:
The text. "" for None.

context_text​

context_text(ctx, context=None, context_key='context')
Return the retrieved context to check a reply against.

In order: the `context` argument, else the span attribute
`context_key`, else the text of every input message (a RAG
call usually puts what it retrieved in its prompt). Cut to
`MAX_TEXT_CHARS`.

Args:
ctx: The evaluation context, or a plain string (then only
`context` counts).
context: The context text, or a list of passages.
context_key: The span attribute that holds the context.

Returns:
The context text, or "".

content_words​

content_words(value)
Return the words of a text that carry content.

Args:
value: The text.

Returns:
Lower-case words of 3 or more characters that are not in
`STOPWORDS`.

find_number​

find_number(value)
Return the first number in a value, as a float.

Args:
value: A number (a bool is not one), or a text holding
one. "1,200" reads as 1200.

Returns:
The number, or None.

find_urls​

find_urls(value, limit=None)
Return the URLs in a text, with trailing punctuation removed.

Args:
value: The text.
limit: The most URLs to return. None: all.

Returns:
A list of URLs, in order.

is_messages​

is_messages(value)
Return whether a value is a list of chat messages.

Args:
value: Anything.

Returns:
True for a list with a dict that has "role" or "parts".

metrics.format​

Import path: oodle_eval.v1.metrics.format

The metrics functions defined in oodle_eval/v1/metrics/format.py. Import them from oodle_eval.v1.metrics.

FunctionSummary
json_validityCheck that the output is JSON and has the shape you expect.
regex_matchCheck that the text matches every required pattern and no forbidden one.
keyword_checkCheck that the text mentions the required words and none of the forbidden ones.
containsCheck that the text contains the given substrings.
length_budgetCheck that the reply stays inside a word and character range.
markdown_structureCheck the reply's Markdown: code fences, tables, headings and lists.
html_validityCheck that the HTML in the reply has balanced tags.
csv_validityCheck that the reply is CSV with the same number of fields in every row.
toml_validityCheck that the reply parses as TOML and has the keys you require.
sql_shapeCheck that generated SQL is only the kind of statement you allow.

metrics.quality​

Import path: oodle_eval.v1.metrics.quality

The metrics functions defined in oodle_eval/v1/metrics/quality.py. Import them from oodle_eval.v1.metrics.

FunctionSummary
readabilityScore Flesch reading ease and check the Flesch-Kincaid grade.
detect_languageReturn the language of a text and a confidence.
language_adherenceCheck that the reply is written in the language you expect.
toneFlag a reply that shouts, piles on exclamation marks, reads as negative, or uses a banned phrase.
sentimentScore the reply's sentiment from -1 (negative) to 1 (positive).

metrics.similarity​

Import path: oodle_eval.v1.metrics.similarity

The metrics functions defined in oodle_eval/v1/metrics/similarity.py. Import them from oodle_eval.v1.metrics.

FunctionSummary
exact_matchCheck that the reply equals a reference text.
levenshtein_ratioScore how close the reply is to a reference, by edit distance.
jaccard_similarityScore the overlap of the reply's and the reference's word sets.
rougeScore ROUGE-L, ROUGE-1 and ROUGE-2 F1 against a reference.
bleuScore sentence BLEU against a reference.
numeric_diffScore how close the first number in the reply is to a reference number.
json_diffScore how close the output's JSON is to a reference JSON value.
bm25_relevanceScore how relevant the reply is to the question or context, by BM25.
tfidf_similarityScore the cosine similarity of the reply and the question or context, by TF-IDF.

metrics.safety​

Import path: oodle_eval.v1.metrics.safety

The metrics functions defined in oodle_eval/v1/metrics/safety.py. Import them from oodle_eval.v1.metrics.

FunctionSummary
pii_leakCheck the reply for an email, phone, card number, national id, IP address or credential.
refusalCheck whether the model declined to do what it was asked.
url_checkCheck the URLs in the text: well formed, allowed scheme and domain, not private.
prompt_injectionFlag text that tries to override the model's instructions.

metrics.conversation​

Import path: oodle_eval.v1.metrics.conversation

The metrics functions defined in oodle_eval/v1/metrics/conversation.py. Import them from oodle_eval.v1.metrics.

FunctionSummary
conversation_degenerationCheck whether the assistant starts to repeat itself, in the reply or across turns.
knowledge_retentionCheck that the reply does not forget what the user already said.

metrics.context​

Import path: oodle_eval.v1.metrics.context

The metrics functions defined in oodle_eval/v1/metrics/context.py. Import them from oodle_eval.v1.metrics.

FunctionSummary
groundingScore the share of the reply's sentences whose words appear in the context.
unsupported_numbersFind numbers and names in the reply that the context does not have.
citation_checkCheck that every citation in the reply points at a real source.

metrics.agent​

Import path: oodle_eval.v1.metrics.agent

The metrics functions defined in oodle_eval/v1/metrics/agent.py. Import them from oodle_eval.v1.metrics.

FunctionSummary
tool_call_validityCheck that the tool calls are well formed, allowed and not stuck in a loop.
tools_usedReturn the names of the tools the agent called, in order.
tool_trajectoryCheck the agent's tool calls against the tools you expect.
token_latency_budgetCheck that the call stays inside its token, duration and time to first token limits.

metrics.trace​

Import path: oodle_eval.v1.metrics.trace

The metrics functions defined in oodle_eval/v1/metrics/trace.py. Import them from oodle_eval.v1.metrics.

FunctionSummary
span_countCount the agent's steps in ctx.trace.spans.
error_spansCount the agent's steps that failed, in ctx.trace.spans.
repeated_tool_callsDetect the agent calling the same tool with the same input again and again.
total_durationMeasure how long the agent ran, from its trace.

Support

If you need assistance or have any questions, please reach out to us through: