oodle_eval Library Reference
Code evaluators can import the oodle_eval library. It holds
the built-in checks as functions, helpers that read text and
tool calls out of a span, and functions that combine scores.
See Code Evaluators for how an evaluator
runs, and the Cookbook for
complete examples.
Version: 1.0.0
from oodle_eval.v1 import metrics, text, combine
Everything under oodle_eval.v1 keeps its behaviour: a later
release changes it only to fix a bug. A change that would move
your scores goes into a new version module, so an evaluator
that imports v1 gives the same scores after an upgrade.
Each function in metrics takes the evaluation context ctx
(or a plain string) as its first argument, takes its settings
as keyword arguments, and returns Scores: a list of Score
objects with the primary score first. Add the lists together
to return scores from several checks.
The evaluation context
| Path | Value |
|---|---|
ctx.observation.input | The span's input: a string, a message list (OpenAI, Anthropic or OpenTelemetry gen_ai shape) or any JSON value. |
ctx.observation.output | The span's output, in the same shapes as the input. A structured output or tool result arrives parsed. |
ctx.observation.metadata | The span's attributes as a dict. A numeric attribute can carry a type suffix, such as duration_ms_int. |
ctx.trace.trace_id | The trace id. Empty for a live span. |
ctx.trace.spans | The agent's steps on an experiment item: dicts with name, kind (llm, tool, agent or span), tool_name, model, input, output, duration_ms and error. Empty for a live span. |
ctx.trace.tools_called() | The tool name of each tool step in ctx.trace.spans, in order. |
ctx.params | The evaluator's settings as a dict: the value set on the evaluator, else the template's default, else None. Every setting the template declares is a key. |
ctx.experiment | The dataset item on an experiment run: item_id, expected_output, item_metadata, run_id and dataset_id. Every field is None (metadata {}) on a live span. |
ctx.expected_output | Short for ctx.experiment.expected_output: the item's expected output, or None on a live span. Checks that compare against a reference use it when their reference setting is empty. |
ctx.scores | The scores other evaluators wrote on the same span, as {evaluator name: Scores}. The evaluator runs after the ones its code names here. Use ctx.scores.get(name): an evaluator with no score on the span is not a key. |
ctx.score(evaluator, score_name=None, default=None) | The value of one score another evaluator wrote: the score called score_name, else its first score, else default. |
Runtime types
Score and EvaluationResult are already in the globals of
your evaluator code. Import the others, and import all of them
in a shared library, from oodle_eval.runtime.
Score
Score(value, data_type='NUMERIC', comment=None, name=None, higher_is_better=None)
One named score.
Args:
value: A number for NUMERIC, True or False for BOOLEAN,
or a string for CATEGORICAL.
data_type: "NUMERIC", "BOOLEAN" or "CATEGORICAL".
comment: A short reason, shown next to the score.
name: The score name. An evaluator that returns more
than one score must name each.
higher_is_better: Which way is good for this score. None
uses the evaluator's own declaration.
EvaluationResult
EvaluationResult(scores)
What `evaluate(ctx)` returns.
Args:
scores: A list of `Score` (a `Scores` is a list).
Scores
Scores(iterable=())
A list of `Score` that also finds a score by its name.
Every built-in check returns one. It is a list, so
`scores + [Score(...)]` and `EvaluationResult(scores=scores)`
work as they do for a list. `scores["name"]` gives the score
of that name (KeyError when there is none); an int index
still gives the score at that place.
EvaluationContext
EvaluationContext(observation, trace=None, params=None, scores=None, experiment=None)
What `evaluate(ctx)` gets.
Args:
observation: The `ObservationContext` under evaluation.
trace: The `TraceContext`. Empty for a live span.
params: The evaluator's settings, resolved: the value set
on the evaluator, else the template's default, else
None. Every setting the template declares is a key.
An empty dict when the template declares none.
scores: The scores other evaluators wrote on the same
span, as {evaluator name: Scores}. Only evaluators
the code names in `ctx.scores[...]`,
`ctx.scores.get(...)` or `ctx.score(...)` are here.
An empty dict when there are none.
experiment: The `ExperimentContext` of an experiment item.
Empty (every field None, metadata {}) on a live span.
ObservationContext
ObservationContext(input=None, output=None, metadata=None)
The span or dataset item under evaluation.
Args:
input: The input, as the span stored it: a string, a
message list or any JSON value.
output: The output, in the same shapes as `input`.
metadata: The span's attributes, as a dict.
TraceContext
TraceContext(trace_id='', spans=None)
The agent's trace, as ordered steps.
Args:
trace_id: The trace id. Empty for a live span.
spans: A list of dicts with name, kind (llm, tool, agent
or span), tool_name, model, input, output,
duration_ms and error. Filled for an experiment
item; empty for a live span.
ExperimentContext
ExperimentContext(item_id=None, expected_output=None, item_metadata=None, run_id=None, dataset_id=None)
The dataset item an experiment runs, empty on a live span.
Args:
item_id: The dataset item's id. None on a live span.
expected_output: The item's expected output, as the
dataset stores it: a string or any JSON value. None
when the item has none.
item_metadata: The item's metadata dict. {} when none.
run_id: The experiment run's id.
dataset_id: The dataset's id.
metrics
Import path: oodle_eval.v1.metrics
Built-in checks. Each returns `Scores`, primary score first.
Every check takes the evaluation context `ctx` first, or a plain
string to check that text. Its settings are keyword arguments,
so a template can pass its settings straight through:
def evaluate(ctx):
return EvaluationResult(scores=metrics.keyword_check(ctx, **ctx.params))
A setting passed as None takes its default, so an evaluator that
leaves a setting empty gets the check's own behaviour. Where the
summary says so, None on a limit means "do not check this limit".
`name` renames the primary score; the other scores keep their
names. Use `combine.rename` to change them all.
Checks that compare against a reference take it as the
`reference` setting. When that is None, they use the experiment
item's expected output (`ctx.expected_output`), so the same check
works on live spans with a fixed reference and in experiments
with the dataset's ground truth.
The checks live in one module per area, and every one is also
here, so `metrics.keyword_check` and
`from oodle_eval.v1.metrics import keyword_check` both work:
- `metrics.format`: JSON, patterns, keywords, length, Markdown,
HTML, CSV, TOML, SQL
- `metrics.quality`: readability, language, tone, sentiment
- `metrics.similarity`: comparison with a reference, a question
or a context
- `metrics.safety`: PII and secrets, refusals, prompt injection,
URLs
- `metrics.conversation`: repetition and forgotten facts
- `metrics.context`: grounding, unsupported numbers, citations
- `metrics.agent`: tool calls, tool trajectory, token and
latency budget
- `metrics.trace`: steps, failures, loops and duration of the
agent's trace
To write a check of your own the same way, see
`oodle_eval.v1.util`.
| Function | Summary |
|---|---|
json_validity | Check that the output is JSON and has the shape you expect. |
regex_match | Check that the text matches every required pattern and no forbidden one. |
keyword_check | Check that the text mentions the required words and none of the forbidden ones. |
contains | Check that the text contains the given substrings. |
length_budget | Check that the reply stays inside a word and character range. |
markdown_structure | Check the reply's Markdown: code fences, tables, headings and lists. |
html_validity | Check that the HTML in the reply has balanced tags. |
csv_validity | Check that the reply is CSV with the same number of fields in every row. |
toml_validity | Check that the reply parses as TOML and has the keys you require. |
sql_shape | Check that generated SQL is only the kind of statement you allow. |
readability | Score Flesch reading ease and check the Flesch-Kincaid grade. |
detect_language | Return the language of a text and a confidence. |
language_adherence | Check that the reply is written in the language you expect. |
tone | Flag a reply that shouts, piles on exclamation marks, reads as negative, or uses a banned phrase. |
sentiment | Score the reply's sentiment from -1 (negative) to 1 (positive). |
exact_match | Check that the reply equals a reference text. |
levenshtein_ratio | Score how close the reply is to a reference, by edit distance. |
jaccard_similarity | Score the overlap of the reply's and the reference's word sets. |
rouge | Score ROUGE-L, ROUGE-1 and ROUGE-2 F1 against a reference. |
bleu | Score sentence BLEU against a reference. |
numeric_diff | Score how close the first number in the reply is to a reference number. |
json_diff | Score how close the output's JSON is to a reference JSON value. |
bm25_relevance | Score how relevant the reply is to the question or context, by BM25. |
tfidf_similarity | Score the cosine similarity of the reply and the question or context, by TF-IDF. |
pii_leak | Check the reply for an email, phone, card number, national id, IP address or credential. |
refusal | Check whether the model declined to do what it was asked. |
url_check | Check the URLs in the text: well formed, allowed scheme and domain, not private. |
prompt_injection | Flag text that tries to override the model's instructions. |
conversation_degeneration | Check whether the assistant starts to repeat itself, in the reply or across turns. |
knowledge_retention | Check that the reply does not forget what the user already said. |
grounding | Score the share of the reply's sentences whose words appear in the context. |
unsupported_numbers | Find numbers and names in the reply that the context does not have. |
citation_check | Check that every citation in the reply points at a real source. |
tool_call_validity | Check that the tool calls are well formed, allowed and not stuck in a loop. |
tools_used | Return the names of the tools the agent called, in order. |
tool_trajectory | Check the agent's tool calls against the tools you expect. |
token_latency_budget | Check that the call stays inside its token, duration and time to first token limits. |
span_count | Count the agent's steps in ctx.trace.spans. |
error_spans | Count the agent's steps that failed, in ctx.trace.spans. |
repeated_tool_calls | Detect the agent calling the same tool with the same input again and again. |
total_duration | Measure how long the agent ran, from its trace. |
json_validity
json_validity(ctx, *, extract_from_code_fence=True, required_keys=(), key_types=None, schema=None, name='is_json')
Score names: is_json, json_schema_ok
Check that the output is JSON and has the shape you expect.
A structured output or tool result that arrives already
parsed counts as JSON with no second parse.
Args:
ctx: The evaluation context (reads the output), a string,
or a parsed JSON value.
extract_from_code_fence: Read the JSON inside a ```json
fenced block when the reply wraps it in prose. False
demands bare JSON.
required_keys: Keys the top-level object must carry.
key_types: {key: type} for top-level keys, where type is
"string", "number", "integer", "boolean", "array",
"object" or "null". A missing key is reported by
`required_keys`, not here.
schema: A JSON Schema subset the value must match: type,
required, properties, items, enum, minimum, maximum,
minLength and maxLength. Other keywords are ignored.
name: Name of the primary score.
Returns:
is_json (BOOLEAN): the output parses as JSON.
json_schema_ok (BOOLEAN): it parses and matches
`required_keys`, `key_types` and `schema`.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.json_validity(ctx))
regex_match
regex_match(ctx, *, must_match=(), must_not_match=(), field='output', name='regex_pass')
Score names: regex_pass
Check that the text matches every required pattern and no forbidden one.
Patterns use Python's re syntax, in multi-line mode. Put
(?i) at the start of a pattern to ignore case. A pattern that
does not compile fails the check and says so.
Args:
ctx: The evaluation context, or the text to check.
must_match: Patterns that must each match.
must_not_match: Patterns that must not match.
field: "output" (the assistant's reply), "input" (the
user's messages) or "both".
name: Name of the primary score.
Returns:
regex_pass (BOOLEAN): every rule held.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.regex_match(ctx))
keyword_check
keyword_check(ctx, *, required=(), forbidden=(), mode='all', case_sensitive=False, field='output', name='keywords_ok')
Score names: keywords_ok, keyword_coverage
Check that the text mentions the required words and none of the forbidden ones.
Matching is on whole words, so "art" does not match "start".
Use `contains` to match any part of a word.
Args:
ctx: The evaluation context, or the text to check.
required: Keywords or phrases the text must mention.
forbidden: Keywords or phrases the text must not mention.
mode: "all": every required keyword must appear. "any":
one is enough.
case_sensitive: Match case exactly.
field: "output" (the assistant's reply), "input" (the
user's messages) or "both".
name: Name of the primary score.
Returns:
keywords_ok (BOOLEAN): the required rule held and no
forbidden keyword appeared.
keyword_coverage (NUMERIC, 0 to 1): the share of required
keywords found. Only when `required` is set.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.keyword_check(ctx))
contains
contains(ctx, *, values=(), mode='any', case_sensitive=False, field='output', name='contains')
Score names: contains
Check that the text contains the given substrings.
Unlike `keyword_check`, a value matches anywhere, also inside
a word. Like OpenAI's `like` / `ilike` string check and Opik's
Contains.
Args:
ctx: The evaluation context, or the text to check.
values: The substrings to look for. Empty: the experiment
item's expected output (`ctx.expected_output`), a text
or a list of texts.
mode: "any": one value is enough. "all": every value
must appear. "none": no value may appear.
case_sensitive: Match case exactly.
field: "output" (the assistant's reply), "input" (the
user's messages) or "both".
name: Name of the primary score.
Returns:
contains (BOOLEAN): the text meets `mode`.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.contains(ctx))
length_budget
length_budget(ctx, *, min_words=1, max_words=300, min_chars=None, max_chars=4000, name='within_length')
Score names: within_length
Check that the reply stays inside a word and character range.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
min_words: Fewest words. None: not checked.
max_words: Most words. None: not checked.
min_chars: Fewest characters. None: not checked.
max_chars: Most characters. None: not checked.
name: Name of the primary score.
Returns:
within_length (BOOLEAN): every bound held.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.length_budget(ctx))
markdown_structure
markdown_structure(ctx, *, required_headings=(), min_headings=0, require_list=False, require_table=False, name='markdown_ok')
Score names: markdown_ok, heading_count
Check the reply's Markdown: code fences, tables, headings and lists.
Code fences (``` or ~~~) must be closed. A table's rows must
have as many cells as its header. Text inside a code fence is
not read as Markdown.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
required_headings: Headings the reply must have, matched
in any case, with no punctuation, inside the heading
text.
min_headings: Fewest headings.
require_list: The reply must have a bullet or numbered
list.
require_table: The reply must have a table.
name: Name of the primary score.
Returns:
markdown_ok (BOOLEAN): every rule held.
heading_count (NUMERIC): how many headings the reply has.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.markdown_structure(ctx))
html_validity
html_validity(ctx, *, allowed_tags=(), name='html_valid')
Score names: html_valid
Check that the HTML in the reply has balanced tags.
Reads a ```html fenced block when there is one, else the
whole reply. Void tags (br, img, ...) need no end tag, and
tags whose end tag HTML makes optional (p, li, td, ...) are
not reported when left open.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
allowed_tags: Tag names allowed. Empty: any tag.
name: Name of the primary score.
Returns:
html_valid (BOOLEAN): every tag is closed in order, and
allowed.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.html_validity(ctx))
csv_validity
csv_validity(ctx, *, delimiter=None, has_header=True, required_columns=(), min_rows=1, name='csv_valid')
Score names: csv_valid, row_count
Check that the reply is CSV with the same number of fields in every row.
Reads a ```csv fenced block when there is one, else the whole
reply. Blank lines are skipped.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
delimiter: The field separator. Empty: the one of , ; tab
and | that the first row uses most.
has_header: The first row names the columns.
required_columns: Column names the header must have.
Needs `has_header`.
min_rows: Fewest data rows, the header not counted.
name: Name of the primary score.
Returns:
csv_valid (BOOLEAN): the rules held.
row_count (NUMERIC): data rows, the header not counted.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.csv_validity(ctx))
toml_validity
toml_validity(ctx, *, required_keys=(), name='toml_valid')
Score names: toml_valid
Check that the reply parses as TOML and has the keys you require.
Reads a ```toml fenced block when there is one, else the
whole reply.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
required_keys: Keys the document must have. A dotted key
such as "server.port" looks inside tables.
name: Name of the primary score.
Returns:
toml_valid (BOOLEAN): it parses and has every required key.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.toml_validity(ctx))
sql_shape
sql_shape(ctx, *, allowed_statements=('SELECT',), allowed_keywords=(), max_statements=1, name='sql_ok')
Score names: sql_ok, statement_type
Check that generated SQL is only the kind of statement you allow.
Reads a ```sql fenced block when there is one, else the whole
reply. Comments and string literals are left out before the
check, so a keyword inside a string does not count. This is a
pattern check, not a SQL parser: use it to catch a model that
writes where it should only read.
Args:
ctx: The evaluation context (reads the reply), or the
SQL text.
allowed_statements: Statement types allowed, such as
"SELECT" or "INSERT". A WITH statement counts as the
write it does, else as SELECT.
allowed_keywords: Write keywords (DROP, DELETE, INSERT,
UPDATE, ALTER, CREATE, TRUNCATE, GRANT, ...) allowed
anywhere. An allowed statement type is also allowed
as a keyword.
max_statements: Most statements, split on ";".
name: Name of the primary score.
Returns:
sql_ok (BOOLEAN): every rule held.
statement_type (CATEGORICAL): the first statement's type,
or "none".
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.sql_shape(ctx))
readability
readability(ctx, *, min_grade=None, max_grade=10, min_words=20, name='readability_ok')
Score names: readability_ok, reading_ease
Score Flesch reading ease and check the Flesch-Kincaid grade.
Syllables are counted with a vowel-group rule, so the grade
is an estimate, close enough to tell a plain answer from a
dense one. Code blocks, URLs and markdown are left out.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
min_grade: Lowest grade allowed. None: not checked.
max_grade: Highest grade allowed. None: not checked.
min_words: Replies shorter than this are too short to
grade and pass.
name: Name of the primary score.
Returns:
readability_ok (BOOLEAN): the grade is in range.
reading_ease (NUMERIC, 0 to 1): Flesch reading ease over
100. Not returned for a reply too short to grade.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.readability(ctx))
detect_language
detect_language(value)
Return the language of a text and a confidence.
Latin-script languages are told apart by their most common
words; other scripts by their characters. Mixed or very short
text can read as "unknown".
Args:
value: The text.
Returns:
(code, confidence): an ISO 639-1 code (en, es, fr, de,
it, pt, nl, ru, uk, el, ar, he, hi, th, zh, ja, ko) or
"unknown", and a confidence from 0 to 1.
language_adherence
language_adherence(ctx, *, expected_language='en', min_words=5, name='language_match')
Score names: language_match, detected_language
Check that the reply is written in the language you expect.
No model is needed, so it runs in the sandbox. See
`detect_language` for how the language is found.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
expected_language: ISO 639-1 code: en, es, fr, de, it,
pt, nl, ru, uk, el, ar, he, hi, th, zh, ja or ko.
min_words: Replies shorter than this pass with no verdict.
name: Name of the primary score.
Returns:
language_match (BOOLEAN): the detected language is the
expected one.
detected_language (CATEGORICAL): the detected code, or
"unknown". Not returned for a reply too short to
detect.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.language_adherence(ctx))
tone
tone(ctx, *, max_uppercase_ratio=0.3, max_exclamations=2, min_sentiment=-0.3, forbidden_phrases=('calm down', 'obviously', 'as i already said', "that's not my problem", 'you should have', 'stupid', 'shut up'), extra_positive_words=(), extra_negative_words=(), name='tone_ok')
Score names: tone_ok, sentiment
Flag a reply that shouts, piles on exclamation marks, reads as negative, or uses a banned phrase.
Sentiment is a small word list with negation, scored the way
VADER normalises its compound score, so it needs no model.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
max_uppercase_ratio: Share of capital letters, over words
of 3 or more letters, above which the reply shouts.
Checked only on 20 or more such letters.
max_exclamations: Most exclamation marks allowed.
min_sentiment: Sentiment runs from -1 to 1. Below this
is negative.
forbidden_phrases: Phrases that fail the check, matched
in lower case anywhere in the reply.
extra_positive_words: Words to add to the positive list.
extra_negative_words: Words to add to the negative list.
name: Name of the primary score.
Returns:
tone_ok (BOOLEAN): no issue found.
sentiment (NUMERIC, -1 to 1): the reply's sentiment.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.tone(ctx))
sentiment
sentiment(ctx, *, extra_positive_words=(), extra_negative_words=(), name='sentiment')
Score names: sentiment, sentiment_label
Score the reply's sentiment from -1 (negative) to 1 (positive).
The same word-list method as `tone`: no model is needed.
Like Opik Sentiment.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
extra_positive_words: Words to add to the positive list.
extra_negative_words: Words to add to the negative list.
name: Name of the primary score.
Returns:
sentiment (NUMERIC, -1 to 1): the compound sentiment.
sentiment_label (CATEGORICAL): "positive" at 0.05 and
above, "negative" at -0.05 and below, else "neutral".
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.sentiment(ctx))
exact_match
exact_match(ctx, *, reference=None, normalize=False, case_sensitive=True, name='exact_match')
Score names: exact_match
Check that the reply equals a reference text.
Space at the start and end is ignored. Like autoevals
ExactMatch, Opik Equals and OpenAI's `eq` string check.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
reference: The expected text. None: the
experiment item's expected output (`ctx.expected_output`).
normalize: Compare after `text.normalize`: lower case,
no punctuation, single spaces.
case_sensitive: Match case exactly. Ignored when
`normalize` is set.
name: Name of the primary score.
Returns:
exact_match (BOOLEAN): the texts are equal.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.exact_match(ctx, reference=ctx.params.get("reference")))
levenshtein_ratio
levenshtein_ratio(ctx, *, reference=None, normalize=False, name='levenshtein_ratio')
Score names: levenshtein_ratio
Score how close the reply is to a reference, by edit distance.
1 minus the Levenshtein distance over the longer length.
Each text is compared on its first 1000 characters. Like
autoevals Levenshtein and Opik LevenshteinRatio.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
reference: The expected text. None: the
experiment item's expected output (`ctx.expected_output`).
normalize: Compare after `text.normalize`.
name: Name of the primary score.
Returns:
levenshtein_ratio (NUMERIC, 0 to 1): 1 is identical.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.levenshtein_ratio(ctx, reference=ctx.params.get("reference")))
jaccard_similarity
jaccard_similarity(ctx, *, reference=None, name='jaccard_similarity')
Score names: jaccard_similarity
Score the overlap of the reply's and the reference's word sets.
Words are compared in lower case with no punctuation.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
reference: The expected text. None: the
experiment item's expected output (`ctx.expected_output`).
name: Name of the primary score.
Returns:
jaccard_similarity (NUMERIC, 0 to 1): shared words over
all words. 1 when both texts have no words.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.jaccard_similarity(ctx, reference=ctx.params.get("reference")))
rouge
rouge(ctx, *, reference=None, name='rouge_l')
Score names: rouge_l, rouge_1, rouge_2
Score ROUGE-L, ROUGE-1 and ROUGE-2 F1 against a reference.
Tokens are normalised words (see `text.normalize`). ROUGE-L
uses the first 400 tokens of each text. Like OpenAI's
text_similarity grader and Opik ROUGE.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
reference: The expected text. None: the
experiment item's expected output (`ctx.expected_output`).
name: Name of the primary score.
Returns:
rouge_l (NUMERIC, 0 to 1): longest common subsequence F1.
rouge_1 (NUMERIC, 0 to 1): word overlap F1.
rouge_2 (NUMERIC, 0 to 1): word pair overlap F1.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.rouge(ctx, reference=ctx.params.get("reference")))
bleu
bleu(ctx, *, reference=None, max_n=4, name='bleu')
Score names: bleu
Score sentence BLEU against a reference.
Geometric mean of the 1- to `max_n`-gram precisions, with
add-one smoothing for n above 1 so a short reply does not
score 0, times the brevity penalty. Tokens are normalised
words. Like Opik's and OpenAI's BLEU.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
reference: The expected text. None: the
experiment item's expected output (`ctx.expected_output`).
max_n: The longest n-gram, 1 to 4.
name: Name of the primary score.
Returns:
bleu (NUMERIC, 0 to 1): 1 is identical.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.bleu(ctx, reference=ctx.params.get("reference")))
numeric_diff
numeric_diff(ctx, *, reference=None, name='numeric_diff')
Score names: numeric_diff
Score how close the first number in the reply is to a reference number.
1 minus the difference over the sum of the two sizes, as
autoevals NumericDiff does. "1,200" reads as 1200.
Args:
ctx: The evaluation context (reads the reply), the text,
or a number.
reference: The expected number, or a text holding it.
None: the experiment item's expected output
(`ctx.expected_output`).
name: Name of the primary score.
Returns:
numeric_diff (NUMERIC, 0 to 1): 1 is equal. 0 when the
reply holds no number.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.numeric_diff(ctx, reference=ctx.params.get("reference")))
json_diff
json_diff(ctx, *, reference=None, name='json_diff')
Score names: json_diff
Score how close the output's JSON is to a reference JSON value.
Objects score the mean over the union of their keys (a key
on one side only scores 0), arrays the mean over the longer
length, strings by edit similarity, numbers as in
`numeric_diff`. Like autoevals JSONDiff.
Args:
ctx: The evaluation context (reads the output), a text,
or a parsed value.
reference: The expected value, or its JSON text. A text
that is not JSON is compared as a string. None: the
experiment item's expected output (`ctx.expected_output`).
name: Name of the primary score.
Returns:
json_diff (NUMERIC, 0 to 1): 1 is the same structure and
values. 0 when the output holds no JSON.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.json_diff(ctx, reference=ctx.params.get("reference")))
bm25_relevance
bm25_relevance(ctx, *, against='input', context=None, context_key='context', k1=1.5, b=0.75, name='bm25_relevance')
Score names: bm25_relevance
Score how relevant the reply is to the question or context, by BM25.
The question's (or context's) content words are the query,
and each sentence of the reply is a document; the best
sentence's BM25 score is divided by the score a sentence of
average length that holds each query word once would get,
and capped at 1. Word frequencies come from the sentences of
both texts. Use it to rank or to flag replies that do not
touch the question; it is not a calibrated probability.
Args:
ctx: The evaluation context, or the reply text (then give
`context`).
against: "input": the user's last message. "context":
the context (see `grounding` for where it comes from).
context: The context text, or a list of passages.
context_key: The span attribute that holds the context
when `context` is not given.
k1: BM25 term saturation.
b: BM25 length normalisation, 0 to 1.
name: Name of the primary score.
Returns:
bm25_relevance (NUMERIC, 0 to 1): higher is more relevant.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.bm25_relevance(ctx))
tfidf_similarity
tfidf_similarity(ctx, *, against='input', context=None, context_key='context', name='tfidf_similarity')
Score names: tfidf_similarity
Score the cosine similarity of the reply and the question or context, by TF-IDF.
Content words only (no stop words). Word weights come from
the sentences of both texts.
Args:
ctx: The evaluation context, or the reply text (then give
`context`).
against: "input": the user's last message. "context":
the context (see `grounding` for where it comes from).
context: The context text, or a list of passages.
context_key: The span attribute that holds the context
when `context` is not given.
name: Name of the primary score.
Returns:
tfidf_similarity (NUMERIC, 0 to 1): 1 is the same words
in the same proportions.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.tfidf_similarity(ctx))
pii_leak
pii_leak(ctx, *, checks=('email', 'phone', 'credit_card', 'us_ssn', 'ip_address', 'secret'), allowed_email_domains=('example.com',), ignore_values_from_input=True, name='pii_free')
Score names: pii_free, pii_count
Check the reply for an email, phone, card number, national id, IP address or credential.
Card numbers are checked with the Luhn sum, so a long order
number does not read as a card. The comment names the kinds
found, never the values, so a score does not copy the leak.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
checks: The kinds to look for: "email", "phone",
"credit_card", "us_ssn", "ip_address", "secret".
allowed_email_domains: Email domains that are fine to
show, such as your support address.
ignore_values_from_input: A value the user typed is not a
leak when the model says it back. False flags every
occurrence.
name: Name of the primary score.
Returns:
pii_free (BOOLEAN): nothing found.
pii_count (NUMERIC): how many values were found.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.pii_leak(ctx))
refusal
refusal(ctx, *, patterns=None, extra_patterns=(), scan_first_chars=400, name='refused')
Score names: refused
Check whether the model declined to do what it was asked.
A refusal is sometimes correct, so this scores the event.
Pair it with a filter on the spans where a refusal is a bug.
Patterns are matched without regard to case.
Args:
ctx: The evaluation context (reads the reply), or the
text to check.
patterns: Patterns that mark a refusal. None uses the
built-in English list.
extra_patterns: Patterns to add to `patterns`.
scan_first_chars: Look only at the start of the reply. A
long answer that says "I can't guarantee this" in
paragraph four is not a refusal.
name: Name of the primary score.
Returns:
refused (BOOLEAN): a pattern matched. Lower is better.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.refusal(ctx))
url_check
url_check(ctx, *, allowed_schemes=('http', 'https'), allowed_domains=(), blocked_domains=(), allow_private=False, require_url=False, field='output', name='urls_ok')
Score names: urls_ok, url_count
Check the URLs in the text: well formed, allowed scheme and domain, not private.
A URL is anything of the form scheme://... . A domain rule
also covers its subdomains. An IP that is not public
(private, loopback, link-local, reserved) and localhost fail
unless `allow_private` is set.
Args:
ctx: The evaluation context, or the text to check.
allowed_schemes: Schemes allowed, such as "https".
allowed_domains: Domains allowed. Empty: any domain.
blocked_domains: Domains that fail the check.
allow_private: Allow private and local addresses.
require_url: Fail when the text has no URL.
field: "output" (the assistant's reply), "input" (the
user's messages) or "both".
name: Name of the primary score.
Returns:
urls_ok (BOOLEAN): every URL passed.
url_count (NUMERIC): how many URLs the text has.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.url_check(ctx))
prompt_injection
prompt_injection(ctx, *, field='input', extra_patterns=(), min_base64_chars=120, threshold=0.5, name='injection_detected')
Score names: injection_detected, injection_risk, injection_signals
Flag text that tries to override the model's instructions.
Signals, each with a weight: "ignore previous instructions"
phrasing (0.6), asking for the system prompt (0.4), role
markers such as "system:" or <|im_start|> (0.4), persona
switches such as "developer mode" (0.3), hidden Unicode tag
characters (0.6), three or more invisible format characters
(0.2), a long base64 blob that decodes (0.3), and each of
your `extra_patterns` (0.5). The risk is the sum of the
signals found, capped at 1. A heuristic: it catches the
common attacks, not a determined one.
Args:
ctx: The evaluation context, or the text to check.
field: "input": the user's and the tools' messages, where
direct and indirect injection arrive. "output": the
reply. "both".
extra_patterns: Python regular expressions to add as
signals.
min_base64_chars: Shortest base64 run that counts.
threshold: Risk at which the text counts as an injection.
name: Name of the primary score.
Returns:
injection_detected (BOOLEAN): risk reached `threshold`.
Lower is better.
injection_risk (NUMERIC, 0 to 1): lower is better.
injection_signals (CATEGORICAL): the signals found,
comma separated, or "none".
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.prompt_injection(ctx))
conversation_degeneration
conversation_degeneration(ctx, *, turns_to_consider=5, min_words=8, threshold=0.6, name='degenerated')
Score names: degenerated, degeneration
Check whether the assistant starts to repeat itself, in the reply or across turns.
Point it at chat spans: their input holds the conversation
so far, and the output is the newest reply. Three signals,
each 0 (healthy) to 1 (degenerate):
- repetition: share of the reply's word 3-grams that occur
more than once in it (a reply that loops on a phrase)
- turn similarity: how close the reply is to the most
similar earlier assistant turn (the same answer again)
- low diversity: 1 minus the word entropy of the reply, over
the most it could be for its length
Args:
ctx: The evaluation context, or the reply text (then only
repetition and low diversity count).
turns_to_consider: Earlier assistant turns to compare
against.
min_words: Replies shorter than this are not compared
across turns: two "Done." replies are not a sign of
degeneration.
threshold: The degeneration score at which the reply
counts as degenerated.
name: Name of the primary score.
Returns:
degenerated (BOOLEAN): degeneration reached `threshold`.
Lower is better.
degeneration (NUMERIC, 0 to 1): the largest of the three
signals. Lower is better. Not returned when the reply
has no text.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.conversation_degeneration(ctx))
knowledge_retention
knowledge_retention(ctx, *, turns_to_consider=5, contradiction_kinds=('email', 'phone', 'name'), name='knowledge_retained')
Score names: knowledge_retained, knowledge_retention
Check that the reply does not forget what the user already said.
Point it at chat spans: their input holds the conversation
so far. It collects facts the user stated in earlier turns
(email, phone, order or ticket number, name), then fails the
reply when it asks for one of those again, or quotes a
different value of the same kind.
Args:
ctx: The evaluation context.
turns_to_consider: Earlier user turns to read facts from.
contradiction_kinds: Kinds where the reply quoting a
different value is a mistake: "email", "phone",
"order_id", "name". Order numbers are left out by
default: an agent that opens a new ticket quotes a
new number, and that is correct.
name: Name of the primary score.
Returns:
knowledge_retained (BOOLEAN): no fact was forgotten.
knowledge_retention (NUMERIC, 0 to 1): the share of facts
kept. Not returned when the user stated no facts.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.knowledge_retention(ctx))
grounding
grounding(ctx, *, context=None, context_key='context', min_overlap=0.6, threshold=0.8, name='grounded')
Score names: grounded, grounding
Score the share of the reply's sentences whose words appear in the context.
A sentence counts as grounded when at least `min_overlap` of
its content words (no stop words, 3 letters or more) appear
in the context. Sentences with no content words are not
counted. A word check, not an entailment model: it catches a
reply that talks about things the context never mentions.
The context is, in order: the `context` argument, the span
attribute `context_key`, else the text of every input
message, where a RAG call usually puts what it retrieved.
Args:
ctx: The evaluation context, or the reply text (then give
`context`).
context: The context text, or a list of passages.
context_key: The span attribute that holds the context.
min_overlap: Share of a sentence's words, 0 to 1, that
must be in the context.
threshold: Share of grounded sentences at which the reply
counts as grounded.
name: Name of the primary score.
Returns:
grounded (BOOLEAN): grounding reached `threshold`.
grounding (NUMERIC, 0 to 1): the share of grounded
sentences.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.grounding(ctx))
unsupported_numbers
unsupported_numbers(ctx, *, context=None, context_key='context', check_entities=True, name='numbers_supported')
Score names: numbers_supported, unsupported_count
Find numbers and names in the reply that the context does not have.
Numbers match by value ("1,200" equals "1200"). One-digit
whole numbers and list positions are ignored. Names are runs
of capitalised words that do not start a sentence, matched in
any case. The context comes from the same places as in
`grounding`.
Args:
ctx: The evaluation context, or the reply text (then give
`context`).
context: The context text, or a list of passages.
context_key: The span attribute that holds the context.
check_entities: Also check names.
name: Name of the primary score.
Returns:
numbers_supported (BOOLEAN): everything found is in the
context.
unsupported_count (NUMERIC): how many are not. Lower is
better.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.unsupported_numbers(ctx))
citation_check
citation_check(ctx, *, sources=None, context=None, context_key='context', require_citation=False, check_urls=True, name='citations_ok')
Score names: citations_ok, citation_count
Check that every citation in the reply points at a real source.
A marker such as [2], [1, 3] or [2-4] must name a source: a
number from 1 to the length of `sources` when given, else a
number the context marks as [n] or as a numbered line. A URL
in the reply must appear in the sources or the context. The
context comes from the same places as in `grounding`.
Args:
ctx: The evaluation context, or the reply text (then give
`context` or `sources`).
sources: The list of sources the reply may cite, in order.
context: The context text, or a list of passages.
context_key: The span attribute that holds the context.
require_citation: Fail a reply with no citation at all.
check_urls: Check the URLs in the reply too.
name: Name of the primary score.
Returns:
citations_ok (BOOLEAN): every citation resolves.
citation_count (NUMERIC): markers and URLs found.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.citation_check(ctx))
tool_call_validity
tool_call_validity(ctx, *, allowed_tools=(), required_args=None, max_identical_calls=2, name='tool_calls_valid')
Score names: tool_calls_valid, tool_call_loop
Check that the tool calls are well formed, allowed and not stuck in a loop.
Checks each call's arguments parse as a JSON object and carry
the arguments you require, that the tool is on the allow
list, and that the model does not repeat a call it already
made with the same arguments earlier in the conversation.
On an experiment item it also reads the agent's tool steps
in ctx.trace.spans.
Args:
ctx: The evaluation context.
allowed_tools: Tool names allowed. Empty: any tool.
required_args: {tool name: [argument names]} each call
of that tool must carry.
max_identical_calls: The same call (name and arguments)
seen this many times counts as a loop.
name: Name of the primary score.
Returns:
tool_calls_valid (BOOLEAN): every call is well formed and
allowed.
tool_call_loop (BOOLEAN): a call repeated. Lower is
better. Not returned when there are no calls.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.tool_call_validity(ctx))
tools_used
tools_used(ctx)
Return the names of the tools the agent called, in order.
Reads the tool steps in ctx.trace.spans. When the trace has
none (a live span), reads the tool calls in the output.
Args:
ctx: The evaluation context.
Returns:
A list of tool names.
tool_trajectory
tool_trajectory(ctx, *, expected=(), mode='strict', name='trajectory_match')
Score names: trajectory_match, tool_correctness
Check the agent's tool calls against the tools you expect.
Like LangChain agentevals trajectory match and DeepEval
ToolCorrectness. Calls are compared by tool name.
Args:
ctx: The evaluation context (see `tools_used`), or a
list of tool names.
expected: The tool names you expect, in order.
mode: "strict": the same calls in the same order.
"unordered": the same calls in any order.
"subset": every call is an expected one (no extra
tool). "superset": every expected call happened
(extra calls allowed). Repeated names count.
name: Name of the primary score.
Returns:
trajectory_match (BOOLEAN): the calls meet `mode`.
tool_correctness (NUMERIC, 0 to 1): the share of expected
calls that happened. 1 when nothing is expected.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.tool_trajectory(ctx))
token_latency_budget
token_latency_budget(ctx, *, max_output_tokens=2000, max_input_tokens=None, max_duration_ms=30000, max_time_to_first_token_ms=5000, name='within_budget')
Score names: within_budget, budget_used
Check that the call stays inside its token, duration and time to first token limits.
Reads the span's own attributes, so it needs no text. A span
that reports none of them passes and says so.
Args:
ctx: The evaluation context, or the metadata dict.
max_output_tokens: None: not checked.
max_input_tokens: None: not checked.
max_duration_ms: None: not checked.
max_time_to_first_token_ms: None: not checked.
name: Name of the primary score.
Returns:
within_budget (BOOLEAN): no limit was passed.
budget_used (NUMERIC): the largest share of any limit;
above 1 is over. Lower is better. Not returned when
the span reports none of the attributes.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.token_latency_budget(ctx))
span_count
span_count(ctx, *, kind=None, max_spans=None, name='span_count')
Score names: span_count, within_span_limit
Count the agent's steps in ctx.trace.spans.
The trace is filled on an experiment item and empty on a live
span, so on a live span the count is 0.
Args:
ctx: The evaluation context, or a list of span dicts.
kind: Count only steps of this kind: "llm", "tool",
"agent" or "span". None: every step.
max_spans: Most steps allowed. None: not checked.
name: Name of the primary score.
Returns:
span_count (NUMERIC): the count. Lower is better.
within_span_limit (BOOLEAN): the count is at most
`max_spans`. Only when `max_spans` is set.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.span_count(ctx))
error_spans
error_spans(ctx, *, max_errors=0, name='no_errors')
Score names: no_errors, error_count
Count the agent's steps that failed, in ctx.trace.spans.
A step failed when its `error` field is set, or its `status`
is "error".
Args:
ctx: The evaluation context, or a list of span dicts.
max_errors: Most failed steps allowed.
name: Name of the primary score.
Returns:
no_errors (BOOLEAN): at most `max_errors` steps failed.
error_count (NUMERIC): failed steps. Lower is better.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.error_spans(ctx))
repeated_tool_calls
repeated_tool_calls(ctx, *, max_repeats=2, consecutive_only=False, name='tool_loop')
Score names: tool_loop, most_repeated
Detect the agent calling the same tool with the same input again and again.
Reads the tool steps in ctx.trace.spans; with none (a live
span), the tool calls in the output.
Args:
ctx: The evaluation context, or a list of span dicts.
max_repeats: The same call seen this many times counts as
a loop.
consecutive_only: Count only calls repeated back to back.
name: Name of the primary score.
Returns:
tool_loop (BOOLEAN): a call repeated `max_repeats` times.
Lower is better.
most_repeated (NUMERIC): the most times one call was
made. Lower is better.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.repeated_tool_calls(ctx))
total_duration
total_duration(ctx, *, max_duration_ms=None, name='total_duration_ms')
Score names: total_duration_ms, within_duration
Measure how long the agent ran, from its trace.
With agent steps in ctx.trace.spans, the longest agent step;
else the sum of every step's duration_ms (steps can overlap,
so the sum is an upper bound); with no steps (a live span),
the span's own duration_ms attribute.
Args:
ctx: The evaluation context, or a list of span dicts.
max_duration_ms: Longest time allowed. None: not checked.
name: Name of the primary score.
Returns:
total_duration_ms (NUMERIC): lower is better.
within_duration (BOOLEAN): at most `max_duration_ms`.
Only when it is set.
Example:
from oodle_eval.v1 import metrics
def evaluate(ctx):
return EvaluationResult(scores=metrics.total_duration(ctx))
text
Import path: oodle_eval.v1.text
Read text, turns and tool calls out of a span, whatever its shape.
A span's input and output arrive as a plain string, one message
dict, or a list of messages. A message holds its text as a
`content` string, as a list of content blocks (OpenAI and
Anthropic), or as OpenTelemetry gen_ai `parts`. The helpers here
read all of these, so a check does not care which SDK wrote the
span.
Every function takes either the evaluation context `ctx` or the
value itself, where the summary says so.
| Function | Summary |
|---|---|
message_text | Return the plain text of a message value. |
reply | Return the text of the assistant's reply in the output. |
user_text | Return the text of every user message in the input. |
turns | Return the conversation as a list of turns. |
last_user_message | Return the text of the last user turn in the input. |
tool_calls | Return every tool call in a value, in order. |
words | Return the words of a text. |
sentences | Return the sentences of a text. |
normalize | Return a text in lower case, with no punctuation and single spaces. |
extract_json | Return the first JSON value found in a text. |
metadata_value | Return the first span attribute found among names. |
metadata_number | Return the first span attribute among names that is a number. |
message_text
message_text(value, role=None, default_role=None)
Return the plain text of a message value.
Args:
value: A string, a message dict, or a list of messages
and strings. Any other JSON value is returned as its
JSON text.
role: Keep only messages of this role ("user",
"assistant", "system", "tool"). None keeps every
message.
default_role: The role of a message that has no `role`
key. None means `role`, so a bare message counts as
the role asked for.
Returns:
The text of the kept messages, joined by new lines. A
string value is returned as it is, whatever `role` is.
reply
reply(ctx)
Return the text of the assistant's reply in the output.
Args:
ctx: The evaluation context, or the output value itself.
Returns:
The assistant messages' text. A message with no role
counts as the assistant's.
user_text
user_text(ctx)
Return the text of every user message in the input.
Args:
ctx: The evaluation context, or the input value itself.
Returns:
The user messages' text. A message with no role counts
as the user's. A plain string input is returned as it is.
turns
turns(ctx, include_reply=True)
Return the conversation as a list of turns.
The shape is the thread shape DeepEval, Ragas and Opik use.
Args:
ctx: The evaluation context.
include_reply: Add the output's messages at the end.
Returns:
A list of {"role": str, "content": str}. An input
message with no role is a user turn; an output message
with no role is an assistant turn. A plain string input
is one user turn. Turns with no text are left out.
last_user_message
last_user_message(ctx)
Return the text of the last user turn in the input.
Args:
ctx: The evaluation context.
Returns:
The text, or "" when the input has no user turn.
tool_calls
tool_calls(ctx_or_value)
Return every tool call in a value, in order.
Reads OpenAI chat `tool_calls`, OpenAI Responses
`function_call` items, OpenTelemetry gen_ai parts of type
`tool_call`, and Anthropic `tool_use` content blocks.
Args:
ctx_or_value: The evaluation context (reads the output)
or a message value.
Returns:
A list of {"id", "name", "arguments"}. `arguments` is
parsed when it is a JSON string, else the value as it
came (a string that is not JSON stays a string). `name`
is "" when the call has none.
words
words(text, lower=True)
Return the words of a text.
Args:
text: The text.
lower: Lower-case the words first.
Returns:
A list of runs of letters, digits and underscores.
sentences
sentences(text)
Return the sentences of a text.
A full stop, question mark or exclamation mark followed by a
space, or a blank line, ends a sentence.
Args:
text: The text.
Returns:
A list of sentences with the space around them removed.
normalize
normalize(text)
Return a text in lower case, with no punctuation and single spaces.
Use it before an exact comparison, so "Paris." equals "paris".
Args:
text: The text.
Returns:
The normalised text.
extract_json
extract_json(text)
Return the first JSON value found in a text.
Tries, in order: the whole text, the first fenced code
block (```json or ```), then the first balanced {...} or
[...] that parses.
Args:
text: The text. A dict or list is returned as it is.
Returns:
The parsed value, or None when no JSON is found.
metadata_value
metadata_value(ctx, *names, default=None)
Return the first span attribute found among `names`.
A numeric attribute can be stored with a type suffix, so
each name is also tried with `_int`, `_float`, `_double`,
`_bool`, `_str` and `_string`.
Args:
ctx: The evaluation context, or the metadata dict.
*names: Attribute names to try, first found wins.
default: Returned when no name is found.
Returns:
The attribute value as stored, or `default`.
metadata_number
metadata_number(ctx, *names, default=None)
Return the first span attribute among `names` that is a number.
Like `metadata_value`, but skips a value that does not read
as a number, and returns a float.
Args:
ctx: The evaluation context, or the metadata dict.
*names: Attribute names to try, first found wins.
default: Returned when no name holds a number.
Returns:
The number as a float, or `default`.
combine
Import path: oodle_eval.v1.combine
Combine several scores into one.
Every function takes a list of scores: a `Scores` from a check,
a plain list, or a list of such lists, which is read as one flat
list. A score can also be a dict with "name", "value",
"data_type" and "higher_is_better", as `ctx.scores` holds.
scores = metrics.tone(ctx) + metrics.readability(ctx)
overall = combine.weighted_mean(scores, weights={"tone_ok": 2})
return EvaluationResult(scores=scores + [overall])
The models for these are OpenAI's multi grader, promptfoo's
weights and assert sets, DeepEval's composite metrics, Opik's
AggregatedMetric and LangSmith's composite scores.
| Function | Summary |
|---|---|
weighted_mean | Return the weighted mean of the scores' values as one NUMERIC score. |
passed | Return whether one score reads as a pass. |
all_pass | Return True when every score passes, as a BOOLEAN score. |
any_fail | Return True when a score fails, as a BOOLEAN score. |
minimum | Return the smallest value among the scores, as a NUMERIC score. |
maximum | Return the largest value among the scores, as a NUMERIC score. |
threshold | Turn one score into a BOOLEAN pass at a line. |
rename | Return copies of the scores with new names. |
weighted_mean
weighted_mean(scores, weights=None, name='combined')
Return the weighted mean of the scores' values as one NUMERIC score.
A BOOLEAN score counts as 1 or 0. CATEGORICAL scores and
scores with no value are skipped. Values are used as they
are: pick scores that read the same way (turn a
lower-is-better score round with `threshold` or `passed`
first).
Args:
scores: The scores to combine.
weights: {score name: weight}, or a list of weights in
the order of `scores`. A name missing from the dict
weighs 1. None: every score weighs 1.
name: Name of the combined score.
Returns:
A NUMERIC Score, higher is better.
Raises:
ValueError: no score has a number to combine.
passed
passed(score, at=None)
Return whether one score reads as a pass.
A BOOLEAN score passes when it is True, or when it is False
for a lower-is-better score (such as `refused`). A NUMERIC
score passes when it is at or above `at`, or at or below
`at` for a lower-is-better score. With no `at`, 0.5 is the
line: at or above for higher-is-better, below for
lower-is-better.
Args:
score: A Score or a score dict.
at: The pass line for a NUMERIC score.
Returns:
True or False.
Raises:
ValueError: the score is CATEGORICAL or has no value.
all_pass
all_pass(scores, name='all_pass', at=None)
Return True when every score passes, as a BOOLEAN score.
Each score is read with `passed`. CATEGORICAL scores and
scores with no value are skipped.
Args:
scores: The scores to check.
at: The pass line for NUMERIC scores, as in `passed`.
name: Name of the result score.
Returns:
A BOOLEAN Score, higher is better. True when there is
nothing to check.
any_fail
any_fail(scores, name='any_fail', at=None)
Return True when a score fails, as a BOOLEAN score.
The opposite of `all_pass`.
Args:
scores: The scores to check.
at: The pass line for NUMERIC scores, as in `passed`.
name: Name of the result score.
Returns:
A BOOLEAN Score, lower is better.
minimum
minimum(scores, name='minimum')
Return the smallest value among the scores, as a NUMERIC score.
A BOOLEAN score counts as 1 or 0; CATEGORICAL scores and
scores with no value are skipped.
Args:
scores: The scores to read.
name: Name of the result score.
Returns:
A NUMERIC Score. Its direction is the one the scores
share, else unset.
Raises:
ValueError: no score has a number.
maximum
maximum(scores, name='maximum')
Return the largest value among the scores, as a NUMERIC score.
A BOOLEAN score counts as 1 or 0; CATEGORICAL scores and
scores with no value are skipped.
Args:
scores: The scores to read.
name: Name of the result score.
Returns:
A NUMERIC Score. Its direction is the one the scores
share, else unset.
Raises:
ValueError: no score has a number.
threshold
threshold(score, at, name=None, above=True)
Turn one score into a BOOLEAN pass at a line.
Args:
score: A Score or a score dict. A BOOLEAN counts as 1
or 0.
at: The line.
above: True: pass at or above `at`. False: pass below
`at`.
name: Name of the result score. None: the score's name
with "_ok" added.
Returns:
A BOOLEAN Score, higher is better. False when the score
has no number.
rename
rename(scores, prefix='', names=None)
Return copies of the scores with new names.
Use it when two checks return scores of the same name, or to
tell apart the scores of one check run twice.
Args:
scores: The scores to rename.
prefix: Text to put before every name.
names: {old name: new name}, applied before `prefix`.
Returns:
A new Scores. The scores passed in do not change.
util
Import path: oodle_eval.v1.util
Building blocks for writing your own checks, the way the built-in ones are written.
from oodle_eval.v1 import text, util
@util.declares("polite", "politeness")
def politeness(ctx, *, words=None, name="polite"):
words = util.opt(words, ["please", "thank"])
reply = util.subject_text(ctx).lower()
found = [w for w in words if w in reply]
return Scores([
util.boolean_score(util.opt(name, "polite"), bool(found), f"found {found}"),
util.numeric_score("politeness", len(found) / len(words), "share of words"),
])
A check built this way takes a context or a plain string, reads
None as "the default", and returns `Scores` whose names the
editor knows before the first run. Message helpers (reply,
turns, tool calls, JSON) are in `oodle_eval.v1.text`.
The sandbox gives each item 5 seconds: cut long text with
`MAX_TEXT_CHARS` and `MAX_SENTENCES` before anything that grows
faster than the text.
| Function | Summary |
|---|---|
declares | Record the score names a check returns, primary first. |
opt | Return value, or default when value is None. |
is_context | Return whether value is an evaluation context. |
subject_text | Return the text a check reads: the reply, the user's turns, or both. |
reference_or_expected | Return the reference setting, else the experiment item's expected output. |
require_reference | Return a reference as text, or say how to set one. |
boolean_score | Return a BOOLEAN Score. |
numeric_score | Return a NUMERIC Score. |
strip_code | Return a text with code blocks, inline code and URLs made spaces. |
fenced_block | Return the first fenced code block of one of langs, else the whole text. |
join_problems | Return problems as one comment, the first limit named. |
flat_text | Return any value as text: a string as it is, a list line by line, else JSON. |
context_text | Return the retrieved context to check a reply against. |
content_words | Return the words of a text that carry content. |
find_number | Return the first number in a value, as a float. |
find_urls | Return the URLs in a text, with trailing punctuation removed. |
is_messages | Return whether a value is a list of chat messages. |
declares
declares(*names)
Record the score names a check returns, primary first.
The library reference reads them from the function, so the
editor can offer a score name before the check has run once.
Args:
*names: The score names, primary first.
Returns:
A decorator that sets `__scores__` on the function and
returns the function unchanged.
opt
opt(value, default)
Return `value`, or `default` when `value` is None.
An evaluator's unset setting arrives as None, so a check
reads each setting through this to get its own default.
Args:
value: The setting as passed.
default: The value to use for None.
Returns:
`value`, or `default`.
is_context
is_context(value)
Return whether `value` is an evaluation context.
Args:
value: Anything.
Returns:
True for an object with `observation` and `trace`.
subject_text
subject_text(ctx, field='output', limit=None)
Return the text a check reads: the reply, the user's turns, or both.
Args:
ctx: The evaluation context, a plain string (returned as
it is), or a message value (its assistant text).
field: "output" (the assistant's reply), "input" (the
user's messages) or "both". Used for a context only.
limit: The most characters to return. None: all.
Returns:
The text.
Raises:
ValueError: `field` is not one of the three.
reference_or_expected
reference_or_expected(ctx, reference)
Return the reference setting, else the experiment item's expected output.
Args:
ctx: The evaluation context, or anything else.
reference: The reference setting.
Returns:
`reference` when set; else `ctx.expected_output` for a
context; else None.
require_reference
require_reference(reference, check)
Return a reference as text, or say how to set one.
Args:
reference: The reference, a string or any JSON value.
check: The check's name, for the error message.
Returns:
The reference itself when it is a string, else its JSON
text.
Raises:
ValueError: `reference` is None.
boolean_score
boolean_score(name, value, comment, higher_is_better=True)
Return a BOOLEAN Score.
Args:
name: The score name.
value: Read as True or False.
comment: A short reason.
higher_is_better: False for a score that flags a problem
when true, such as `refused`.
Returns:
A Score.
numeric_score
numeric_score(name, value, comment, higher_is_better=True)
Return a NUMERIC Score.
Args:
name: The score name.
value: A number.
comment: A short reason.
higher_is_better: False for a count of problems or a
duration.
Returns:
A Score.
strip_code
strip_code(value)
Return a text with code blocks, inline code and URLs made spaces.
Use it before a check of prose, so code does not read as
words.
Args:
value: The text.
Returns:
The text without code and URLs.
fenced_block
fenced_block(value, *langs)
Return the first fenced code block of one of `langs`, else the whole text.
Args:
value: The text.
*langs: Fence labels to accept, such as "json". A block
with no label is also accepted.
Returns:
The block's content, or `value` when there is none.
join_problems
join_problems(problems, limit=10)
Return problems as one comment, the first `limit` named.
Args:
problems: A list of short texts.
limit: The most to name; the rest are counted.
Returns:
"a; b; and 3 more", or "" for no problems.
flat_text
flat_text(value)
Return any value as text: a string as it is, a list line by line, else JSON.
Args:
value: A string, a list, or any JSON value.
Returns:
The text. "" for None.
context_text
context_text(ctx, context=None, context_key='context')
Return the retrieved context to check a reply against.
In order: the `context` argument, else the span attribute
`context_key`, else the text of every input message (a RAG
call usually puts what it retrieved in its prompt). Cut to
`MAX_TEXT_CHARS`.
Args:
ctx: The evaluation context, or a plain string (then only
`context` counts).
context: The context text, or a list of passages.
context_key: The span attribute that holds the context.
Returns:
The context text, or "".
content_words
content_words(value)
Return the words of a text that carry content.
Args:
value: The text.
Returns:
Lower-case words of 3 or more characters that are not in
`STOPWORDS`.
find_number
find_number(value)
Return the first number in a value, as a float.
Args:
value: A number (a bool is not one), or a text holding
one. "1,200" reads as 1200.
Returns:
The number, or None.
find_urls
find_urls(value, limit=None)
Return the URLs in a text, with trailing punctuation removed.
Args:
value: The text.
limit: The most URLs to return. None: all.
Returns:
A list of URLs, in order.
is_messages
is_messages(value)
Return whether a value is a list of chat messages.
Args:
value: Anything.
Returns:
True for a list with a dict that has "role" or "parts".
metrics.format
Import path: oodle_eval.v1.metrics.format
The metrics functions defined in oodle_eval/v1/metrics/format.py. Import them from oodle_eval.v1.metrics.
| Function | Summary |
|---|---|
json_validity | Check that the output is JSON and has the shape you expect. |
regex_match | Check that the text matches every required pattern and no forbidden one. |
keyword_check | Check that the text mentions the required words and none of the forbidden ones. |
contains | Check that the text contains the given substrings. |
length_budget | Check that the reply stays inside a word and character range. |
markdown_structure | Check the reply's Markdown: code fences, tables, headings and lists. |
html_validity | Check that the HTML in the reply has balanced tags. |
csv_validity | Check that the reply is CSV with the same number of fields in every row. |
toml_validity | Check that the reply parses as TOML and has the keys you require. |
sql_shape | Check that generated SQL is only the kind of statement you allow. |
metrics.quality
Import path: oodle_eval.v1.metrics.quality
The metrics functions defined in oodle_eval/v1/metrics/quality.py. Import them from oodle_eval.v1.metrics.
| Function | Summary |
|---|---|
readability | Score Flesch reading ease and check the Flesch-Kincaid grade. |
detect_language | Return the language of a text and a confidence. |
language_adherence | Check that the reply is written in the language you expect. |
tone | Flag a reply that shouts, piles on exclamation marks, reads as negative, or uses a banned phrase. |
sentiment | Score the reply's sentiment from -1 (negative) to 1 (positive). |
metrics.similarity
Import path: oodle_eval.v1.metrics.similarity
The metrics functions defined in oodle_eval/v1/metrics/similarity.py. Import them from oodle_eval.v1.metrics.
| Function | Summary |
|---|---|
exact_match | Check that the reply equals a reference text. |
levenshtein_ratio | Score how close the reply is to a reference, by edit distance. |
jaccard_similarity | Score the overlap of the reply's and the reference's word sets. |
rouge | Score ROUGE-L, ROUGE-1 and ROUGE-2 F1 against a reference. |
bleu | Score sentence BLEU against a reference. |
numeric_diff | Score how close the first number in the reply is to a reference number. |
json_diff | Score how close the output's JSON is to a reference JSON value. |
bm25_relevance | Score how relevant the reply is to the question or context, by BM25. |
tfidf_similarity | Score the cosine similarity of the reply and the question or context, by TF-IDF. |
metrics.safety
Import path: oodle_eval.v1.metrics.safety
The metrics functions defined in oodle_eval/v1/metrics/safety.py. Import them from oodle_eval.v1.metrics.
| Function | Summary |
|---|---|
pii_leak | Check the reply for an email, phone, card number, national id, IP address or credential. |
refusal | Check whether the model declined to do what it was asked. |
url_check | Check the URLs in the text: well formed, allowed scheme and domain, not private. |
prompt_injection | Flag text that tries to override the model's instructions. |
metrics.conversation
Import path: oodle_eval.v1.metrics.conversation
The metrics functions defined in oodle_eval/v1/metrics/conversation.py. Import them from oodle_eval.v1.metrics.
| Function | Summary |
|---|---|
conversation_degeneration | Check whether the assistant starts to repeat itself, in the reply or across turns. |
knowledge_retention | Check that the reply does not forget what the user already said. |
metrics.context
Import path: oodle_eval.v1.metrics.context
The metrics functions defined in oodle_eval/v1/metrics/context.py. Import them from oodle_eval.v1.metrics.
| Function | Summary |
|---|---|
grounding | Score the share of the reply's sentences whose words appear in the context. |
unsupported_numbers | Find numbers and names in the reply that the context does not have. |
citation_check | Check that every citation in the reply points at a real source. |
metrics.agent
Import path: oodle_eval.v1.metrics.agent
The metrics functions defined in oodle_eval/v1/metrics/agent.py. Import them from oodle_eval.v1.metrics.
| Function | Summary |
|---|---|
tool_call_validity | Check that the tool calls are well formed, allowed and not stuck in a loop. |
tools_used | Return the names of the tools the agent called, in order. |
tool_trajectory | Check the agent's tool calls against the tools you expect. |
token_latency_budget | Check that the call stays inside its token, duration and time to first token limits. |
metrics.trace
Import path: oodle_eval.v1.metrics.trace
The metrics functions defined in oodle_eval/v1/metrics/trace.py. Import them from oodle_eval.v1.metrics.
| Function | Summary |
|---|---|
span_count | Count the agent's steps in ctx.trace.spans. |
error_spans | Count the agent's steps that failed, in ctx.trace.spans. |
repeated_tool_calls | Detect the agent calling the same tool with the same input again and again. |
total_duration | Measure how long the agent ran, from its trace. |
Support
If you need assistance or have any questions, please reach out to us through:
- Email at [email protected]