Coming from Another Evaluation Tool
If you wrote evaluators for Opik, Braintrust, LangSmith or Langfuse, most of what you know carries over. This page maps their concepts to Oodle's.
Concepts in common
| Concept | In Oodle |
|---|---|
| The function that scores | evaluate(ctx) in a code template |
| What it gets | ctx: the span's input, output and attributes, the agent's steps in an experiment, the evaluator's settings, and other evaluators' scores. See the ctx object |
| What it returns | EvaluationResult(scores=[Score(...)]), one or more named scores |
| Score types | NUMERIC, BOOLEAN, CATEGORICAL, and higher_is_better on each score |
| Pre-built metrics | Built-in code checks, and the same checks as functions in the oodle_eval library |
| Parameters of a scorer | Template settings, with values per evaluator |
| Code that many scorers share | Shared libraries |
| Combined or composite scores | combine in the library, or an evaluator that reads other evaluators' scores |
| Online evaluation | An evaluator (rule) with filters, sampling and an hourly cap |
| Offline evaluation | Experiments on a dataset |
| Expected output | ctx.expected_output in an experiment. The reference checks use it when you set no reference. On a live span, pass a reference as a setting |
The sandbox runs Python with the standard library
modules that are useful for scoring and the
oodle_eval library. Each span gets 5 seconds and
128 MB, and the code has no network access. Rewrite a
scorer that calls a model as an
LLM-as-Judge evaluator, and combine
its score with code checks by
reading its scores.
Opik
| Opik | Oodle |
|---|---|
BaseMetric subclass with score(...) | evaluate(ctx) |
score_result.ScoreResult(name, value, reason) | Score(name=..., value=..., comment=...) |
A list of ScoreResult from one metric | A list of Score in one EvaluationResult |
Equals, Contains, RegexMatch, IsJson | metrics.exact_match, metrics.contains, metrics.regex_match, metrics.json_validity |
LevenshteinRatio, SentenceBLEU, ROUGE, Sentiment | metrics.levenshtein_ratio, metrics.bleu, metrics.rouge, metrics.sentiment |
ConversationDegenerationMetric, KnowledgeRetentionMetric | metrics.conversation_degeneration, metrics.knowledge_retention |
AggregatedMetric(metrics=[...], aggregator=...) | combine.weighted_mean(scores), or your own arithmetic on the list |
| Online evaluation rule | Evaluator with filters and a sampling rate |
A metric that you constructed with arguments, such as
Contains(case_sensitive=True), is a call with keyword
arguments: metrics.contains(ctx, values=[...], case_sensitive=True).
The cookbook
shows an Opik custom metric and the same evaluator in
Oodle side by side.
Braintrust
| Braintrust | Oodle |
|---|---|
Code scorer handler(input, output, expected, metadata) | evaluate(ctx). Read ctx.observation.input, ctx.observation.output and ctx.observation.metadata |
A number, or {"name", "score", "metadata"} | Score(name=..., value=..., comment=...) in an EvaluationResult |
autoevals ExactMatch, Levenshtein, NumericDiff, JSONDiff | metrics.exact_match, metrics.levenshtein_ratio, metrics.numeric_diff, metrics.json_diff |
Scorer parameters | Template settings |
project.scorers.create(...) and pushing functions from the CLI | A code template, created in the UI, with oodle genai templates create, or through the API |
| Online scoring rule | Evaluator with filters and a sampling rate |
Where a Braintrust scorer reads expected, read
ctx.expected_output in an experiment. On live spans,
give the reference as a setting (reference=).
LangSmith
| LangSmith | Oodle |
|---|---|
perform_eval(run, example) | evaluate(ctx) |
run["inputs"], run["outputs"] | ctx.observation.input, ctx.observation.output |
example["outputs"] (reference outputs) | ctx.expected_output in an experiment, or a reference setting |
{"metric_name": score} with several keys | Several named Score objects |
| Composite evaluator (weighted average or sum of feedback keys) | An evaluator that reads other evaluators' scores and uses combine.weighted_mean, or arithmetic of your own |
| Online evaluation rule | Evaluator with filters and a sampling rate |
A LangSmith composite score has no value when one of its inputs is missing on a run. In Oodle, an evaluator that reads scores runs only on the spans that every input evaluator scored, which gives the same result.
Langfuse
Oodle's code evaluator contract follows Langfuse's:
evaluate(ctx) returns an EvaluationResult of Score
objects with name, value, data_type and
comment. Most Langfuse Python code evaluators run in
Oodle with few changes.
| Langfuse | Oodle |
|---|---|
ctx.observation.input, .output, .metadata | The same |
ctx.observation.tool_calls | text.tool_calls(ctx), which gives {"id", "name", "arguments"} dicts |
ctx.experiment.item_expected_output | ctx.experiment.expected_output, or the shortcut ctx.expected_output |
ctx.experiment.item_metadata | ctx.experiment.item_metadata |
Score(name, value, data_type, comment) | The same, plus higher_is_better |
| One evaluator's scores in another | ctx.scores and ctx.score(...) |
If your evaluator defines its own Score or
EvaluationResult dataclasses so that it runs outside
Langfuse, keep them: Oodle reads the fields that it
needs from your classes.
Move an evaluator
- Find the checks that the library already has, and
use them. Most heuristic metrics are in
metrics. - Move each constant that differs between uses (a keyword list, a threshold) into a setting.
- Move helper code that several scorers share into a shared library.
- Run the template on sample spans in the editor, or
with
try_codein MCP, before you create evaluators.
Support
If you need assistance or have any questions, please reach out to us through:
- Email at [email protected]