Skip to main content

Coming from Another Evaluation Tool

If you wrote evaluators for Opik, Braintrust, LangSmith or Langfuse, most of what you know carries over. This page maps their concepts to Oodle's.

Concepts in common​

ConceptIn Oodle
The function that scoresevaluate(ctx) in a code template
What it getsctx: the span's input, output and attributes, the agent's steps in an experiment, the evaluator's settings, and other evaluators' scores. See the ctx object
What it returnsEvaluationResult(scores=[Score(...)]), one or more named scores
Score typesNUMERIC, BOOLEAN, CATEGORICAL, and higher_is_better on each score
Pre-built metricsBuilt-in code checks, and the same checks as functions in the oodle_eval library
Parameters of a scorerTemplate settings, with values per evaluator
Code that many scorers shareShared libraries
Combined or composite scorescombine in the library, or an evaluator that reads other evaluators' scores
Online evaluationAn evaluator (rule) with filters, sampling and an hourly cap
Offline evaluationExperiments on a dataset
Expected outputctx.expected_output in an experiment. The reference checks use it when you set no reference. On a live span, pass a reference as a setting

The sandbox runs Python with the standard library modules that are useful for scoring and the oodle_eval library. Each span gets 5 seconds and 128 MB, and the code has no network access. Rewrite a scorer that calls a model as an LLM-as-Judge evaluator, and combine its score with code checks by reading its scores.

Opik​

OpikOodle
BaseMetric subclass with score(...)evaluate(ctx)
score_result.ScoreResult(name, value, reason)Score(name=..., value=..., comment=...)
A list of ScoreResult from one metricA list of Score in one EvaluationResult
Equals, Contains, RegexMatch, IsJsonmetrics.exact_match, metrics.contains, metrics.regex_match, metrics.json_validity
LevenshteinRatio, SentenceBLEU, ROUGE, Sentimentmetrics.levenshtein_ratio, metrics.bleu, metrics.rouge, metrics.sentiment
ConversationDegenerationMetric, KnowledgeRetentionMetricmetrics.conversation_degeneration, metrics.knowledge_retention
AggregatedMetric(metrics=[...], aggregator=...)combine.weighted_mean(scores), or your own arithmetic on the list
Online evaluation ruleEvaluator with filters and a sampling rate

A metric that you constructed with arguments, such as Contains(case_sensitive=True), is a call with keyword arguments: metrics.contains(ctx, values=[...], case_sensitive=True). The cookbook shows an Opik custom metric and the same evaluator in Oodle side by side.

Braintrust​

BraintrustOodle
Code scorer handler(input, output, expected, metadata)evaluate(ctx). Read ctx.observation.input, ctx.observation.output and ctx.observation.metadata
A number, or {"name", "score", "metadata"}Score(name=..., value=..., comment=...) in an EvaluationResult
autoevals ExactMatch, Levenshtein, NumericDiff, JSONDiffmetrics.exact_match, metrics.levenshtein_ratio, metrics.numeric_diff, metrics.json_diff
Scorer parametersTemplate settings
project.scorers.create(...) and pushing functions from the CLIA code template, created in the UI, with oodle genai templates create, or through the API
Online scoring ruleEvaluator with filters and a sampling rate

Where a Braintrust scorer reads expected, read ctx.expected_output in an experiment. On live spans, give the reference as a setting (reference=).

LangSmith​

LangSmithOodle
perform_eval(run, example)evaluate(ctx)
run["inputs"], run["outputs"]ctx.observation.input, ctx.observation.output
example["outputs"] (reference outputs)ctx.expected_output in an experiment, or a reference setting
{"metric_name": score} with several keysSeveral named Score objects
Composite evaluator (weighted average or sum of feedback keys)An evaluator that reads other evaluators' scores and uses combine.weighted_mean, or arithmetic of your own
Online evaluation ruleEvaluator with filters and a sampling rate

A LangSmith composite score has no value when one of its inputs is missing on a run. In Oodle, an evaluator that reads scores runs only on the spans that every input evaluator scored, which gives the same result.

Langfuse​

Oodle's code evaluator contract follows Langfuse's: evaluate(ctx) returns an EvaluationResult of Score objects with name, value, data_type and comment. Most Langfuse Python code evaluators run in Oodle with few changes.

LangfuseOodle
ctx.observation.input, .output, .metadataThe same
ctx.observation.tool_callstext.tool_calls(ctx), which gives {"id", "name", "arguments"} dicts
ctx.experiment.item_expected_outputctx.experiment.expected_output, or the shortcut ctx.expected_output
ctx.experiment.item_metadatactx.experiment.item_metadata
Score(name, value, data_type, comment)The same, plus higher_is_better
One evaluator's scores in anotherctx.scores and ctx.score(...)

If your evaluator defines its own Score or EvaluationResult dataclasses so that it runs outside Langfuse, keep them: Oodle reads the fields that it needs from your classes.

Move an evaluator​

  1. Find the checks that the library already has, and use them. Most heuristic metrics are in metrics.
  2. Move each constant that differs between uses (a keyword list, a threshold) into a setting.
  3. Move helper code that several scorers share into a shared library.
  4. Run the template on sample spans in the editor, or with try_code in MCP, before you create evaluators.

Support

If you need assistance or have any questions, please reach out to us through: