Built-in Code Checks
Built-in code checks are code evaluators that Oodle writes and maintains. You write no code: pick a check, set its settings, and choose the spans it scores. Like all code evaluators, they are fast and make no LLM call.
Use a built-in check
- Go to Agent Observability → Evaluators (ap1, us1)
- Click Create Evaluator and pick a check from Built-in checks
- Set the check's settings, for example the keywords that the reply must mention
- Set the filters, the sampling rate and the hourly cap, as for any evaluator
- Click Save
Each check is a managed template with the id
oodle-managed-code-<check>-v1. Use that id as the
template when you create an evaluator through the API
(evaluatorId), the CLI or MCP (templateId):
{"action": "create_evaluator", "name": "Support replies mention refunds",
"templateId": "oodle-managed-code-keyword-check-v1",
"params": {"required": ["refund"], "mode": "any"}, "enabled": false}
The settings of each check are the keyword arguments of its library function. The library reference gives each argument and its default. A setting that you do not set takes that default.
The checks
The first score of each check is its primary score.
Format
| Check | Template id | What it scores | Scores | Function |
|---|---|---|---|---|
| JSON validity | oodle-managed-code-json-validity-v1 | Output parses as JSON and has the keys, value types and schema you set | is_json, json_schema_ok | json_validity |
| Regex match | oodle-managed-code-regex-match-v1 | Output matches every required pattern and none of the forbidden ones | regex_pass | regex_match |
| Keyword check | oodle-managed-code-keyword-check-v1 | Output mentions the required keywords and none of the forbidden ones | keywords_ok, keyword_coverage | keyword_check |
| Contains | oodle-managed-code-contains-v1 | Output contains the values you list, anywhere in the text | contains | contains |
| Length budget | oodle-managed-code-length-budget-v1 | Output stays inside a word and character range | within_length | length_budget |
| Markdown structure | oodle-managed-code-markdown-structure-v1 | Code fences are closed, tables are even, and the headings, lists and tables you require are there | markdown_ok, heading_count | markdown_structure |
| HTML validity | oodle-managed-code-html-validity-v1 | The HTML in the output has balanced tags, and only the tags you allow | html_valid | html_validity |
| CSV validity | oodle-managed-code-csv-validity-v1 | Output is CSV with the same number of fields in every row, and the columns you require | csv_valid, row_count | csv_validity |
| TOML validity | oodle-managed-code-toml-validity-v1 | Output parses as TOML and has the keys you require | toml_valid | toml_validity |
| SQL shape | oodle-managed-code-sql-shape-v1 | Generated SQL is only the statement types you allow, with no write keyword you did not allow | sql_ok, statement_type | sql_shape |
Text quality
| Check | Template id | What it scores | Scores | Function |
|---|---|---|---|---|
| Readability | oodle-managed-code-readability-v1 | Flesch reading ease and Flesch-Kincaid grade, with a grade range to enforce | readability_ok, reading_ease | readability |
| Language adherence | oodle-managed-code-language-adherence-v1 | Output is written in the language you expect | language_match, detected_language | language_adherence |
Compare to a reference
These checks compare the output with a Reference setting. In an experiment, a check with no reference set uses the dataset item's expected output. On a live span, set the reference.
| Check | Template id | What it scores | Scores | Function |
|---|---|---|---|---|
| Exact match | oodle-managed-code-exact-match-v1 | Output equals a reference text | exact_match | exact_match |
| Edit similarity | oodle-managed-code-levenshtein-ratio-v1 | How close the output is to a reference text, by edit distance | levenshtein_ratio | levenshtein_ratio |
| Word overlap | oodle-managed-code-jaccard-similarity-v1 | Share of words the output and a reference text have in common | jaccard_similarity | jaccard_similarity |
| ROUGE | oodle-managed-code-rouge-v1 | ROUGE-L, ROUGE-1 and ROUGE-2 F1 of the output against a reference text | rouge_l, rouge_1, rouge_2 | rouge |
| BLEU | oodle-managed-code-bleu-v1 | Sentence BLEU of the output against a reference text | bleu | bleu |
| Numeric difference | oodle-managed-code-numeric-diff-v1 | How close the first number in the output is to a reference number | numeric_diff | numeric_diff |
| JSON difference | oodle-managed-code-json-diff-v1 | How close the output's JSON is to a reference JSON value, key by key | json_diff | json_diff |
Safety and tone
| Check | Template id | What it scores | Scores | Function |
|---|---|---|---|---|
| Tone | oodle-managed-code-tone-v1 | Flags shouting, too many exclamation marks, negative sentiment and banned phrases | tone_ok, sentiment | tone |
| Sentiment | oodle-managed-code-sentiment-v1 | Sentiment of the output from -1 (negative) to 1 (positive), from a word list | sentiment, sentiment_label | sentiment |
| PII and secret leak | oodle-managed-code-pii-leak-v1 | Output contains an email, phone, card number, national id, IP address or credential | pii_free, pii_count | pii_leak |
| URL check | oodle-managed-code-url-check-v1 | URLs are well formed, use an allowed scheme and domain, and do not point at a private address | urls_ok, url_count | url_check |
| Prompt injection | oodle-managed-code-prompt-injection-v1 | The input tries to override the model's instructions, directly or through a tool result | injection_detected, injection_risk, injection_signals | prompt_injection |
| Refusal | oodle-managed-code-refusal-v1 | The model declined to do what it was asked | refused | refusal |
Conversation
| Check | Template id | What it scores | Scores | Function |
|---|---|---|---|---|
| Conversation degeneration | oodle-managed-code-conversation-degeneration-v1 | The assistant repeats itself, within the reply or across turns | degenerated, degeneration | conversation_degeneration |
| Knowledge retention | oodle-managed-code-knowledge-retention-v1 | The reply asks again for, or contradicts, facts the user already gave | knowledge_retained, knowledge_retention | knowledge_retention |
Context and relevance
These checks compare the reply with the retrieved
context. They find the context in this order: the
context setting, then the span attribute that
context_key names (context by default), then the
text of the input messages. bm25_relevance and
tfidf_similarity compare with the question by
default; set against to context to compare with
the context.
| Check | Template id | What it scores | Scores | Function |
|---|---|---|---|---|
| Grounding | oodle-managed-code-grounding-v1 | Share of the reply's sentences whose words are in the retrieved context | grounded, grounding | grounding |
| Unsupported numbers | oodle-managed-code-unsupported-numbers-v1 | Numbers and names in the reply that the retrieved context does not have | numbers_supported, unsupported_count | unsupported_numbers |
| Citation check | oodle-managed-code-citation-check-v1 | Every [n] marker and URL in the reply points at a source the context has | citations_ok, citation_count | citation_check |
| BM25 relevance | oodle-managed-code-bm25-relevance-v1 | How relevant the reply is to the question or the context, by BM25 | bm25_relevance | bm25_relevance |
| TF-IDF similarity | oodle-managed-code-tfidf-similarity-v1 | Cosine similarity of the reply and the question or the context | tfidf_similarity | tfidf_similarity |
Agent and cost
The step checks (span_count, error_spans,
repeated_tool_calls, total_duration) read the
agent's steps in ctx.trace.spans, which an
experiment fills.
| Check | Template id | What it scores | Scores | Function |
|---|---|---|---|---|
| Tool call validity | oodle-managed-code-tool-call-validity-v1 | Tool calls have JSON arguments, use allowed tools, and do not loop | tool_calls_valid, tool_call_loop | tool_call_validity |
| Tool trajectory | oodle-managed-code-tool-trajectory-v1 | The agent called the tools you expect, in order or as a set | trajectory_match, tool_correctness | tool_trajectory |
| Step count | oodle-managed-code-span-count-v1 | How many steps the agent took, from its trace | span_count, within_span_limit | span_count |
| Failed steps | oodle-managed-code-error-spans-v1 | How many of the agent's steps failed, from its trace | no_errors, error_count | error_spans |
| Tool call loop | oodle-managed-code-repeated-tool-calls-v1 | The agent calls the same tool with the same input again and again | tool_loop, most_repeated | repeated_tool_calls |
| Total duration | oodle-managed-code-total-duration-v1 | How long the agent ran, from its trace | total_duration_ms, within_duration | total_duration |
| Token and latency budget | oodle-managed-code-token-latency-budget-v1 | The call stays inside its token, duration and time to first token limits | within_budget, budget_used | token_latency_budget |
list_templates in MCP and oodle genai templates list
in the CLI show these templates with the other managed
templates.
When you need more
- Change a check's logic. Start a code template from the check's starter (the template form offers each check as a starter), then change the code.
- Run several checks in one evaluator, or add your own score. Call the library functions from your own template. See the cookbook.
- Compare with a value from the span. Call a
reference check from your own template and pass
reference=fromctx. See the cookbook.
Built-in checks need the same account access as other code evaluators.
Support
If you need assistance or have any questions, please reach out to us through:
- Email at [email protected]