Skip to main content

Built-in Code Checks

Built-in code checks are code evaluators that Oodle writes and maintains. You write no code: pick a check, set its settings, and choose the spans it scores. Like all code evaluators, they are fast and make no LLM call.

Use a built-in check​

  1. Go to Agent Observability → Evaluators (ap1, us1)
  2. Click Create Evaluator and pick a check from Built-in checks
  3. Set the check's settings, for example the keywords that the reply must mention
  4. Set the filters, the sampling rate and the hourly cap, as for any evaluator
  5. Click Save

Each check is a managed template with the id oodle-managed-code-<check>-v1. Use that id as the template when you create an evaluator through the API (evaluatorId), the CLI or MCP (templateId):

{"action": "create_evaluator", "name": "Support replies mention refunds",
"templateId": "oodle-managed-code-keyword-check-v1",
"params": {"required": ["refund"], "mode": "any"}, "enabled": false}

The settings of each check are the keyword arguments of its library function. The library reference gives each argument and its default. A setting that you do not set takes that default.

The checks​

The first score of each check is its primary score.

Format​

CheckTemplate idWhat it scoresScoresFunction
JSON validityoodle-managed-code-json-validity-v1Output parses as JSON and has the keys, value types and schema you setis_json, json_schema_okjson_validity
Regex matchoodle-managed-code-regex-match-v1Output matches every required pattern and none of the forbidden onesregex_passregex_match
Keyword checkoodle-managed-code-keyword-check-v1Output mentions the required keywords and none of the forbidden oneskeywords_ok, keyword_coveragekeyword_check
Containsoodle-managed-code-contains-v1Output contains the values you list, anywhere in the textcontainscontains
Length budgetoodle-managed-code-length-budget-v1Output stays inside a word and character rangewithin_lengthlength_budget
Markdown structureoodle-managed-code-markdown-structure-v1Code fences are closed, tables are even, and the headings, lists and tables you require are theremarkdown_ok, heading_countmarkdown_structure
HTML validityoodle-managed-code-html-validity-v1The HTML in the output has balanced tags, and only the tags you allowhtml_validhtml_validity
CSV validityoodle-managed-code-csv-validity-v1Output is CSV with the same number of fields in every row, and the columns you requirecsv_valid, row_countcsv_validity
TOML validityoodle-managed-code-toml-validity-v1Output parses as TOML and has the keys you requiretoml_validtoml_validity
SQL shapeoodle-managed-code-sql-shape-v1Generated SQL is only the statement types you allow, with no write keyword you did not allowsql_ok, statement_typesql_shape

Text quality​

CheckTemplate idWhat it scoresScoresFunction
Readabilityoodle-managed-code-readability-v1Flesch reading ease and Flesch-Kincaid grade, with a grade range to enforcereadability_ok, reading_easereadability
Language adherenceoodle-managed-code-language-adherence-v1Output is written in the language you expectlanguage_match, detected_languagelanguage_adherence

Compare to a reference​

These checks compare the output with a Reference setting. In an experiment, a check with no reference set uses the dataset item's expected output. On a live span, set the reference.

CheckTemplate idWhat it scoresScoresFunction
Exact matchoodle-managed-code-exact-match-v1Output equals a reference textexact_matchexact_match
Edit similarityoodle-managed-code-levenshtein-ratio-v1How close the output is to a reference text, by edit distancelevenshtein_ratiolevenshtein_ratio
Word overlapoodle-managed-code-jaccard-similarity-v1Share of words the output and a reference text have in commonjaccard_similarityjaccard_similarity
ROUGEoodle-managed-code-rouge-v1ROUGE-L, ROUGE-1 and ROUGE-2 F1 of the output against a reference textrouge_l, rouge_1, rouge_2rouge
BLEUoodle-managed-code-bleu-v1Sentence BLEU of the output against a reference textbleubleu
Numeric differenceoodle-managed-code-numeric-diff-v1How close the first number in the output is to a reference numbernumeric_diffnumeric_diff
JSON differenceoodle-managed-code-json-diff-v1How close the output's JSON is to a reference JSON value, key by keyjson_diffjson_diff

Safety and tone​

CheckTemplate idWhat it scoresScoresFunction
Toneoodle-managed-code-tone-v1Flags shouting, too many exclamation marks, negative sentiment and banned phrasestone_ok, sentimenttone
Sentimentoodle-managed-code-sentiment-v1Sentiment of the output from -1 (negative) to 1 (positive), from a word listsentiment, sentiment_labelsentiment
PII and secret leakoodle-managed-code-pii-leak-v1Output contains an email, phone, card number, national id, IP address or credentialpii_free, pii_countpii_leak
URL checkoodle-managed-code-url-check-v1URLs are well formed, use an allowed scheme and domain, and do not point at a private addressurls_ok, url_counturl_check
Prompt injectionoodle-managed-code-prompt-injection-v1The input tries to override the model's instructions, directly or through a tool resultinjection_detected, injection_risk, injection_signalsprompt_injection
Refusaloodle-managed-code-refusal-v1The model declined to do what it was askedrefusedrefusal

Conversation​

CheckTemplate idWhat it scoresScoresFunction
Conversation degenerationoodle-managed-code-conversation-degeneration-v1The assistant repeats itself, within the reply or across turnsdegenerated, degenerationconversation_degeneration
Knowledge retentionoodle-managed-code-knowledge-retention-v1The reply asks again for, or contradicts, facts the user already gaveknowledge_retained, knowledge_retentionknowledge_retention

Context and relevance​

These checks compare the reply with the retrieved context. They find the context in this order: the context setting, then the span attribute that context_key names (context by default), then the text of the input messages. bm25_relevance and tfidf_similarity compare with the question by default; set against to context to compare with the context.

CheckTemplate idWhat it scoresScoresFunction
Groundingoodle-managed-code-grounding-v1Share of the reply's sentences whose words are in the retrieved contextgrounded, groundinggrounding
Unsupported numbersoodle-managed-code-unsupported-numbers-v1Numbers and names in the reply that the retrieved context does not havenumbers_supported, unsupported_countunsupported_numbers
Citation checkoodle-managed-code-citation-check-v1Every [n] marker and URL in the reply points at a source the context hascitations_ok, citation_countcitation_check
BM25 relevanceoodle-managed-code-bm25-relevance-v1How relevant the reply is to the question or the context, by BM25bm25_relevancebm25_relevance
TF-IDF similarityoodle-managed-code-tfidf-similarity-v1Cosine similarity of the reply and the question or the contexttfidf_similaritytfidf_similarity

Agent and cost​

The step checks (span_count, error_spans, repeated_tool_calls, total_duration) read the agent's steps in ctx.trace.spans, which an experiment fills.

CheckTemplate idWhat it scoresScoresFunction
Tool call validityoodle-managed-code-tool-call-validity-v1Tool calls have JSON arguments, use allowed tools, and do not looptool_calls_valid, tool_call_looptool_call_validity
Tool trajectoryoodle-managed-code-tool-trajectory-v1The agent called the tools you expect, in order or as a settrajectory_match, tool_correctnesstool_trajectory
Step countoodle-managed-code-span-count-v1How many steps the agent took, from its tracespan_count, within_span_limitspan_count
Failed stepsoodle-managed-code-error-spans-v1How many of the agent's steps failed, from its traceno_errors, error_counterror_spans
Tool call loopoodle-managed-code-repeated-tool-calls-v1The agent calls the same tool with the same input again and againtool_loop, most_repeatedrepeated_tool_calls
Total durationoodle-managed-code-total-duration-v1How long the agent ran, from its tracetotal_duration_ms, within_durationtotal_duration
Token and latency budgetoodle-managed-code-token-latency-budget-v1The call stays inside its token, duration and time to first token limitswithin_budget, budget_usedtoken_latency_budget

list_templates in MCP and oodle genai templates list in the CLI show these templates with the other managed templates.

When you need more​

  • Change a check's logic. Start a code template from the check's starter (the template form offers each check as a starter), then change the code.
  • Run several checks in one evaluator, or add your own score. Call the library functions from your own template. See the cookbook.
  • Compare with a value from the span. Call a reference check from your own template and pass reference= from ctx. See the cookbook.

Built-in checks need the same account access as other code evaluators.


Support

If you need assistance or have any questions, please reach out to us through: