Skip to main content

Alerting on Logs and Traces

Oodle monitors evaluate PromQL. To alert on logs or on traces, you make the signal a metric, then you write a monitor on that metric.

SignalMetricSetup
Log lines, structured or plain textoodle_logs_<name> from a LogMetrics ruleCreate a rule.
Spans of all servicesoodle_trace_metricsNone. Oodle makes it from every span.
GenAI spans (model calls, tools, tokens, cost)oodle_genai_*None. Oodle makes it from GenAI spans.

The labels of the metric become the labels of the alert. A label that a LogMetrics regex takes from a log line, for example a user name, shows in the alert and in the notification.

Alert on Logs​

  1. Create a LogMetrics rule. The rule selects the log lines and emits a metric. It can take labels and numbers from any field, and from plain-text lines with a regular expression. See LogMetrics.
  2. Create a monitor on the metric. Use increase() or rate() on counters, and histogram_quantile() on histograms. Aggregate with sum by (...) and keep the labels that you want in the alert.

A rule applies to logs that arrive after you save it. The monitor can fire after the metric has data for the monitor time window.

Example: Alert When a User Has Too Many Failed Logins​

Log lines:

WARN auth Login failed for user [email protected]: bad password

Step 1: LogMetrics rule. Save this rule as login-failures.json and run oodle log-metrics create -f login-failures.json, or create it in Logs > LogMetrics:

{
"name": "login-failures-by-user",
"filter": {
"all": [
{ "field": "message", "operator": "contains", "value": "Login failed for user" }
]
},
"labels": [
{
"name": "user",
"valueExtractor": { "field": "message", "regex": "Login failed for user (\\S+?):" }
}
],
"metricDefinitions": [
{ "name": "login_failures", "type": "log_count" }
]
}

The rule emits the counter oodle_logs_login_failures with a user label.

Step 2: Monitor.

FieldValue
NameToo many failed logins for {{ $labels.user }}
Querysum by (user) (increase(oodle_logs_login_failures[10m]))
ConditionCritical: Above 5
Message{{ $labels.user }} had {{ $value }} failed logins in 10 minutes.
GroupingBy Labels, with user, to get one notification for each user

The same monitor in Terraform:

resource "oodle_monitor" "login_failures" {
name = "Too many failed logins for {{ $labels.user }}"
promql_query = "sum by (user) (increase(oodle_logs_login_failures[10m]))"
conditions = {
critical = {
value = 5
operation = ">"
for = "1m"
}
}
}

See the oodle_monitor resource for all the fields.

More Log Alert Queries​

Each query uses the metric of a LogMetrics rule with the name in the query. The slow_tool_calls and tool_latency_ms rules are examples on the LogMetrics page.

Error lines for each service, above 50 in 5 minutes:

sum by (service) (increase(oodle_logs_error_lines[5m])) > 50

p95 tool latency from plain-text in <n> ms lines, above 5 seconds:

histogram_quantile(
0.95,
sum by (tool, le) (rate(oodle_logs_tool_latency_ms_bucket[5m]))
) > 5000

Tool calls of 10 s or more, from a log_count rule with a threshold regex in its filter:

sum by (tool) (increase(oodle_logs_slow_tool_calls[15m])) > 0

No log lines from a job in the last 30 minutes (the job stopped):

(sum(increase(oodle_logs_backup_completed[30m])) or vector(0)) < 1

When no line matches, the metric has no series and sum returns no value, so the comparison has nothing to compare. or vector(0) gives the value 0 in that case, and the monitor fires.

Alert on Traces​

Oodle makes span metrics from all spans. You can write monitors on them without setup.

Error ratio of a service above 5%:

sum by (service_name) (increase(oodle_trace_metrics{service_name="checkout", span_status="Error"}[5m]))
/
sum by (service_name) (increase(oodle_trace_metrics{service_name="checkout"}[5m]))
> 0.05

More than 20 requests of 1 s or more in 5 minutes, for each operation:

sum by (service_name, span_name) (
increase(oodle_trace_metrics{
service_name="checkout",
is_first_span_of_service="true",
duration_ns_bucket=~"1000000000|2000000000|4000000000|8000000000|16000000000|32000000000|64000000000|128000000000"
}[5m])
) > 20

duration_ns_bucket is the lower edge of a power-of-two bucket in nanoseconds. See Duration Buckets for the edges.

GenAI Alerts​

GenAI tool calls slower than 8 s, for each tool:

(
sum by (tool_name) (increase(oodle_genai_tool_latency_seconds{le="+Inf"}[15m]))
-
sum by (tool_name) (increase(oodle_genai_tool_latency_seconds{le="8"}[15m]))
) > 0

Tool error ratio above 10%:

sum by (tool_name) (increase(oodle_genai_tool_error_total[15m]))
/
sum by (tool_name) (increase(oodle_genai_tool_call_total[15m]))
> 0.1

Model call p95 latency above 30 s:

histogram_quantile(
0.95,
sum by (model, le) (rate(oodle_genai_llm_call_latency_seconds[10m]))
) > 30

LLM cost above 50 US dollars in one hour:

sum(increase(oodle_genai_generation_cost_dollars[1h])) > 50

Choose the Right Tool​

QuestionUse
Tell me when a log pattern occurs, now and in the future.LogMetrics rule and a monitor.
Tell me when a service has errors or slow requests.Monitor on oodle_trace_metrics.
Tell me when agent tools fail or are slow, or when LLM cost grows.Monitor on oodle_genai_*.
How many matching log lines were there last week?Log search aggregation.
How many spans matched a filter last week?TraceQL metrics query.
Which traces match this filter?TraceQL search or the Trace Explorer.

Answer Questions About the Past​

LogMetrics rules do not backfill. Use log search and TraceQL to answer questions about data that Oodle already stored, and to test a filter before you make a rule from it.

Log Search Aggregations​

The Logs Query API, the log explorer and the Oodle MCP server accept OpenSearch Query DSL with these aggregations: terms, date_histogram, range, filters, avg, min, max, sum, percentiles, percentile_ranks and cardinality.

Failed logins for each user in one week, from a structured user field. The time range is in epoch milliseconds:

{"index": "<INDEX_PATTERN>"}
{"size": 0, "query": {"bool": {"filter": [{"match_phrase": {"message": "Login failed"}}, {"range": {"timestamp": {"gte": 1759190400000, "lte": 1759795200000, "format": "epoch_millis"}}}]}}, "aggs": {"by_user": {"terms": {"field": "user.keyword", "size": 20}}}}

Error lines in each hour:

{"index": "<INDEX_PATTERN>"}
{"size": 0, "query": {"bool": {"filter": [{"match_phrase": {"level.keyword": "ERROR"}}, {"range": {"timestamp": {"gte": 1759190400000, "lte": 1759795200000, "format": "epoch_millis"}}}]}}, "aggs": {"per_hour": {"date_histogram": {"field": "timestamp", "fixed_interval": "1h"}}}}

Aggregations group and compute on log fields. When the value is inside the text of a message, for example in 1250 ms, make it a field first:

  • A log transform extracts the value into a field when Oodle ingests the log. Aggregations can then use the field.
  • A LogMetrics rule extracts the value into a metric with a regex. Use it for alerts and for long-term trends.

To count past lines that contain a text pattern, filter on the text with match_phrase and count with date_histogram or the hit total.

TraceQL​

TraceQL searches the stored spans and computes metrics from them, with any span attribute. For example, the number of GenAI tool calls over 10 s for each tool:

{ name =~ "execute_tool.*" && duration > 10s } | count_over_time() by (span.gen_ai.tool.name)

Run TraceQL in Grafana Explore, with the HTTP API, with oodle traces traceql, or with the query_traceql MCP tool.


Support

If you need assistance or have any questions, please reach out to us through: