Alerting on Logs and Traces
Oodle monitors evaluate PromQL. To alert on logs or on traces, you make the signal a metric, then you write a monitor on that metric.
| Signal | Metric | Setup |
|---|---|---|
| Log lines, structured or plain text | oodle_logs_<name> from a LogMetrics rule | Create a rule. |
| Spans of all services | oodle_trace_metrics | None. Oodle makes it from every span. |
| GenAI spans (model calls, tools, tokens, cost) | oodle_genai_* | None. Oodle makes it from GenAI spans. |
The labels of the metric become the labels of the alert. A label that a LogMetrics regex takes from a log line, for example a user name, shows in the alert and in the notification.
Alert on Logs
- Create a LogMetrics rule. The rule selects the log lines and emits a metric. It can take labels and numbers from any field, and from plain-text lines with a regular expression. See LogMetrics.
- Create a monitor on the metric. Use
increase()orrate()on counters, andhistogram_quantile()on histograms. Aggregate withsum by (...)and keep the labels that you want in the alert.
A rule applies to logs that arrive after you save it. The monitor can fire after the metric has data for the monitor time window.
Example: Alert When a User Has Too Many Failed Logins
Log lines:
WARN auth Login failed for user [email protected]: bad password
Step 1: LogMetrics rule. Save this rule as login-failures.json and
run oodle log-metrics create -f login-failures.json, or create it in
Logs > LogMetrics:
{
"name": "login-failures-by-user",
"filter": {
"all": [
{ "field": "message", "operator": "contains", "value": "Login failed for user" }
]
},
"labels": [
{
"name": "user",
"valueExtractor": { "field": "message", "regex": "Login failed for user (\\S+?):" }
}
],
"metricDefinitions": [
{ "name": "login_failures", "type": "log_count" }
]
}
The rule emits the counter oodle_logs_login_failures with a user
label.
Step 2: Monitor.
| Field | Value |
|---|---|
| Name | Too many failed logins for {{ $labels.user }} |
| Query | sum by (user) (increase(oodle_logs_login_failures[10m])) |
| Condition | Critical: Above 5 |
| Message | {{ $labels.user }} had {{ $value }} failed logins in 10 minutes. |
| Grouping | By Labels, with user, to get one notification for each user |
The same monitor in Terraform:
resource "oodle_monitor" "login_failures" {
name = "Too many failed logins for {{ $labels.user }}"
promql_query = "sum by (user) (increase(oodle_logs_login_failures[10m]))"
conditions = {
critical = {
value = 5
operation = ">"
for = "1m"
}
}
}
See the
oodle_monitor resource
for all the fields.
More Log Alert Queries
Each query uses the metric of a LogMetrics rule with the name in the
query. The slow_tool_calls and tool_latency_ms rules are examples on
the LogMetrics page.
Error lines for each service, above 50 in 5 minutes:
sum by (service) (increase(oodle_logs_error_lines[5m])) > 50
p95 tool latency from plain-text in <n> ms lines, above 5 seconds:
histogram_quantile(
0.95,
sum by (tool, le) (rate(oodle_logs_tool_latency_ms_bucket[5m]))
) > 5000
Tool calls of 10 s or more, from a log_count rule with a threshold
regex in its filter:
sum by (tool) (increase(oodle_logs_slow_tool_calls[15m])) > 0
No log lines from a job in the last 30 minutes (the job stopped):
(sum(increase(oodle_logs_backup_completed[30m])) or vector(0)) < 1
When no line matches, the metric has no series and sum returns no
value, so the comparison has nothing to compare. or vector(0) gives
the value 0 in that case, and the monitor fires.
Alert on Traces
Oodle makes span metrics from all spans. You can write monitors on them without setup.
Error ratio of a service above 5%:
sum by (service_name) (increase(oodle_trace_metrics{service_name="checkout", span_status="Error"}[5m]))
/
sum by (service_name) (increase(oodle_trace_metrics{service_name="checkout"}[5m]))
> 0.05
More than 20 requests of 1 s or more in 5 minutes, for each operation:
sum by (service_name, span_name) (
increase(oodle_trace_metrics{
service_name="checkout",
is_first_span_of_service="true",
duration_ns_bucket=~"1000000000|2000000000|4000000000|8000000000|16000000000|32000000000|64000000000|128000000000"
}[5m])
) > 20
duration_ns_bucket is the lower edge of a power-of-two bucket in
nanoseconds. See Duration Buckets
for the edges.
GenAI Alerts
GenAI tool calls slower than 8 s, for each tool:
(
sum by (tool_name) (increase(oodle_genai_tool_latency_seconds{le="+Inf"}[15m]))
-
sum by (tool_name) (increase(oodle_genai_tool_latency_seconds{le="8"}[15m]))
) > 0
Tool error ratio above 10%:
sum by (tool_name) (increase(oodle_genai_tool_error_total[15m]))
/
sum by (tool_name) (increase(oodle_genai_tool_call_total[15m]))
> 0.1
Model call p95 latency above 30 s:
histogram_quantile(
0.95,
sum by (model, le) (rate(oodle_genai_llm_call_latency_seconds[10m]))
) > 30
LLM cost above 50 US dollars in one hour:
sum(increase(oodle_genai_generation_cost_dollars[1h])) > 50
Choose the Right Tool
| Question | Use |
|---|---|
| Tell me when a log pattern occurs, now and in the future. | LogMetrics rule and a monitor. |
| Tell me when a service has errors or slow requests. | Monitor on oodle_trace_metrics. |
| Tell me when agent tools fail or are slow, or when LLM cost grows. | Monitor on oodle_genai_*. |
| How many matching log lines were there last week? | Log search aggregation. |
| How many spans matched a filter last week? | TraceQL metrics query. |
| Which traces match this filter? | TraceQL search or the Trace Explorer. |
Answer Questions About the Past
LogMetrics rules do not backfill. Use log search and TraceQL to answer questions about data that Oodle already stored, and to test a filter before you make a rule from it.
Log Search Aggregations
The Logs Query API, the log explorer and the
Oodle MCP server accept OpenSearch Query DSL with these aggregations:
terms, date_histogram, range, filters, avg, min, max,
sum, percentiles, percentile_ranks and cardinality.
Failed logins for each user in one week, from a structured user
field. The time range is in epoch milliseconds:
{"index": "<INDEX_PATTERN>"}
{"size": 0, "query": {"bool": {"filter": [{"match_phrase": {"message": "Login failed"}}, {"range": {"timestamp": {"gte": 1759190400000, "lte": 1759795200000, "format": "epoch_millis"}}}]}}, "aggs": {"by_user": {"terms": {"field": "user.keyword", "size": 20}}}}
Error lines in each hour:
{"index": "<INDEX_PATTERN>"}
{"size": 0, "query": {"bool": {"filter": [{"match_phrase": {"level.keyword": "ERROR"}}, {"range": {"timestamp": {"gte": 1759190400000, "lte": 1759795200000, "format": "epoch_millis"}}}]}}, "aggs": {"per_hour": {"date_histogram": {"field": "timestamp", "fixed_interval": "1h"}}}}
Aggregations group and compute on log fields. When the value is inside
the text of a message, for example in 1250 ms, make it a field first:
- A log transform extracts the value into a field when Oodle ingests the log. Aggregations can then use the field.
- A LogMetrics rule extracts the value into a metric with a regex. Use it for alerts and for long-term trends.
To count past lines that contain a text pattern, filter on the text
with match_phrase and count with date_histogram or the hit total.
TraceQL
TraceQL searches the stored spans and computes metrics from them, with any span attribute. For example, the number of GenAI tool calls over 10 s for each tool:
{ name =~ "execute_tool.*" && duration > 10s } | count_over_time() by (span.gen_ai.tool.name)
Run TraceQL in Grafana Explore, with the HTTP API, with
oodle traces traceql, or with the query_traceql MCP tool.
Support
If you need assistance or have any questions, please reach out to us through:
- Email at [email protected]