Skip to main content

Span Metrics

Oodle makes Prometheus metrics from the spans that you send. You do not need to configure them. You can query them with PromQL in dashboards, in the metrics explorer and in monitors. For alert examples, see Alerting on Logs and Traces.

Metric familyMade from
oodle_trace_metricsEvery span of every service.
oodle_genai_*Spans that have GenAI attributes (gen_ai.system, gen_ai.operation.name or a tool name).

oodle_trace_metrics​

oodle_trace_metrics counts spans. Oodle groups the spans by one-minute intervals of the span start time. Query it with increase() or rate().

LabelValue
service_nameThe service name of the span.
span_nameThe span name (operation).
span_statusOk, Error or Unset.
duration_ns_bucketThe lower edge of the duration bucket, in nanoseconds. See the table below.
is_root_spantrue for the root span of a trace, else false.
is_first_span_of_servicetrue for the first span in a service, for example the server span of a request. Else empty.
clusterThe Kubernetes cluster, when the span has it.
namespaceThe Kubernetes namespace, when the span has it.
envThe deployment environment, when the span has it.

Duration Buckets​

duration_ns_bucket is the largest bucket edge that is less than or equal to the span duration. A span of 1.5 s has duration_ns_bucket="1000000000".

duration_ns_bucketSpan duration
0less than 0.1 ms
1000000.1 ms to 0.2 ms
2000000.2 ms to 0.5 ms
5000000.5 ms to 1 ms
10000001 ms to 2 ms
20000002 ms to 4 ms
40000004 ms to 8 ms
80000008 ms to 16 ms
1600000016 ms to 32 ms
3200000032 ms to 64 ms
6400000064 ms to 128 ms
128000000128 ms to 256 ms
256000000256 ms to 512 ms
512000000512 ms to 1 s
10000000001 s to 2 s
20000000002 s to 4 s
40000000004 s to 8 s
80000000008 s to 16 s
1600000000016 s to 32 s
3200000000032 s to 64 s
6400000000064 s to 128 s
128000000000128 s or more

To select spans slower than a bucket edge, use a regex on the edges at or above it. For example, spans of 1 s or more:

duration_ns_bucket=~"1000000000|2000000000|4000000000|8000000000|16000000000|32000000000|64000000000|128000000000"

Example Queries​

Requests for each operation of a service in the last 5 minutes:

sum by (span_name) (
increase(oodle_trace_metrics{service_name="checkout", is_first_span_of_service="true"}[5m])
)

Error ratio for each service:

sum by (service_name) (increase(oodle_trace_metrics{span_status="Error"}[5m]))
/
sum by (service_name) (increase(oodle_trace_metrics[5m]))

Spans of 1 s or more for each operation:

sum by (service_name, span_name) (
increase(oodle_trace_metrics{
service_name="checkout",
duration_ns_bucket=~"1000000000|2000000000|4000000000|8000000000|16000000000|32000000000|64000000000|128000000000"
}[5m])
)

Share of requests that take 1 s or more:

sum(increase(oodle_trace_metrics{service_name="checkout", is_first_span_of_service="true", duration_ns_bucket=~"1000000000|2000000000|4000000000|8000000000|16000000000|32000000000|64000000000|128000000000"}[10m]))
/
sum(increase(oodle_trace_metrics{service_name="checkout", is_first_span_of_service="true"}[10m]))

GenAI Metrics: oodle_genai_*​

Oodle makes these metrics from GenAI spans, for example spans from OpenTelemetry GenAI instrumentation, OpenInference or OpenLLMetry. See Agent Observability to send these spans.

MetricTypeWhat it countsLabels
oodle_genai_span_totalcounterEvery GenAI span: model calls, agent and workflow spans, tool calls.env, gen_ai_system, model, operation, service_name, span_name, status (ok or error)
oodle_genai_span_latency_secondshistogram buckets (le)Duration of every GenAI span.env, gen_ai_system, model, operation, service_name, span_name, le
oodle_genai_span_error_totalcounterGenAI spans with an error.env, error_type, model, service_name, span_name
oodle_genai_llm_call_totalcounterDirect model calls only (chat, completion, embeddings).env, gen_ai_system, model, operation, service_name, status
oodle_genai_llm_call_latency_secondshistogram buckets (le)Duration of direct model calls.env, gen_ai_system, model, operation, service_name, le
oodle_genai_generation_tokens_totalcounterTokens of model calls.env, gen_ai_system, model, service_name, type (input, output, cache_read, cache_write)
oodle_genai_generation_cost_dollarscounterCost of model calls, in US dollars.env, gen_ai_system, model, service_name
oodle_genai_generation_ttft_secondshistogram buckets (le)Time to first token of streaming model calls.env, gen_ai_system, model, service_name, le
oodle_genai_tool_call_totalcounterTool calls.env, model, service_name, tool_name
oodle_genai_tool_error_totalcounterTool calls with an error.env, model, service_name, tool_name
oodle_genai_tool_latency_secondshistogram buckets (le)Duration of tool calls.env, model, service_name, tool_name, le
oodle_genai_tool_result_size_byteshistogram buckets (le)Size of tool results.env, model, service_name, tool_name, le

All GenAI metrics also have agent_name when the span reports an agent. The span, model call, token, cost and time-to-first-token metrics also have prompt_name and prompt_version when the span uses a managed prompt.

The histograms use cumulative le buckets:

Metricle buckets
oodle_genai_span_latency_seconds, oodle_genai_llm_call_latency_seconds, oodle_genai_tool_latency_secondsSeconds: 0.1, 0.25, 0.5, 1, 2, 4, 8, 16, 32, 64, 128, 256, 512, +Inf
oodle_genai_generation_ttft_secondsSeconds: 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2, 4, 8, 16, 32, +Inf
oodle_genai_tool_result_size_bytesBytes: 256, 1024, 4096, 16384, 65536, 262144, 1048576, 4194304, +Inf

The bucket series have the metric name itself, with no _bucket suffix. Use them with histogram_quantile(), or subtract two buckets to count calls above an edge.

Example Queries​

Tool calls slower than 8 s in the last 15 minutes, for each tool:

sum by (tool_name) (increase(oodle_genai_tool_latency_seconds{le="+Inf"}[15m]))
-
sum by (tool_name) (increase(oodle_genai_tool_latency_seconds{le="8"}[15m]))

p95 tool latency, in seconds:

histogram_quantile(
0.95,
sum by (tool_name, le) (rate(oodle_genai_tool_latency_seconds[10m]))
)

Tool error ratio:

sum by (tool_name) (increase(oodle_genai_tool_error_total[15m]))
/
sum by (tool_name) (increase(oodle_genai_tool_call_total[15m]))

Model call error ratio for each model:

sum by (model) (increase(oodle_genai_llm_call_total{status="error"}[15m]))
/
sum by (model) (increase(oodle_genai_llm_call_total[15m]))

LLM cost in the last hour, for each model:

sum by (model) (increase(oodle_genai_generation_cost_dollars[1h]))

To count spans above a threshold that is not a bucket edge, for example tool calls over exactly 10 s, use a TraceQL metrics query:

{ name =~ "execute_tool.*" && duration > 10s } | count_over_time() by (span.gen_ai.tool.name)

Support

If you need assistance or have any questions, please reach out to us through: