Span Metrics
Oodle makes Prometheus metrics from the spans that you send. You do not need to configure them. You can query them with PromQL in dashboards, in the metrics explorer and in monitors. For alert examples, see Alerting on Logs and Traces.
| Metric family | Made from |
|---|---|
oodle_trace_metrics | Every span of every service. |
oodle_genai_* | Spans that have GenAI attributes (gen_ai.system, gen_ai.operation.name or a tool name). |
oodle_trace_metrics
oodle_trace_metrics counts spans. Oodle groups the spans by one-minute
intervals of the span start time. Query it with increase() or
rate().
| Label | Value |
|---|---|
service_name | The service name of the span. |
span_name | The span name (operation). |
span_status | Ok, Error or Unset. |
duration_ns_bucket | The lower edge of the duration bucket, in nanoseconds. See the table below. |
is_root_span | true for the root span of a trace, else false. |
is_first_span_of_service | true for the first span in a service, for example the server span of a request. Else empty. |
cluster | The Kubernetes cluster, when the span has it. |
namespace | The Kubernetes namespace, when the span has it. |
env | The deployment environment, when the span has it. |
Duration Buckets
duration_ns_bucket is the largest bucket edge that is less than or
equal to the span duration. A span of 1.5 s has
duration_ns_bucket="1000000000".
duration_ns_bucket | Span duration |
|---|---|
0 | less than 0.1 ms |
100000 | 0.1 ms to 0.2 ms |
200000 | 0.2 ms to 0.5 ms |
500000 | 0.5 ms to 1 ms |
1000000 | 1 ms to 2 ms |
2000000 | 2 ms to 4 ms |
4000000 | 4 ms to 8 ms |
8000000 | 8 ms to 16 ms |
16000000 | 16 ms to 32 ms |
32000000 | 32 ms to 64 ms |
64000000 | 64 ms to 128 ms |
128000000 | 128 ms to 256 ms |
256000000 | 256 ms to 512 ms |
512000000 | 512 ms to 1 s |
1000000000 | 1 s to 2 s |
2000000000 | 2 s to 4 s |
4000000000 | 4 s to 8 s |
8000000000 | 8 s to 16 s |
16000000000 | 16 s to 32 s |
32000000000 | 32 s to 64 s |
64000000000 | 64 s to 128 s |
128000000000 | 128 s or more |
To select spans slower than a bucket edge, use a regex on the edges at or above it. For example, spans of 1 s or more:
duration_ns_bucket=~"1000000000|2000000000|4000000000|8000000000|16000000000|32000000000|64000000000|128000000000"
Example Queries
Requests for each operation of a service in the last 5 minutes:
sum by (span_name) (
increase(oodle_trace_metrics{service_name="checkout", is_first_span_of_service="true"}[5m])
)
Error ratio for each service:
sum by (service_name) (increase(oodle_trace_metrics{span_status="Error"}[5m]))
/
sum by (service_name) (increase(oodle_trace_metrics[5m]))
Spans of 1 s or more for each operation:
sum by (service_name, span_name) (
increase(oodle_trace_metrics{
service_name="checkout",
duration_ns_bucket=~"1000000000|2000000000|4000000000|8000000000|16000000000|32000000000|64000000000|128000000000"
}[5m])
)
Share of requests that take 1 s or more:
sum(increase(oodle_trace_metrics{service_name="checkout", is_first_span_of_service="true", duration_ns_bucket=~"1000000000|2000000000|4000000000|8000000000|16000000000|32000000000|64000000000|128000000000"}[10m]))
/
sum(increase(oodle_trace_metrics{service_name="checkout", is_first_span_of_service="true"}[10m]))
GenAI Metrics: oodle_genai_*
Oodle makes these metrics from GenAI spans, for example spans from OpenTelemetry GenAI instrumentation, OpenInference or OpenLLMetry. See Agent Observability to send these spans.
| Metric | Type | What it counts | Labels |
|---|---|---|---|
oodle_genai_span_total | counter | Every GenAI span: model calls, agent and workflow spans, tool calls. | env, gen_ai_system, model, operation, service_name, span_name, status (ok or error) |
oodle_genai_span_latency_seconds | histogram buckets (le) | Duration of every GenAI span. | env, gen_ai_system, model, operation, service_name, span_name, le |
oodle_genai_span_error_total | counter | GenAI spans with an error. | env, error_type, model, service_name, span_name |
oodle_genai_llm_call_total | counter | Direct model calls only (chat, completion, embeddings). | env, gen_ai_system, model, operation, service_name, status |
oodle_genai_llm_call_latency_seconds | histogram buckets (le) | Duration of direct model calls. | env, gen_ai_system, model, operation, service_name, le |
oodle_genai_generation_tokens_total | counter | Tokens of model calls. | env, gen_ai_system, model, service_name, type (input, output, cache_read, cache_write) |
oodle_genai_generation_cost_dollars | counter | Cost of model calls, in US dollars. | env, gen_ai_system, model, service_name |
oodle_genai_generation_ttft_seconds | histogram buckets (le) | Time to first token of streaming model calls. | env, gen_ai_system, model, service_name, le |
oodle_genai_tool_call_total | counter | Tool calls. | env, model, service_name, tool_name |
oodle_genai_tool_error_total | counter | Tool calls with an error. | env, model, service_name, tool_name |
oodle_genai_tool_latency_seconds | histogram buckets (le) | Duration of tool calls. | env, model, service_name, tool_name, le |
oodle_genai_tool_result_size_bytes | histogram buckets (le) | Size of tool results. | env, model, service_name, tool_name, le |
All GenAI metrics also have agent_name when the span reports an agent.
The span, model call, token, cost and time-to-first-token metrics also
have prompt_name and prompt_version when the span uses a
managed prompt.
The histograms use cumulative le buckets:
| Metric | le buckets |
|---|---|
oodle_genai_span_latency_seconds, oodle_genai_llm_call_latency_seconds, oodle_genai_tool_latency_seconds | Seconds: 0.1, 0.25, 0.5, 1, 2, 4, 8, 16, 32, 64, 128, 256, 512, +Inf |
oodle_genai_generation_ttft_seconds | Seconds: 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2, 4, 8, 16, 32, +Inf |
oodle_genai_tool_result_size_bytes | Bytes: 256, 1024, 4096, 16384, 65536, 262144, 1048576, 4194304, +Inf |
The bucket series have the metric name itself, with no _bucket
suffix. Use them with histogram_quantile(), or subtract two
buckets to count calls above an edge.
Example Queries
Tool calls slower than 8 s in the last 15 minutes, for each tool:
sum by (tool_name) (increase(oodle_genai_tool_latency_seconds{le="+Inf"}[15m]))
-
sum by (tool_name) (increase(oodle_genai_tool_latency_seconds{le="8"}[15m]))
p95 tool latency, in seconds:
histogram_quantile(
0.95,
sum by (tool_name, le) (rate(oodle_genai_tool_latency_seconds[10m]))
)
Tool error ratio:
sum by (tool_name) (increase(oodle_genai_tool_error_total[15m]))
/
sum by (tool_name) (increase(oodle_genai_tool_call_total[15m]))
Model call error ratio for each model:
sum by (model) (increase(oodle_genai_llm_call_total{status="error"}[15m]))
/
sum by (model) (increase(oodle_genai_llm_call_total[15m]))
LLM cost in the last hour, for each model:
sum by (model) (increase(oodle_genai_generation_cost_dollars[1h]))
To count spans above a threshold that is not a bucket edge, for example tool calls over exactly 10 s, use a TraceQL metrics query:
{ name =~ "execute_tool.*" && duration > 10s } | count_over_time() by (span.gen_ai.tool.name)
Support
If you need assistance or have any questions, please reach out to us through:
- Email at [email protected]