Observability
Measure request latency, token throughput, and pipeline health with Prometheus.
The worker that serves the API (rank 0) also serves Prometheus metrics at
GET /metrics. Request, generation, engine, and transport metrics all appear
on this one endpoint. Other ranks ignore --api and have no HTTP listener,
so each pipeline has exactly one scrape target: its rank 0 node. If you run
several independent pipelines, scrape each one’s rank 0 node.
curl -s http://127.0.0.1:8000/metrics | headPrometheus scrape config, with one target per pipeline:
scrape_configs: - job_name: cascadia static_configs: - targets: ["pipeline-a-rank0:8000", "pipeline-b-rank0:8000"]Metrics with labels appear after their first sample is recorded, which is standard Prometheus client behavior. Metrics without labels appear from startup.
Request metrics
Section titled “Request metrics”| Name | Type | Labels | Description |
|---|---|---|---|
cascadia_http_requests_total |
counter | endpoint, status |
Requests by route template (such as /v1/cancel/:task_id, never the raw URI, to keep cardinality bounded) and status code. Streaming responses are counted when headers are sent, so an engine failure partway through a stream still shows as 200 here. It appears in cascadia_tasks_failed_total instead. Routes for the CLI’s built-in dashboard aren’t counted. |
cascadia_http_request_duration_seconds |
histogram | endpoint |
Request latency. For streaming responses, this measures time until headers are sent. For full generation time, use cascadia_generation_duration_seconds. |
cascadia_inflight_tasks |
gauge | Generation requests currently running: admitted, with a response that hasn’t finished or been dropped. | |
cascadia_api_rejected_total |
counter | reason |
Requests rejected before generation started. See rejection reasons. |
Rejection reasons
Section titled “Rejection reasons”reason |
Status | Cause and fix |
|---|---|---|
capacity |
503 | The concurrency limit or the engine’s pending queue is full. Add workers. |
empty_prompt, multi_prompt |
400 | Malformed request. |
prompt_too_large |
413 | Prompt is over the API’s size limit. Raise --api-max-body-mb. |
prompt_over_window |
413 | Prompt is longer than the engine can handle for one request, such as a packed slot’s KV region. Resize slots or increase the context length. |
engine_unavailable |
503 | The engine isn’t loaded or a peer is unreachable. Start the missing stage. |
engine_error |
500 | The engine failed to accept the request. Check that node’s logs. |
invalid_request |
400 or 404 | Bad config, rejected shard, or unknown model. |
Requests rejected before reaching a handler, such as a malformed JSON body or one
over --api-max-body-mb, aren’t counted in cascadia_api_rejected_total.
Find them by status code in cascadia_http_requests_total.
A request the engine refuses never starts generating, so
cascadia_tasks_failed_total can’t see it. The only way to tell a worker
that fails every request apart from an idle one is engine_unavailable and
engine_error. Set alerts on both.
There’s no metric for queued requests. The API rejects a request with 503
as soon as the concurrency limit is reached instead of queueing it. Engines
do keep a bounded pending queue after that point, and a full queue is the
other cause of reason="capacity", but its depth isn’t exported. Watch
cascadia_api_rejected_total{reason="capacity"} to see pressure at either
point.
Generation metrics
Section titled “Generation metrics”These metrics are recorded as the response streams out, so they cover
streaming and non-streaming requests on /v1/chat/completions and
/v1/completions. The model label is the shard’s model_id.
| Name | Type | Labels | Description |
|---|---|---|---|
cascadia_generation_ttft_seconds |
histogram | model |
Time from request admission until the first token reaches the client. |
cascadia_generation_inter_token_seconds |
histogram | model |
Time between consecutive tokens in one generation, measured at delivery. |
cascadia_generation_duration_seconds |
histogram | model, finish_reason |
Time from admission to the end of the generation. finish_reason is stop, length, cancelled (client disconnected or cancelled), error (engine fault), or teardown (server shut down mid-generation). Every admitted generation records exactly one sample, so this count matches admissions, including across restarts. |
cascadia_tokens_generated_total |
counter | model |
Tokens delivered to clients. Counts multi-token chunks correctly for speculative decoding and ov-genai. Tokens generated after a client disconnects, before the cancellation takes effect, aren’t counted. |
cascadia_tokens_prompt_total |
counter | model |
Prompt tokens, for engines that report them. |
cascadia_tasks_cancelled_total |
counter | model |
Generations stopped early by /v1/cancel or a client disconnect. Generations interrupted by a server shutdown aren’t counted here. They’re recorded with finish_reason="teardown". |
cascadia_tasks_failed_total |
counter | model |
Generations that ended with an engine error. |
ov-genai without --cb is the default worker configuration. It returns the
whole response at once, so it records one time-to-first-token sample equal to
the full generation time and no inter-token samples. In this mode, read
cascadia_generation_ttft_seconds as total generation latency. Engines
that stream token by token (ov-genai --cb, ov-runtime, sparse-moe, and
mock) fill both histograms as described.
Timing is measured when tokens reach the client, not when the engine
produces them. With a single stream, the two are close. When several streams
share one engine, tokens can wait in a buffer, and that wait is included.
For streaming requests that use tool calls, the API buffers the whole
generation before responding, so cascadia_http_request_duration_seconds
covers the full generation for those requests.
Engine metrics
Section titled “Engine metrics”Set once when the stage starts.
| Name | Type | Labels | Description |
|---|---|---|---|
cascadia_engine_model_load_duration_seconds |
gauge | model, device |
Time to load weights and build the engine. Excludes connecting to other stages, which can include waiting for a peer to start. |
cascadia_engine_warmup_duration_seconds |
gauge | model, device |
Engine warmup time. |
Transport metrics
Section titled “Transport metrics”These cover traffic between stages for every engine. kind is tensor for
activation tensors (including headers) or raw for speculative-decoding
control messages.
| Name | Type | Labels | Description |
|---|---|---|---|
cascadia_transport_sent_bytes_total |
counter | kind |
Bytes sent between stages, counting complete frames only. |
cascadia_transport_recv_bytes_total |
counter | kind |
Bytes received between stages, counting complete frames only. |
cascadia_transport_send_seconds |
histogram | Time to send a tensor frame, including flush. | |
cascadia_transport_recv_payload_seconds |
histogram | Time from receiving a tensor frame’s header to receiving its full payload. Time spent waiting for a frame to start isn’t included, since idle stages wait there between requests. |
Only complete frames are counted. Bytes from a send or receive that fails
partway, such as a timeout on a bad link, aren’t recorded, and neither are
their durations. A link that stalls mid-frame produces no histogram sample at
all, so recv_payload_seconds keeps showing normal percentiles while
recv_bytes_total stops increasing. To catch a stalled link, alert on the
rate of recv_bytes_total, not on the histogram.
You can’t compare sent_bytes on one stage with recv_bytes on the next,
because only rank 0 serves /metrics.
Histogram buckets
Section titled “Histogram buckets”Buckets are tuned for LLM serving and defined in
crates/cascadia-metrics/src/lib.rs:
- Time to first token:
0.05 0.1 0.25 0.5 1 2 5 10 30 60 - Inter-token:
0.01 0.025 0.05 0.1 0.25 0.5 1 2 5 10 - Generation and HTTP durations:
0.1 0.5 1 5 10 30 60 120 300 600 - Transport frames:
0.0001 0.001 0.01 0.05 0.1 0.5 1
Example queries
Section titled “Example queries”# Request rate by endpointsum by (endpoint) (rate(cascadia_http_requests_total[5m]))
# p95 time to first tokenhistogram_quantile(0.95, sum by (le) (rate(cascadia_generation_ttft_seconds_bucket[5m])))
# Tokens per second across all pipelinessum(rate(cascadia_tokens_generated_total[1m]))
# Capacity rejections (503s)rate(cascadia_api_rejected_total{reason="capacity"}[5m])
# Throughput between stagesrate(cascadia_transport_sent_bytes_total[1m])Limitations
Section titled “Limitations”- Ranks other than 0 don’t expose metrics.
- There are no per-layer expert metrics for
sparse-moe. - OTLP export and trace exemplars aren’t supported.