Skip to content
CascadiaCascadiaDocsOpen source
Operations

Observability

Measure request latency, token throughput, and pipeline health with Prometheus.

The worker that serves the API (rank 0) also serves Prometheus metrics at GET /metrics. Request, generation, engine, and transport metrics all appear on this one endpoint. Other ranks ignore --api and have no HTTP listener, so each pipeline has exactly one scrape target: its rank 0 node. If you run several independent pipelines, scrape each one’s rank 0 node.

Terminal window
curl -s http://127.0.0.1:8000/metrics | head

Prometheus scrape config, with one target per pipeline:

scrape_configs:
- job_name: cascadia
static_configs:
- targets: ["pipeline-a-rank0:8000", "pipeline-b-rank0:8000"]

Metrics with labels appear after their first sample is recorded, which is standard Prometheus client behavior. Metrics without labels appear from startup.

Name Type Labels Description
cascadia_http_requests_total counter endpoint, status Requests by route template (such as /v1/cancel/:task_id, never the raw URI, to keep cardinality bounded) and status code. Streaming responses are counted when headers are sent, so an engine failure partway through a stream still shows as 200 here. It appears in cascadia_tasks_failed_total instead. Routes for the CLI’s built-in dashboard aren’t counted.
cascadia_http_request_duration_seconds histogram endpoint Request latency. For streaming responses, this measures time until headers are sent. For full generation time, use cascadia_generation_duration_seconds.
cascadia_inflight_tasks gauge Generation requests currently running: admitted, with a response that hasn’t finished or been dropped.
cascadia_api_rejected_total counter reason Requests rejected before generation started. See rejection reasons.
reason Status Cause and fix
capacity 503 The concurrency limit or the engine’s pending queue is full. Add workers.
empty_prompt, multi_prompt 400 Malformed request.
prompt_too_large 413 Prompt is over the API’s size limit. Raise --api-max-body-mb.
prompt_over_window 413 Prompt is longer than the engine can handle for one request, such as a packed slot’s KV region. Resize slots or increase the context length.
engine_unavailable 503 The engine isn’t loaded or a peer is unreachable. Start the missing stage.
engine_error 500 The engine failed to accept the request. Check that node’s logs.
invalid_request 400 or 404 Bad config, rejected shard, or unknown model.

Requests rejected before reaching a handler, such as a malformed JSON body or one over --api-max-body-mb, aren’t counted in cascadia_api_rejected_total. Find them by status code in cascadia_http_requests_total.

A request the engine refuses never starts generating, so cascadia_tasks_failed_total can’t see it. The only way to tell a worker that fails every request apart from an idle one is engine_unavailable and engine_error. Set alerts on both.

There’s no metric for queued requests. The API rejects a request with 503 as soon as the concurrency limit is reached instead of queueing it. Engines do keep a bounded pending queue after that point, and a full queue is the other cause of reason="capacity", but its depth isn’t exported. Watch cascadia_api_rejected_total{reason="capacity"} to see pressure at either point.

These metrics are recorded as the response streams out, so they cover streaming and non-streaming requests on /v1/chat/completions and /v1/completions. The model label is the shard’s model_id.

Name Type Labels Description
cascadia_generation_ttft_seconds histogram model Time from request admission until the first token reaches the client.
cascadia_generation_inter_token_seconds histogram model Time between consecutive tokens in one generation, measured at delivery.
cascadia_generation_duration_seconds histogram model, finish_reason Time from admission to the end of the generation. finish_reason is stop, length, cancelled (client disconnected or cancelled), error (engine fault), or teardown (server shut down mid-generation). Every admitted generation records exactly one sample, so this count matches admissions, including across restarts.
cascadia_tokens_generated_total counter model Tokens delivered to clients. Counts multi-token chunks correctly for speculative decoding and ov-genai. Tokens generated after a client disconnects, before the cancellation takes effect, aren’t counted.
cascadia_tokens_prompt_total counter model Prompt tokens, for engines that report them.
cascadia_tasks_cancelled_total counter model Generations stopped early by /v1/cancel or a client disconnect. Generations interrupted by a server shutdown aren’t counted here. They’re recorded with finish_reason="teardown".
cascadia_tasks_failed_total counter model Generations that ended with an engine error.

ov-genai without --cb is the default worker configuration. It returns the whole response at once, so it records one time-to-first-token sample equal to the full generation time and no inter-token samples. In this mode, read cascadia_generation_ttft_seconds as total generation latency. Engines that stream token by token (ov-genai --cb, ov-runtime, sparse-moe, and mock) fill both histograms as described.

Timing is measured when tokens reach the client, not when the engine produces them. With a single stream, the two are close. When several streams share one engine, tokens can wait in a buffer, and that wait is included. For streaming requests that use tool calls, the API buffers the whole generation before responding, so cascadia_http_request_duration_seconds covers the full generation for those requests.

Set once when the stage starts.

Name Type Labels Description
cascadia_engine_model_load_duration_seconds gauge model, device Time to load weights and build the engine. Excludes connecting to other stages, which can include waiting for a peer to start.
cascadia_engine_warmup_duration_seconds gauge model, device Engine warmup time.

These cover traffic between stages for every engine. kind is tensor for activation tensors (including headers) or raw for speculative-decoding control messages.

Name Type Labels Description
cascadia_transport_sent_bytes_total counter kind Bytes sent between stages, counting complete frames only.
cascadia_transport_recv_bytes_total counter kind Bytes received between stages, counting complete frames only.
cascadia_transport_send_seconds histogram Time to send a tensor frame, including flush.
cascadia_transport_recv_payload_seconds histogram Time from receiving a tensor frame’s header to receiving its full payload. Time spent waiting for a frame to start isn’t included, since idle stages wait there between requests.

Only complete frames are counted. Bytes from a send or receive that fails partway, such as a timeout on a bad link, aren’t recorded, and neither are their durations. A link that stalls mid-frame produces no histogram sample at all, so recv_payload_seconds keeps showing normal percentiles while recv_bytes_total stops increasing. To catch a stalled link, alert on the rate of recv_bytes_total, not on the histogram.

You can’t compare sent_bytes on one stage with recv_bytes on the next, because only rank 0 serves /metrics.

Buckets are tuned for LLM serving and defined in crates/cascadia-metrics/src/lib.rs:

  • Time to first token: 0.05 0.1 0.25 0.5 1 2 5 10 30 60
  • Inter-token: 0.01 0.025 0.05 0.1 0.25 0.5 1 2 5 10
  • Generation and HTTP durations: 0.1 0.5 1 5 10 30 60 120 300 600
  • Transport frames: 0.0001 0.001 0.01 0.05 0.1 0.5 1
# Request rate by endpoint
sum by (endpoint) (rate(cascadia_http_requests_total[5m]))
# p95 time to first token
histogram_quantile(0.95, sum by (le) (rate(cascadia_generation_ttft_seconds_bucket[5m])))
# Tokens per second across all pipelines
sum(rate(cascadia_tokens_generated_total[1m]))
# Capacity rejections (503s)
rate(cascadia_api_rejected_total{reason="capacity"}[5m])
# Throughput between stages
rate(cascadia_transport_sent_bytes_total[1m])
  • Ranks other than 0 don’t expose metrics.
  • There are no per-layer expert metrics for sparse-moe.
  • OTLP export and trace exemplars aren’t supported.