# Observability

> Measure request latency, token throughput, and pipeline health with Prometheus.

The worker that serves the API (rank 0) also serves Prometheus metrics at
`GET /metrics`. Request, generation, engine, and transport metrics all appear
on this one endpoint. Other ranks ignore `--api` and have no HTTP listener,
so each pipeline has exactly one scrape target: its rank 0 node. If you run
several independent pipelines, scrape each one's rank 0 node.

```bash
curl -s http://127.0.0.1:8000/metrics | head
```

Prometheus scrape config, with one target per pipeline:

```yaml
scrape_configs:
  - job_name: cascadia
    static_configs:
      - targets: ["pipeline-a-rank0:8000", "pipeline-b-rank0:8000"]
```

Metrics with labels appear after their first sample is recorded, which is
standard Prometheus client behavior. Metrics without labels appear from
startup.

## Request metrics

| Name | Type | Labels | Description |
|---|---|---|---|
| `cascadia_http_requests_total` | counter | `endpoint`, `status` | Requests by route template (such as `/v1/cancel/:task_id`, never the raw URI, to keep cardinality bounded) and status code. Streaming responses are counted when headers are sent, so an engine failure partway through a stream still shows as 200 here. It appears in `cascadia_tasks_failed_total` instead. Routes for the CLI's built-in dashboard aren't counted. |
| `cascadia_http_request_duration_seconds` | histogram | `endpoint` | Request latency. For streaming responses, this measures time until headers are sent. For full generation time, use `cascadia_generation_duration_seconds`. |
| `cascadia_inflight_tasks` | gauge | | Generation requests currently running: admitted, with a response that hasn't finished or been dropped. |
| `cascadia_api_rejected_total` | counter | `reason` | Requests rejected before generation started. See [rejection reasons](#rejection-reasons). |

### Rejection reasons

| `reason` | Status | Cause and fix |
|---|---|---|
| `capacity` | 503 | The concurrency limit or the engine's pending queue is full. Add workers. |
| `empty_prompt`, `multi_prompt` | 400 | Malformed request. |
| `prompt_too_large` | 413 | Prompt is over the API's size limit. Raise `--api-max-body-mb`. |
| `prompt_over_window` | 413 | Prompt is longer than the engine can handle for one request, such as a packed slot's KV region. Resize slots or increase the context length. |
| `engine_unavailable` | 503 | The engine isn't loaded or a peer is unreachable. Start the missing stage. |
| `engine_error` | 500 | The engine failed to accept the request. Check that node's logs. |
| `invalid_request` | 400 or 404 | Bad config, rejected shard, or unknown model. |

Requests rejected before reaching a handler, such as a malformed JSON body or one
over `--api-max-body-mb`, aren't counted in `cascadia_api_rejected_total`.
Find them by status code in `cascadia_http_requests_total`.

A request the engine refuses never starts generating, so
`cascadia_tasks_failed_total` can't see it. The only way to tell a worker
that fails every request apart from an idle one is `engine_unavailable` and
`engine_error`. Set alerts on both.

There's no metric for queued requests. The API rejects a request with 503
as soon as the concurrency limit is reached instead of queueing it. Engines
do keep a bounded pending queue after that point, and a full queue is the
other cause of `reason="capacity"`, but its depth isn't exported. Watch
`cascadia_api_rejected_total{reason="capacity"}` to see pressure at either
point.

## Generation metrics

These metrics are recorded as the response streams out, so they cover
streaming and non-streaming requests on `/v1/chat/completions` and
`/v1/completions`. The `model` label is the shard's `model_id`.

| Name | Type | Labels | Description |
|---|---|---|---|
| `cascadia_generation_ttft_seconds` | histogram | `model` | Time from request admission until the first token reaches the client. |
| `cascadia_generation_inter_token_seconds` | histogram | `model` | Time between consecutive tokens in one generation, measured at delivery. |
| `cascadia_generation_duration_seconds` | histogram | `model`, `finish_reason` | Time from admission to the end of the generation. `finish_reason` is `stop`, `length`, `cancelled` (client disconnected or cancelled), `error` (engine fault), or `teardown` (server shut down mid-generation). Every admitted generation records exactly one sample, so this count matches admissions, including across restarts. |
| `cascadia_tokens_generated_total` | counter | `model` | Tokens delivered to clients. Counts multi-token chunks correctly for speculative decoding and `ov-genai`. Tokens generated after a client disconnects, before the cancellation takes effect, aren't counted. |
| `cascadia_tokens_prompt_total` | counter | `model` | Prompt tokens, for engines that report them. |
| `cascadia_tasks_cancelled_total` | counter | `model` | Generations stopped early by `/v1/cancel` or a client disconnect. Generations interrupted by a server shutdown aren't counted here. They're recorded with `finish_reason="teardown"`. |
| `cascadia_tasks_failed_total` | counter | `model` | Generations that ended with an engine error. |

`ov-genai` without `--cb` is the default worker configuration. It returns the
whole response at once, so it records one time-to-first-token sample equal to
the full generation time and no inter-token samples. In this mode, read
`cascadia_generation_ttft_seconds` as total generation latency. Engines
that stream token by token (`ov-genai --cb`, `ov-runtime`, `sparse-moe`, and
`mock`) fill both histograms as described.

Timing is measured when tokens reach the client, not when the engine
produces them. With a single stream, the two are close. When several streams
share one engine, tokens can wait in a buffer, and that wait is included.
For streaming requests that use tool calls, the API buffers the whole
generation before responding, so `cascadia_http_request_duration_seconds`
covers the full generation for those requests.

## Engine metrics

Set once when the stage starts.

| Name | Type | Labels | Description |
|---|---|---|---|
| `cascadia_engine_model_load_duration_seconds` | gauge | `model`, `device` | Time to load weights and build the engine. Excludes connecting to other stages, which can include waiting for a peer to start. |
| `cascadia_engine_warmup_duration_seconds` | gauge | `model`, `device` | Engine warmup time. |

## Transport metrics

These cover traffic between stages for every engine. `kind` is `tensor` for
activation tensors (including headers) or `raw` for speculative-decoding
control messages.

| Name | Type | Labels | Description |
|---|---|---|---|
| `cascadia_transport_sent_bytes_total` | counter | `kind` | Bytes sent between stages, counting complete frames only. |
| `cascadia_transport_recv_bytes_total` | counter | `kind` | Bytes received between stages, counting complete frames only. |
| `cascadia_transport_send_seconds` | histogram | | Time to send a tensor frame, including flush. |
| `cascadia_transport_recv_payload_seconds` | histogram | | Time from receiving a tensor frame's header to receiving its full payload. Time spent waiting for a frame to start isn't included, since idle stages wait there between requests. |

Only complete frames are counted. Bytes from a send or receive that fails
partway, such as a timeout on a bad link, aren't recorded, and neither are
their durations. A link that stalls mid-frame produces no histogram sample at
all, so `recv_payload_seconds` keeps showing normal percentiles while
`recv_bytes_total` stops increasing. To catch a stalled link, alert on the
rate of `recv_bytes_total`, not on the histogram.

You can't compare `sent_bytes` on one stage with `recv_bytes` on the next,
because only rank 0 serves `/metrics`.

## Histogram buckets

Buckets are tuned for LLM serving and defined in
`crates/cascadia-metrics/src/lib.rs`:

- Time to first token: `0.05 0.1 0.25 0.5 1 2 5 10 30 60`
- Inter-token: `0.01 0.025 0.05 0.1 0.25 0.5 1 2 5 10`
- Generation and HTTP durations: `0.1 0.5 1 5 10 30 60 120 300 600`
- Transport frames: `0.0001 0.001 0.01 0.05 0.1 0.5 1`

## Example queries

```promql
# Request rate by endpoint
sum by (endpoint) (rate(cascadia_http_requests_total[5m]))

# p95 time to first token
histogram_quantile(0.95, sum by (le) (rate(cascadia_generation_ttft_seconds_bucket[5m])))

# Tokens per second across all pipelines
sum(rate(cascadia_tokens_generated_total[1m]))

# Capacity rejections (503s)
rate(cascadia_api_rejected_total{reason="capacity"}[5m])

# Throughput between stages
rate(cascadia_transport_sent_bytes_total[1m])
```

## Limitations

- Ranks other than 0 don't expose metrics.
- There are no per-layer expert metrics for `sparse-moe`.
- OTLP export and trace exemplars aren't supported.
