Skip to content
CascadiaCascadiaDocsOpen source
Reference

HTTP API

An OpenAI-compatible inference interface, served by your Cascadia worker.

The standard local endpoint is http://localhost:8000. Only the first stage of a distributed pipeline serves the API.

Method Path Purpose
GET /health Check pipeline readiness
GET /metrics Read Prometheus metrics
GET /v1/models List the configured model
POST /v1/chat/completions Generate a chat completion, optionally streamed
POST /v1/completions Generate a text completion
POST /v1/cancel/:task_id Cancel an inference task
Terminal window
curl http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "mock-model",
"messages": [{"role": "user", "content": "Hello, Cascadia!"}],
"max_tokens": 128,
"stream": false
}'

This example uses the mock server. For real inference, use the configured model ID returned by /v1/models.

Field Type Description
model string Configured model ID
messages array Chat messages with roles and content
max_tokens integer Maximum generated tokens
temperature number Sampling temperature
stream boolean Return server-sent events when true
top_p, top_k number Sampling controls, depending on engine support
stop string or array Stop sequences
seed integer Sampling seed, depending on engine support

Additional tool, log-probability, and reasoning fields are defined in the upstream request schema. Behavior depends on the model and engine.

Set stream to true to receive SSE chunks. Clients should read the stream until data: [DONE]. Use curl -N to disable output buffering when testing.

Terminal window
curl http://localhost:8000/health
curl http://localhost:8000/metrics

Health reflects pipeline readiness. The metrics endpoint exposes request and inference measurements; see the metric inventory.

The HTTP server is plaintext and unauthenticated. Bind to loopback for local development, or place TLS and authentication at a reverse proxy for wider access. See security.

This page summarizes the routes in cascadia-api.