Skip to content
CascadiaCascadiaDocsOpen source
Inference

Make a request

Send requests to Cascadia's OpenAI-compatible API from curl, an SDK, or an existing tool.

Every Cascadia server, on one machine or several, exposes an OpenAI-compatible HTTP API. It implements chat completions, text completions, model listing, and SSE streaming, so most tools and SDKs built for OpenAI can use it.

In any client or tool that lets you set a custom base URL, use your Cascadia server’s /v1 endpoint:

http://localhost:8000/v1

Replace localhost:8000 with the address you passed to --api. Use the model ID returned by /v1/models.

These examples use the mock server from Try the API without hardware.

Terminal window
curl http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "mock-model",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 128
}'

For real inference, replace mock-model with the ID returned by /v1/models.

Add "stream": true to the request body and use curl -N to display server-sent events as they arrive. Streaming ends with a data: [DONE] event.

Terminal window
curl -N http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"mock-model","messages":[{"role":"user","content":"Hello!"}],"stream":true}'

This is a compatible inference surface, not the complete OpenAI platform. Sampling, log probabilities, tool calls, and reasoning behavior depend on the selected model and engine. Check the HTTP API reference and engine notes for details.

The open-source server does not authenticate API keys. If your client requires an API key value, a placeholder does not add access control. For access outside loopback, follow the security guide.