# Make a request

> Send requests to Cascadia's OpenAI-compatible API from curl, an SDK, or an existing tool.

Every Cascadia server, on one machine or several, exposes an OpenAI-compatible HTTP API. It implements chat completions, text completions, model listing, and SSE streaming, so most tools and SDKs built for OpenAI can use it.

## Point your client at Cascadia

In any client or tool that lets you set a custom base URL, use your Cascadia server’s `/v1` endpoint:

```text
http://localhost:8000/v1
```

Replace `localhost:8000` with the address you passed to `--api`. Use the model ID returned by `/v1/models`.

These examples use the mock server from [Try the API without hardware](/getting-started/try-without-hardware/).

## Make a chat request

```bash
curl http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "mock-model",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 128
  }'
```

For real inference, replace `mock-model` with the ID returned by `/v1/models`.

## Stream the response

Add `"stream": true` to the request body and use `curl -N` to display server-sent events as they arrive. Streaming ends with a `data: [DONE]` event.

```bash
curl -N http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"mock-model","messages":[{"role":"user","content":"Hello!"}],"stream":true}'
```

## Compatibility boundaries

This is a compatible inference surface, not the complete OpenAI platform. Sampling, log probabilities, tool calls, and reasoning behavior depend on the selected model and engine. Check the [HTTP API reference](/reference/api/) and [engine notes](/reference/cli/#engines) for details.

The open-source server does not authenticate API keys. If your client requires an API key value, a placeholder does not add access control. For access outside loopback, follow the [security guide](/guides/security/).
