# Single-machine inference

> An open model, your Intel machine, and a local API endpoint.

## Start with one machine

Cascadia serves models from a local directory through an OpenAI-compatible API. `cascadia run` is the simplest path: it starts a single-stage worker and exposes an API on port 8000.

## Check the prerequisites

Install a [release bundle](/getting-started/installation/) for Linux or Windows, then run:

```bash
cascadia doctor
```

From an unpacked bundle, use `./cascadia` or `.\cascadia.exe`, or add its directory to your PATH. On Linux, GPU inference also needs the Intel GPU runtime stack described in the installation guide.

## Get a model

Download a [ready-to-run model](https://huggingface.co/communitylabs), or [export your own](/features/export-a-model/) with `--num-stages 1`. The examples below use a one-stage Llama 3.1 8B export in `~/cascadia/llama-8b-1stage`.

`run` loads a local directory; it never downloads models.

## Serve the model

```bash
cascadia run ~/cascadia/llama-8b-1stage \
  --engine ov-runtime --device GPU --api 127.0.0.1:8000
```

Use `ov-runtime` for a Cascadia shard tree. If you already have a whole-model OpenVINO IR directory, use the default `ov-genai` engine instead:

```bash
cascadia run ~/models/llama-3.1-8b-int4-ov --api 127.0.0.1:8000
```

## Connect your application

Find the served model ID with `curl http://localhost:8000/v1/models`, then send requests to `/v1/chat/completions`. See [Make a request](/features/make-a-request/) for examples.

Ready for a larger model? [Split it across machines](/features/multi-machine/).
