Skip to content
CascadiaCascadiaDocsOpen source
Inference

Single-machine inference

An open model, your Intel machine, and a local API endpoint.

Cascadia serves models from a local directory through an OpenAI-compatible API. cascadia run is the simplest path: it starts a single-stage worker and exposes an API on port 8000.

Install a release bundle for Linux or Windows, then run:

Terminal window
cascadia doctor

From an unpacked bundle, use ./cascadia or .\cascadia.exe, or add its directory to your PATH. On Linux, GPU inference also needs the Intel GPU runtime stack described in the installation guide.

Download a ready-to-run model, or export your own with --num-stages 1. The examples below use a one-stage Llama 3.1 8B export in ~/cascadia/llama-8b-1stage.

run loads a local directory; it never downloads models.

Terminal window
cascadia run ~/cascadia/llama-8b-1stage \
--engine ov-runtime --device GPU --api 127.0.0.1:8000

Use ov-runtime for a Cascadia shard tree. If you already have a whole-model OpenVINO IR directory, use the default ov-genai engine instead:

Terminal window
cascadia run ~/models/llama-3.1-8b-int4-ov --api 127.0.0.1:8000

Find the served model ID with curl http://localhost:8000/v1/models, then send requests to /v1/chat/completions. See Make a request for examples.

Ready for a larger model? Split it across machines.