Single-machine inference
An open model, your Intel machine, and a local API endpoint.
Start with one machine
Section titled “Start with one machine”Cascadia serves models from a local directory through an OpenAI-compatible API. cascadia run is the simplest path: it starts a single-stage worker and exposes an API on port 8000.
Check the prerequisites
Section titled “Check the prerequisites”Install a release bundle for Linux or Windows, then run:
cascadia doctorFrom an unpacked bundle, use ./cascadia or .\cascadia.exe, or add its directory to your PATH. On Linux, GPU inference also needs the Intel GPU runtime stack described in the installation guide.
Get a model
Section titled “Get a model”Download a ready-to-run model, or export your own with --num-stages 1. The examples below use a one-stage Llama 3.1 8B export in ~/cascadia/llama-8b-1stage.
run loads a local directory; it never downloads models.
Serve the model
Section titled “Serve the model”cascadia run ~/cascadia/llama-8b-1stage \ --engine ov-runtime --device GPU --api 127.0.0.1:8000Use ov-runtime for a Cascadia shard tree. If you already have a whole-model OpenVINO IR directory, use the default ov-genai engine instead:
cascadia run ~/models/llama-3.1-8b-int4-ov --api 127.0.0.1:8000Connect your application
Section titled “Connect your application”Find the served model ID with curl http://localhost:8000/v1/models, then send requests to /v1/chat/completions. See Make a request for examples.
Ready for a larger model? Split it across machines.