Skip to content
CascadiaCascadiaDocsOpen source
Start here

Quickstart

Serve a real model on your Intel machine and send it your first request, with no build or export step.

You need:

  • An Intel machine on a supported hardware target, running Linux (Ubuntu 22.04 or newer) or Windows.
  • The Hugging Face CLI, hf. Install it before you start.
  • About 2 GB of free disk space for the model.

You don’t need Git, Rust, or Python. No Intel hardware? Try the API without hardware instead.

Download the archive for your operating system from the latest GitHub release. Each archive includes the OpenVINO runtime.

Unpack the archive and move into its folder:

Terminal window
tar -xzf cascadia-*-linux-x86_64.tar.gz
cd cascadia-*-linux-x86_64

Cascadia needs the OpenCL loader to start, even on CPU:

Terminal window
sudo apt install ocl-icd-libopencl1

GPU inference also needs Intel’s GPU runtime stack. Follow the GPU runtime instructions if you haven’t installed it yet.

Terminal window
./cascadia doctor

Look for a GPU in the device list. If doctor reports only a CPU, OpenVINO can’t see your GPU. Fix that before you continue: see troubleshooting.

Download Phi-4-mini Instruct, already exported for Cascadia and quantized to INT4 (2.0 GB):

Terminal window
hf download communitylabs/cascadia-phi-4-mini-int4 --local-dir ./phi-4-mini

Cascadia never downloads models while it serves them. run loads a local directory.

Terminal window
./cascadia run ./phi-4-mini/int4/stages-1 \
--engine ov-runtime --device GPU --api 127.0.0.1:8000

Keep this terminal open. The first start compiles GPU kernels, so it is slower than later starts.

In a second terminal, find the served model ID:

Terminal window
curl http://127.0.0.1:8000/v1/models

Then send a chat request, replacing <MODEL_ID> with the id from that response:

Terminal window
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "<MODEL_ID>",
"messages": [{"role": "user", "content": "Hello, Cascadia!"}]
}'

The response is a standard OpenAI-style chat completion, generated on your own hardware.