Quickstart
Serve a real model on your Intel machine and send it your first request, with no build or export step.
Before you begin
Section titled “Before you begin”You need:
- An Intel machine on a supported hardware target, running Linux (Ubuntu 22.04 or newer) or Windows.
- The Hugging Face CLI,
hf. Install it before you start. - About 2 GB of free disk space for the model.
You don’t need Git, Rust, or Python. No Intel hardware? Try the API without hardware instead.
1. Download Cascadia
Section titled “1. Download Cascadia”Download the archive for your operating system from the latest GitHub release. Each archive includes the OpenVINO runtime.
Unpack the archive and move into its folder:
tar -xzf cascadia-*-linux-x86_64.tar.gzcd cascadia-*-linux-x86_64Cascadia needs the OpenCL loader to start, even on CPU:
sudo apt install ocl-icd-libopencl1GPU inference also needs Intel’s GPU runtime stack. Follow the GPU runtime instructions if you haven’t installed it yet.
Extract the .zip file, then open PowerShell in the folder that contains cascadia.exe.
Windows needs only a current Intel graphics driver. The GPU runtime ships inside the driver.
2. Check your hardware
Section titled “2. Check your hardware”./cascadia doctor.\cascadia.exe doctorLook for a GPU in the device list. If doctor reports only a CPU, OpenVINO can’t see your GPU. Fix that before you continue: see troubleshooting.
3. Download a model
Section titled “3. Download a model”Download Phi-4-mini Instruct, already exported for Cascadia and quantized to INT4 (2.0 GB):
hf download communitylabs/cascadia-phi-4-mini-int4 --local-dir ./phi-4-miniCascadia never downloads models while it serves them. run loads a local directory.
4. Start the server
Section titled “4. Start the server”./cascadia run ./phi-4-mini/int4/stages-1 \ --engine ov-runtime --device GPU --api 127.0.0.1:8000.\cascadia.exe run .\phi-4-mini\int4\stages-1 ` --engine ov-runtime --device GPU --api 127.0.0.1:8000Keep this terminal open. The first start compiles GPU kernels, so it is slower than later starts.
5. Send your first request
Section titled “5. Send your first request”In a second terminal, find the served model ID:
curl http://127.0.0.1:8000/v1/modelsInvoke-RestMethod http://127.0.0.1:8000/v1/modelsThen send a chat request, replacing <MODEL_ID> with the id from that response:
curl http://127.0.0.1:8000/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "<MODEL_ID>", "messages": [{"role": "user", "content": "Hello, Cascadia!"}] }'Invoke-RestMethod http://127.0.0.1:8000/v1/chat/completions ` -Method Post -ContentType 'application/json' ` -Body '{"model": "<MODEL_ID>", "messages": [{"role": "user", "content": "Hello, Cascadia!"}]}'The response is a standard OpenAI-style chat completion, generated on your own hardware.
What comes next?
Section titled “What comes next?”- Connect an existing application to the API.
- Serve your own model by exporting it from Hugging Face.
- Split a model across machines.