Requirements
Check whether your Intel machine and operating system can run Cascadia, and how large a model it can hold.
Supported hardware
Section titled “Supported hardware”Cascadia is in alpha. It runs on any Intel GPU, NPU, or CPU that OpenVINO can see, but only some hardware is tested regularly.
| Hardware | Devices | Status |
|---|---|---|
| Intel Lunar Lake AI PCs | Arc iGPU (Xe2), NPU 4, CPU | Tested |
| Intel Panther Lake AI PCs | Arc B390 iGPU (Xe3), NPU 5, CPU | Tested |
| Intel Arc Pro B70 (Battlemage) | Discrete GPU | Tested |
| Intel Arrow Lake AI PCs | iGPU, NPU, CPU | Expected to work |
| Other Intel Arc B-series discrete GPUs | Discrete GPU | Expected to work |
| Intel Arc A-series discrete GPUs | Discrete GPU | Roadmap |
| Intel Xeon servers, CPU only | CPU | Roadmap for dense models; used for the CPU sparse-moe engine |
| Any other machine | None | Mock engine only, for API and development testing |
Tested means the Cascadia team runs it on that hardware. Expected to work means it uses the same OpenVINO drivers and devices as tested hardware, but isn’t tested regularly.
Operating systems
Section titled “Operating systems”Windows is the most tested operating system; hardware tests run on Windows AI PCs. Linux (Ubuntu 22.04 or newer) is supported and needs Intel’s GPU runtime packages: see installation.
Check your device
Section titled “Check your device”cascadia doctorDoctor lists every device OpenVINO can reach, with its full name, so you can confirm it found your iGPU or discrete GPU. If it reports only a CPU, OpenVINO can’t see your GPU, even if your graphics driver works. Fix that before running models: inference falls back to the CPU silently and runs several times slower.
Choose a device
Section titled “Choose a device”- GPU (
--device GPU) is the default choice: integrated Arc graphics on AI PCs, or a discrete Arc card. - NPU (
--device NPU) needs a static-shape export and has its own constraints. Read NPU sharding before choosing it. - CPU works everywhere but is the slowest for dense models.
How large a model can I run?
Section titled “How large a model can I run?”At INT4, model weights take about 0.55 GB per billion parameters: about 2 GB for a 4B model, 4 GB for 8B, and 17 GB for 32B. Add about 0.5 GB of OpenVINO workspace per stage, plus the KV cache, which grows with context length.
An iGPU shares system memory, so leave room for the operating system. If a model doesn’t fit on one machine, split it across several.
Ready-to-run models are on Hugging Face. To export another model, see Export a model.