Skip to content
CascadiaCascadiaDocsOpen source
Start here

Requirements

Check whether your Intel machine and operating system can run Cascadia, and how large a model it can hold.

Cascadia is in alpha. It runs on any Intel GPU, NPU, or CPU that OpenVINO can see, but only some hardware is tested regularly.

Hardware Devices Status
Intel Lunar Lake AI PCs Arc iGPU (Xe2), NPU 4, CPU Tested
Intel Panther Lake AI PCs Arc B390 iGPU (Xe3), NPU 5, CPU Tested
Intel Arc Pro B70 (Battlemage) Discrete GPU Tested
Intel Arrow Lake AI PCs iGPU, NPU, CPU Expected to work
Other Intel Arc B-series discrete GPUs Discrete GPU Expected to work
Intel Arc A-series discrete GPUs Discrete GPU Roadmap
Intel Xeon servers, CPU only CPU Roadmap for dense models; used for the CPU sparse-moe engine
Any other machine None Mock engine only, for API and development testing

Tested means the Cascadia team runs it on that hardware. Expected to work means it uses the same OpenVINO drivers and devices as tested hardware, but isn’t tested regularly.

Windows is the most tested operating system; hardware tests run on Windows AI PCs. Linux (Ubuntu 22.04 or newer) is supported and needs Intel’s GPU runtime packages: see installation.

Terminal window
cascadia doctor

Doctor lists every device OpenVINO can reach, with its full name, so you can confirm it found your iGPU or discrete GPU. If it reports only a CPU, OpenVINO can’t see your GPU, even if your graphics driver works. Fix that before running models: inference falls back to the CPU silently and runs several times slower.

  • GPU (--device GPU) is the default choice: integrated Arc graphics on AI PCs, or a discrete Arc card.
  • NPU (--device NPU) needs a static-shape export and has its own constraints. Read NPU sharding before choosing it.
  • CPU works everywhere but is the slowest for dense models.

At INT4, model weights take about 0.55 GB per billion parameters: about 2 GB for a 4B model, 4 GB for 8B, and 17 GB for 32B. Add about 0.5 GB of OpenVINO workspace per stage, plus the KV cache, which grows with context length.

An iGPU shares system memory, so leave room for the operating system. If a model doesn’t fit on one machine, split it across several.

Ready-to-run models are on Hugging Face. To export another model, see Export a model.