Skip to content
CascadiaCascadiaDocsOpen source
Start here

Installation

Install a prebuilt release, or build Cascadia from source.

There are two ways to install Cascadia:

Option Use it to You need
Prebuilt release (recommended) Run models on Intel hardware An Intel graphics driver; on Linux, the GPU runtime stack
Build from source Develop Cascadia, test it in CI, or build against a specific OpenVINO version Rust; for real inference, also a C++ toolchain and the OpenVINO GenAI SDK

Whichever you choose, finish by running cascadia doctor. It checks your setup and tells you whether OpenVINO can actually see your GPU, which otherwise fails silently.

Each GitHub release ships self-contained archives for Linux and Windows. The OpenVINO runtime is included, so there is no SDK to install and no INTEL_OPENVINO_DIR to set.

Download cascadia-<version>-linux-x86_64.tar.gz, then unpack it:

Terminal window
tar -xzf cascadia-*-linux-x86_64.tar.gz
cd cascadia-*-linux-x86_64

Bundled libraries load from lib/ beside the binary. The bundle needs glibc 2.35 or newer (Ubuntu 22.04 or newer).

The binary is not added to your PATH. Run it from the unpacked folder (./cascadia or .\cascadia.exe), or add that folder to your PATH.

This is the step most people miss. OpenVINO GPU inference needs the Intel Compute Runtime, the OpenCL ICD, and the Level Zero loader, and your user must be in the render group. Without these, OpenVINO silently sees only the CPU, even with a working driver and a healthy clinfo.

Install them from Intel’s graphics repository (Ubuntu 22.04 and 24.04):

Terminal window
sudo apt-get install -y ca-certificates gnupg wget
wget -qO- https://repositories.intel.com/gpu/intel-graphics.key \
| sudo gpg --yes --dearmor -o /usr/share/keyrings/intel-graphics.gpg
. /etc/os-release # noble on 24.04, jammy on 22.04
echo "deb [arch=amd64 signed-by=/usr/share/keyrings/intel-graphics.gpg] \
https://repositories.intel.com/gpu/ubuntu $UBUNTU_CODENAME unified" \
| sudo tee /etc/apt/sources.list.d/intel-gpu.list
sudo apt-get update
sudo apt-get install -y ocl-icd-libopencl1 intel-opencl-icd libze-intel-gpu1 libze1
sudo usermod -a -G render "$USER" # then log out and back in

Use Intel’s repository, not the distro packages. Ubuntu 24.04 ships Compute Runtime 23.43, which predates Lunar Lake and Arc B-series; on that hardware OpenVINO may see no GPU at all. Intel publishes the repository for Ubuntu only. On other distributions, install the equivalent packages by hand.

ocl-icd-libopencl1 is required even for CPU-only use: the bundled OpenVINO library imports libOpenCL.so.1, so without it the binary does not start.

Terminal window
./cascadia doctor

Doctor should list a GPU device, not just the CPU. It also prints the bundled OpenVINO GenAI version.

Next: serve your first model.

Choose the build that matches what you need:

Build What works Platforms You need
Stub The mock engine: the full API, transport, and multi-stage plumbing, without real inference Linux, Windows, macOS Rust
OpenVINO Real inference on Intel hardware Linux, Windows Rust, a C++ toolchain, the OpenVINO GenAI SDK, and on Linux the GPU runtime

Install Rust 1.89 or newer with rustup, then:

Terminal window
rustup default stable
git clone https://github.com/labscommunity/cascadia.git
cd cascadia
Terminal window
cargo build --release -p cascadia
./target/release/cascadia doctor

The OpenVINO-backed engines return a clean runtime error in this build. Use --engine mock; see Try the API without hardware.

The binary is statically linked apart from the OpenVINO libraries. Copy it together with:

  • Linux: $INTEL_OPENVINO_DIR/runtime/lib/intel64/ and runtime/3rdparty/tbb/lib/
  • Windows: runtime/bin/intel64/Release/ and runtime/3rdparty/tbb/bin/

TBB ships beside the runtime, not inside it, and OpenVINO imports it.

The browser dashboard served at / by --api workers is compiled in with the dashboard-embed feature. Build the dashboard first (Node 20 or newer):

Terminal window
cd crates/cascadia-dashboard/web && npm ci && npm run build && cd -
cargo build --release -p cascadia --features dashboard-embed # stub
cargo build --release -p cascadia --features openvino,dashboard-embed # OpenVINO

Without the feature, / serves a pointer page and only the JSON API is live. Release archives newer than v0.1.8 include the dashboard.

cascadia shard runs a bundled Python exporter to turn a Hugging Face model into per-stage OpenVINO IR. Python is needed only for export, not for inference, and not on workers.

  • From a source checkout: pip install -r tools/requirements.txt
  • From a prebuilt release: run cascadia doctor, which prints the exact pinned pip install command for that binary.

nncf is optional and enables INT4 quantization; without it, sharding falls back to FP16. Exporting a whole-model IR for --engine ov-genai uses Intel’s exporter instead: pip install "optimum-intel[openvino]".

The repository’s Dockerfile bundles the OpenVINO and Level Zero stack, so you can skip the host install. GPU inference needs the host GPU passed in (--device /dev/dri) and a matching host driver. See the comments at the top of the Dockerfile.

Every release archive carries one OpenVINO GenAI version: the newest stable release validated on the Cascadia fleet. cascadia doctor prints it. OpenVINO’s C++ ABI is not stable across releases, so you can’t swap the runtime underneath a binary. To use another version:

  • Variant archives. A release may include extra archives named cascadia-<version>-<os>-x86_64-ov<openvino-version>.*, built against a newer or pre-release OpenVINO.
  • Build from source. scripts/ov_sdk.py fetch <version> accepts any published SDK: stable (2026.4.1.0), beta (2026.5.0.0beta1), or nightly (2026.5.0.0.dev20260925).