Multi-machine inference
Give a model room to run by sharing it across your Intel machines.
How the pipeline works
Section titled “How the pipeline works”Cascadia divides a model into stages. Each worker holds a contiguous set of transformer layers. The first stage receives the API request, intermediate activations travel over TCP, and the last stage produces tokens that return through the chain.
Your application │ ▲ │ │ OpenAI-compatible HTTP ▼ │┌─────────────────┐ activations ┌─────────────────┐│ Node A · rank 0 │ ──────────────▶ │ Node B · rank 1 ││ First layers │ │ Final layers ││ API :8000 │ ◀────────────── │ Relay :9100 │└─────────────────┘ tokens └─────────────────┘1. Get a two-stage model
Section titled “1. Get a two-stage model”On both nodes, download the ready-to-run Llama 3.1 8B, which includes a two-stage preset:
hf download communitylabs/Llama-3.1-8B-cascadia-int4 \ --local-dir ~/cascadia/llama-8bThe two-stage model is in ~/cascadia/llama-8b/int4/stages-2. To use a different model, export it with --num-stages 2 and copy the output to both nodes.
2. Start the last stage first
Section titled “2. Start the last stage first”On Node B:
cascadia worker --rank 1 --total 2 --engine ov-runtime --device GPU \ --model ~/cascadia/llama-8b/int4/stages-2 --listen :91003. Find Node B’s address
Section titled “3. Find Node B’s address”Running workers advertise themselves over mDNS on _cascadia._tcp.local.. On Node A, list the workers on your LAN:
cascadia discoverThe output includes each peer’s address. If Node B doesn’t appear, check that its worker is running, both machines are on the same LAN, and the network allows multicast. If mDNS is blocked, use Node B’s known address directly.
Discovery only finds addresses. It doesn’t connect stages or authenticate peers. See cascadia discover for namespaces and timeouts.
4. Start the first stage
Section titled “4. Start the first stage”On Node A, replace 10.0.0.2 with Node B’s address:
cascadia worker --rank 0 --total 2 --engine ov-runtime --device GPU \ --model ~/cascadia/llama-8b/int4/stages-2 \ --next 10.0.0.2:9100 --api :8000Your application connects to Node A. The request format is the same as single-machine inference.
What to know
Section titled “What to know”- Placement is manual: set each worker’s rank, total, and downstream address. Discovery doesn’t form a pipeline automatically.
- The API and relay are plaintext and unauthenticated. Use a trusted network and read the security model.
- More stages expand capacity; network transfer also adds latency. Measure your workload.
For speculative decoding across stages, see the ov-dist-spec engine.