Skip to content
CascadiaCascadiaDocsOpen source
Inference

Multi-machine inference

Give a model room to run by sharing it across your Intel machines.

Cascadia divides a model into stages. Each worker holds a contiguous set of transformer layers. The first stage receives the API request, intermediate activations travel over TCP, and the last stage produces tokens that return through the chain.

Your application
│ ▲
│ │ OpenAI-compatible HTTP
▼ │
┌─────────────────┐ activations ┌─────────────────┐
│ Node A · rank 0 │ ──────────────▶ │ Node B · rank 1 │
│ First layers │ │ Final layers │
│ API :8000 │ ◀────────────── │ Relay :9100 │
└─────────────────┘ tokens └─────────────────┘

On both nodes, download the ready-to-run Llama 3.1 8B, which includes a two-stage preset:

Terminal window
hf download communitylabs/Llama-3.1-8B-cascadia-int4 \
--local-dir ~/cascadia/llama-8b

The two-stage model is in ~/cascadia/llama-8b/int4/stages-2. To use a different model, export it with --num-stages 2 and copy the output to both nodes.

On Node B:

Terminal window
cascadia worker --rank 1 --total 2 --engine ov-runtime --device GPU \
--model ~/cascadia/llama-8b/int4/stages-2 --listen :9100

Running workers advertise themselves over mDNS on _cascadia._tcp.local.. On Node A, list the workers on your LAN:

Terminal window
cascadia discover

The output includes each peer’s address. If Node B doesn’t appear, check that its worker is running, both machines are on the same LAN, and the network allows multicast. If mDNS is blocked, use Node B’s known address directly.

Discovery only finds addresses. It doesn’t connect stages or authenticate peers. See cascadia discover for namespaces and timeouts.

On Node A, replace 10.0.0.2 with Node B’s address:

Terminal window
cascadia worker --rank 0 --total 2 --engine ov-runtime --device GPU \
--model ~/cascadia/llama-8b/int4/stages-2 \
--next 10.0.0.2:9100 --api :8000

Your application connects to Node A. The request format is the same as single-machine inference.

  • Placement is manual: set each worker’s rank, total, and downstream address. Discovery doesn’t form a pipeline automatically.
  • The API and relay are plaintext and unauthenticated. Use a trusted network and read the security model.
  • More stages expand capacity; network transfer also adds latency. Measure your workload.

For speculative decoding across stages, see the ov-dist-spec engine.