# Multi-machine inference

> Give a model room to run by sharing it across your Intel machines.

## How the pipeline works

Cascadia divides a model into stages. Each worker holds a contiguous set of transformer layers. The first stage receives the API request, intermediate activations travel over TCP, and the last stage produces tokens that return through the chain.

```text
Your application
   │  ▲
   │  │  OpenAI-compatible HTTP
   ▼  │
┌─────────────────┐   activations   ┌─────────────────┐
│ Node A · rank 0 │ ──────────────▶ │ Node B · rank 1 │
│ First layers    │                 │ Final layers    │
│ API :8000       │ ◀────────────── │ Relay :9100     │
└─────────────────┘     tokens      └─────────────────┘
```

## 1. Get a two-stage model

On both nodes, download the [ready-to-run](https://huggingface.co/communitylabs) Llama 3.1 8B, which includes a two-stage preset:

```bash
hf download communitylabs/Llama-3.1-8B-cascadia-int4 \
  --local-dir ~/cascadia/llama-8b
```

The two-stage model is in `~/cascadia/llama-8b/int4/stages-2`. To use a different model, [export it](/features/export-a-model/) with `--num-stages 2` and copy the output to both nodes.

## 2. Start the last stage first

On Node B:

```bash
cascadia worker --rank 1 --total 2 --engine ov-runtime --device GPU \
  --model ~/cascadia/llama-8b/int4/stages-2 --listen :9100
```

## 3. Find Node B’s address

Running workers advertise themselves over mDNS on `_cascadia._tcp.local.`. On Node A, list the workers on your LAN:

```bash
cascadia discover
```

The output includes each peer’s address. If Node B doesn’t appear, check that its worker is running, both machines are on the same LAN, and the network allows multicast. If mDNS is blocked, use Node B’s known address directly.

Discovery only finds addresses. It doesn’t connect stages or authenticate peers. See [`cascadia discover`](/reference/cli/#cascadia-discover) for namespaces and timeouts.

## 4. Start the first stage

On Node A, replace `10.0.0.2` with Node B’s address:

```bash
cascadia worker --rank 0 --total 2 --engine ov-runtime --device GPU \
  --model ~/cascadia/llama-8b/int4/stages-2 \
  --next 10.0.0.2:9100 --api :8000
```

Your application connects to Node A. The request format is the same as single-machine inference.

## What to know

- Placement is manual: set each worker’s rank, total, and downstream address. Discovery doesn’t form a pipeline automatically.
- The API and relay are plaintext and unauthenticated. Use a trusted network and read the [security model](/guides/security/).
- More stages expand capacity; network transfer also adds latency. Measure your workload.

For speculative decoding across stages, see the [ov-dist-spec engine](https://github.com/labscommunity/cascadia/blob/main/docs/engines/ov-dist-spec.md).
