# Export a model

> Turn a Hugging Face model into quantized OpenVINO stages for one machine or several.

Cascadia serves models from a local directory of OpenVINO stages. `cascadia shard` creates that directory from a Hugging Face model: one stage to serve on a single machine, or several to split the model across machines. Export once, then copy the output to each machine that serves it.

:::tip[Skip this step with a ready-to-run model]
The [ready-to-run models](https://huggingface.co/communitylabs) are already exported and quantized. Download one with `hf download` and go straight to [single-machine](/features/single-machine/) or [multi-machine](/features/multi-machine/) inference.
:::

## Before you start

- **Check that your model can be exported.** Support varies by architecture. See the [supported architectures](https://github.com/labscommunity/cascadia/blob/main/docs/SHARDING.md#supported-architectures).
- **Install the export dependencies** on the machine doing the export. Follow [Python, for exporting models](/getting-started/installation/#python-for-exporting-models). Python is needed only for export, not on the machines that serve the model.

## Export the model

Set `--num-stages` to the number of machines that will serve the model. For one machine:

```bash
cascadia shard --model unsloth/Meta-Llama-3.1-8B-Instruct \
  --output-dir ~/cascadia/llama-8b-1stage \
  --num-stages 1 --quantization int4
```

For a two-machine pipeline:

```bash
cascadia shard --model unsloth/Meta-Llama-3.1-8B-Instruct \
  --output-dir ~/cascadia/llama-8b-2stage \
  --num-stages 2 --quantization int4
```

`shard` is the only command that downloads from Hugging Face. `run` and `worker` load the exported directory and never download models.

## Understand the output

```text
llama-8b-2stage/
├── pipeline_config.json
├── tokenizer/
├── stage_0/
│   ├── openvino_model.xml
│   ├── openvino_model.bin
│   └── stage_config.json
└── stage_1/
    ├── openvino_model.xml
    ├── openvino_model.bin
    └── stage_config.json
```

The first stage holds the embeddings and the last holds the final normalization and language-model head. Copy the whole tree to each machine; each worker loads only its own stage.

## Match the export to your hardware

- **Stages:** use the fewest that fit in memory. Each extra stage adds a network hop per token. See [How large a model can I run?](/getting-started/requirements/#how-large-a-model-can-i-run)
- **Layer split:** stages get an even split by default. `--layer-split` puts more layers on a machine with more memory.
- **Quantization:** `--quantization` selects `int4` (default), `int4-asym`, `int8`, or `fp16`.
- **NPU:** NPU workers need a static-shape export with `--target npu`. Read the [NPU requirements](https://github.com/labscommunity/cascadia/blob/main/docs/NPU_SHARDING.md).

All flags are listed under [`cascadia shard`](/reference/cli/#cascadia-shard).

## Next steps

- One stage: [serve it on one machine](/features/single-machine/).
- Several stages: [start a pipeline across machines](/features/multi-machine/).
