Export a model
Turn a Hugging Face model into quantized OpenVINO stages for one machine or several.
Cascadia serves models from a local directory of OpenVINO stages. cascadia shard creates that directory from a Hugging Face model: one stage to serve on a single machine, or several to split the model across machines. Export once, then copy the output to each machine that serves it.
Before you start
Section titled “Before you start”- Check that your model can be exported. Support varies by architecture. See the supported architectures.
- Install the export dependencies on the machine doing the export. Follow Python, for exporting models. Python is needed only for export, not on the machines that serve the model.
Export the model
Section titled “Export the model”Set --num-stages to the number of machines that will serve the model. For one machine:
cascadia shard --model unsloth/Meta-Llama-3.1-8B-Instruct \ --output-dir ~/cascadia/llama-8b-1stage \ --num-stages 1 --quantization int4For a two-machine pipeline:
cascadia shard --model unsloth/Meta-Llama-3.1-8B-Instruct \ --output-dir ~/cascadia/llama-8b-2stage \ --num-stages 2 --quantization int4shard is the only command that downloads from Hugging Face. run and worker load the exported directory and never download models.
Understand the output
Section titled “Understand the output”llama-8b-2stage/├── pipeline_config.json├── tokenizer/├── stage_0/│ ├── openvino_model.xml│ ├── openvino_model.bin│ └── stage_config.json└── stage_1/ ├── openvino_model.xml ├── openvino_model.bin └── stage_config.jsonThe first stage holds the embeddings and the last holds the final normalization and language-model head. Copy the whole tree to each machine; each worker loads only its own stage.
Match the export to your hardware
Section titled “Match the export to your hardware”- Stages: use the fewest that fit in memory. Each extra stage adds a network hop per token. See How large a model can I run?
- Layer split: stages get an even split by default.
--layer-splitputs more layers on a machine with more memory. - Quantization:
--quantizationselectsint4(default),int4-asym,int8, orfp16. - NPU: NPU workers need a static-shape export with
--target npu. Read the NPU requirements.
All flags are listed under cascadia shard.
Next steps
Section titled “Next steps”- One stage: serve it on one machine.
- Several stages: start a pipeline across machines.