# Quickstart

> Serve a real model on your Intel machine and send it your first request, with no build or export step.

## Before you begin

You need:

- An Intel machine on a [supported hardware target](/getting-started/requirements/), running Linux (Ubuntu 22.04 or newer) or Windows.
- The Hugging Face CLI, `hf`. [Install it](https://huggingface.co/docs/huggingface_hub/en/guides/cli) before you start.
- About 2 GB of free disk space for the model.

You don't need Git, Rust, or Python. No Intel hardware? [Try the API without hardware](/getting-started/try-without-hardware/) instead.

## 1. Download Cascadia

Download the archive for your operating system from the [latest GitHub release](https://github.com/labscommunity/cascadia/releases/latest). Each archive includes the OpenVINO runtime.

**Linux**

Unpack the archive and move into its folder:

```bash
tar -xzf cascadia-*-linux-x86_64.tar.gz
cd cascadia-*-linux-x86_64
```

Cascadia needs the OpenCL loader to start, even on CPU:

```bash
sudo apt install ocl-icd-libopencl1
```

GPU inference also needs Intel's GPU runtime stack. Follow the [GPU runtime instructions](/getting-started/installation/#2-install-the-gpu-runtime) if you haven't installed it yet.

**Windows**

Extract the `.zip` file, then open PowerShell in the folder that contains `cascadia.exe`.

Windows needs only a current Intel graphics driver. The GPU runtime ships inside the driver.

## 2. Check your hardware

**Linux**

```bash
./cascadia doctor
```

**Windows**

```powershell
.\cascadia.exe doctor
```

Look for a GPU in the device list. If doctor reports only a CPU, OpenVINO can't see your GPU. Fix that before you continue: see [troubleshooting](/reference/troubleshooting/).

## 3. Download a model

Download Phi-4-mini Instruct, already exported for Cascadia and quantized to INT4 (2.0 GB):

```bash
hf download communitylabs/cascadia-phi-4-mini-int4 --local-dir ./phi-4-mini
```

Cascadia never downloads models while it serves them. `run` loads a local directory.

## 4. Start the server

**Linux**

```bash
./cascadia run ./phi-4-mini/int4/stages-1 \
  --engine ov-runtime --device GPU --api 127.0.0.1:8000
```

**Windows**

```powershell
.\cascadia.exe run .\phi-4-mini\int4\stages-1 `
  --engine ov-runtime --device GPU --api 127.0.0.1:8000
```

Keep this terminal open. The first start compiles GPU kernels, so it is slower than later starts.

## 5. Send your first request

In a second terminal, find the served model ID:

**Linux**

```bash
curl http://127.0.0.1:8000/v1/models
```

**Windows**

```powershell
Invoke-RestMethod http://127.0.0.1:8000/v1/models
```

Then send a chat request, replacing `<MODEL_ID>` with the `id` from that response:

**Linux**

```bash
curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "<MODEL_ID>",
    "messages": [{"role": "user", "content": "Hello, Cascadia!"}]
  }'
```

**Windows**

```powershell
Invoke-RestMethod http://127.0.0.1:8000/v1/chat/completions `
  -Method Post -ContentType 'application/json' `
  -Body '{"model": "<MODEL_ID>", "messages": [{"role": "user", "content": "Hello, Cascadia!"}]}'
```

The response is a standard OpenAI-style chat completion, generated on your own hardware.

## What comes next?

- [Connect an existing application](/features/make-a-request/) to the API.
- [Serve your own model](/features/single-machine/) by exporting it from Hugging Face.
- [Split a model across machines](/features/multi-machine/).
