Skip to content
CascadiaCascadiaDocsOpen source
Reference

Rust crates

What each crate in the Cascadia Rust workspace does.

Cascadia is a Cargo workspace at the repo root. Each crate has a single responsibility and a stable interface; engines and discovery backends are swappable.

OpenAI-compatible HTTP server (axum). Routes: /health, /v1/models, /v1/chat/completions (non-streaming + SSE streaming), /v1/cancel/<task_id>. Backpressure via a concurrent-request semaphore (default 16); request body cap and rendered-prompt cap (both --api-max-body-mb, default 1 MiB, on every engine) enforce 413 / 503 responses on oversized or over-capacity input.

Per-stage Runner. Connects upstream + downstream transports, loads weights, builds the engine, warms it up, and exposes submit / generate / cancel. Concurrent-safe — multiple generate() callers share one engine through a Mutex; chunks for other tasks emitted during one caller’s step() are buffered for their owners.

Two trait definitions — the plugin seam:

  • Engine: warmup, submit, step, cancel, close. submit returns EngineError::QueueFull when the per-engine pending cap is reached.
  • Builder: configure_listen, connect, load, build, close.

Five engines:

  • ov-genai — single-stage openvino_genai.LLMPipeline via the C++ FFI shim. FastDraft + Prompt Lookup variants.
  • ov-runtime — multi-stage stateful KV cache. Pre-exported per-stage v3+ shards; each stage owns its layer range and runs SDPA attention with internal RoPE.
  • ov-dist-spec — multi-stage spec decode with mask-based KV-cache rewind on rejected drafts. v5 shards (canonical optimum-style inputs).
  • gemma4 — Gemma 4 multi-stage: per-layer-type attention, KV-sharing, per-layer-input embeddings. gemma4_cached_v1.x shards.
  • qwen35 (alias qwen36-moe) — Qwen3.5-family staged chain (GatedDeltaNet; Qwen3.6 MoE or dense Qwen3.8) from qwen3_5* IR-surgery shards; single-box or N-rank pipeline; in-process prefix cache. See architectures/qwen36-moe-support.md.

Deterministic word-echo engine — splits the prompt and emits one word per step(). Used by API / runner / CLI tests.

CPU-targeted sparse mixture-of-experts engine (Kimi K2.6-style models, MiniMax-M2). Runs attention/norm shells natively in Rust (default; OV IR shells are an optional backend) and dispatches only the top-k experts the router selects each step. Experts execute as per-(layer, expert) OV IRs by default, or through the cascadia-int4-gemm AVX-512 kernels against packed int4 weight binaries (int4_bin backend).

Hand-rolled AVX-512 INT4 GEMM kernels for the sparse-MoE expert path — group-32 symmetric quantization with bf16 scales, matching the compressed-tensors on-disk format.

Dashboard HTTP routes (/api/topology, /api/stats) and an embedded Vite SPA (behind the embed-spa feature) for visualizing a cluster; without the feature, / serves a built-in pointer page explaining how to enable the UI. Kept separate from cascadia-api so the OpenAI surface doesn’t grow a topology dependency or bundled static assets.

C++ FFI shim wrapping openvino-genai. extern "C" only; every entry point catches ... so a C++ exception cannot unwind into Rust UB. Stub mode (no link) is the default for dev / CI; --features openvino links against the real OV GenAI 2026.2+ SDK.

Engine-agnostic core types (serde-serializable): generation tasks and chunks, shard descriptions, peer layout. Only serde + thiserror as dependencies, so downstream crates share vocabulary without version-lockstep.

TCP activation relay between pipeline stages. Wire format: 20-byte header (payload_len, dtype, dim0, dim1, dim2) then row-major payload. dtype codes: 0=f32, 1=f16, 2=i8, 3=i32, 4=i64. Caps incoming payloads at 256 MiB and applies a 60 s read timeout per recv.

Topology graph with per-link latency and bandwidth measurements.

mDNS peer discovery via the mdns-sd crate. Advertises _cascadia._tcp.local. and browses for siblings in the same namespace (a TXT-record field; peers in other namespaces are ignored). Zero-config: workers on the same LAN list each other, but you still connect stages yourself with --next.

Model registry plus on-demand HuggingFace pull. Not wired into the CLI — no crate depends on it; workers never download (only cascadia shard fetches). Registry lives at ~/.cache/cascadia/registry.json; writes are atomic (.tmp + fsync + rename). Symlinks at the registry path are rejected to prevent path-substitution attacks.

cascadia worker --rank N --total M --engine <name> --model <dir> ... is the core serving subcommand; run is its single-machine sugar. Other subcommands: shard (bundled exporter), doctor (environment checks), discover (mDNS browse), engines, completions (shell completions), profile-devices / profile-stages / place / run-placement (placement tooling). The cascadia crate is the binary entry point and depends on cascadia-cli.