Skip to content
CascadiaCascadiaDocsOpen source
Operations

Run as a background service

Keep workers running with systemd, Windows services, or launchd.

Cascadia runs as a long-lived CLI process and doesn’t daemonize itself. Run it under your platform’s usual supervisor: systemd on Linux, NSSM or Task Scheduler on Windows, or launchd on macOS.

Use a release bundle where you can. Its OpenVINO libraries sit next to the binary, so the service finds them without extra setup.

A binary built from source links against the OpenVINO SDK, and a service doesn’t inherit your shell’s PATH. Point the service at the runtime libraries and at TBB, which is next to the runtime folder rather than inside it. Otherwise the worker exits immediately with a missing libtbb or tbb12 error.

  • Windows: nssm set <svc> AppEnvironmentExtra PATH=<sdk>\runtime\bin\intel64\Release;<sdk>\runtime\3rdparty\tbb\bin;…
  • Linux: Environment=LD_LIBRARY_PATH=<sdk>/runtime/lib/intel64:<sdk>/runtime/3rdparty/tbb/lib

Start from the template unit, cascadia-worker.service. It expects a cascadia user, the cascadia binary on PATH (or edit ExecStart), and model shards in /opt/cascadia/shards/.

Set CASCADIA_API for rank 0. A worker started without --api reads from stdin, gets end-of-file under systemd, and exits with status 0. Restart=on-failure doesn’t restart a clean exit, so the worker stays down. Other ranks ignore --api and keep running in the relay loop, so one template works for every rank.

Terminal window
# Edit the Environment= lines in the unit file, then:
sudo cp docs/deploy/cascadia-worker.service /etc/systemd/system/cascadia-worker@.service
sudo systemctl daemon-reload
# Start stage 0 and stage 1, one per host. Two instances on the same host
# would both try to bind CASCADIA_LISTEN.
sudo systemctl enable --now cascadia-worker@0.service
sudo systemctl enable --now cascadia-worker@1.service
# Check status and follow logs:
sudo systemctl status cascadia-worker@0.service
sudo journalctl -u cascadia-worker@0.service -f

The unit uses Type=simple, so systemd tracks the worker process directly and no PID file is needed. It restarts the worker on failure, up to 3 times a minute.

On SIGTERM, rank 0 closes its sockets and releases GPU contexts, then exits with status 0. Other ranks don’t handle the signal, so they stop without a clean shutdown. They don’t store anything, so nothing is lost, but they won’t log a shutdown message.

Terminal window
# Install NSSM (https://nssm.cc), then:
nssm install cascadia-worker-0 "C:\cascadia\cascadia.exe" `
"worker --rank 0 --total 2 --engine ov-runtime --device GPU " `
"--model C:\cascadia\shards --next 10.0.0.2:9100 --listen :9100 " `
"--api :8000 --log-level info"
nssm set cascadia-worker-0 AppStdout C:\ProgramData\cascadia\worker-0.log
nssm set cascadia-worker-0 AppStderr C:\ProgramData\cascadia\worker-0.log
nssm set cascadia-worker-0 AppExit Default Restart
nssm start cascadia-worker-0

Rank 0 needs --api here too. Without it, the worker reads from stdin, gets end-of-file, and exits with status 0, and AppExit Default Restart turns that into a restart loop. Other ranks don’t need it.

When stopping a service, NSSM sends a console event, then WM_CLOSE, then TerminateProcess, so the worker gets a chance to shut down cleanly.

Don’t start production workers over SSH with start /B. Windows OpenSSH runs them in the services session, and they stop when the SSH session closes. Use NSSM or Task Scheduler with /RU SYSTEM instead.

macOS has no Intel GPU runtime, so use launchd for development only, with the mock engine or a stub build. A minimal ~/Library/LaunchAgents/com.cascadia.worker.plist:

<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
"http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key><string>com.cascadia.worker</string>
<key>ProgramArguments</key>
<array>
<string>/usr/local/bin/cascadia</string>
<string>worker</string>
<string>--rank</string><string>0</string>
<string>--total</string><string>1</string>
<string>--engine</string><string>mock</string>
<string>--model</string><string>mock-model</string>
<string>--api</string><string>:8000</string>
</array>
<key>KeepAlive</key><true/>
<key>StandardOutPath</key><string>/tmp/cascadia-worker.out</string>
<key>StandardErrorPath</key><string>/tmp/cascadia-worker.err</string>
</dict>
</plist>

Load it with launchctl load ~/Library/LaunchAgents/com.cascadia.worker.plist.

The HTTP API on rank 0 has two endpoints for health checks:

  • GET /health returns {"status": "ok"}.
  • GET /v1/models lists the served model IDs.

Point TCP or HTTP probes at the API port. Other ranks don’t serve an API, so supervise them by process state and exit code, which systemd does automatically with Type=simple. A TCP probe of a worker’s --listen port also works as a liveness check.

Cascadia writes plain-text logs to stdout and stderr. Set the level with --log-level (default info):

2026-07-02T17:41:23.189Z INFO cascadia_runner: runner ready

Cascadia doesn’t include a JSON log formatter. To ship structured logs to a service like Loki or CloudWatch, wrap cascadia worker in whatever format your log shipper expects.