Multi-model llama.cpp router
Serve many GGUF models from one endpoint. The router loads and switches models on demand and queues incoming requests during a swap instead of failing them.
Llama Manager turns a Strix Halo (gfx1151) box into a self-contained LLM server — a multi-model llama.cpp router with optimized engines — shipped as an Ubuntu 24.04 appliance ISO and a signed APT repository.
Built specifically for the AMD Ryzen AI Max platform and its large unified-memory iGPU.
Serve many GGUF models from one endpoint. The router loads and switches models on demand and queues incoming requests during a swap instead of failing them.
Beyond stock llama.cpp, purpose-built inference engines are tuned for the gfx1151 iGPU and its large unified memory, selected automatically per model.
Boots straight into a full-screen local dashboard. Manage models, presets, and routing from the browser, with a local System Login escape hatch for shell access.
Everything runs on the box. No cloud dependency and no telemetry required — models and inference stay entirely local to your hardware.
Ships with a signed APT repository preconfigured, so updates install in place with standard apt tooling and GPG verification against the archive key.
Strix Halo's shared 128 GB GTT pool lets a single box hold large models that would not fit in a typical discrete-GPU VRAM budget.