Local LLM appliance · AMD Ryzen AI Max

Run frontier models on your own silicon.

Llama Manager turns a Strix Halo (gfx1151) box into a self-contained LLM server — a multi-model llama.cpp router with optimized engines — shipped as an Ubuntu 24.04 appliance ISO and a signed APT repository.

GPG-signed ISO & APTReproducible buildsFully offlineUbuntu 24.04.4 Desktop
Target hardware
AMD Ryzen AI Max / Strix Halo
GPU architecture
gfx1151 iGPU
Unified memory
Up to 128 GB shared GTT
Base OS
Ubuntu 24.04.4 Desktop
What's inside

One appliance, everything local.

Built specifically for the AMD Ryzen AI Max platform and its large unified-memory iGPU.

Multi-model llama.cpp router

Serve many GGUF models from one endpoint. The router loads and switches models on demand and queues incoming requests during a swap instead of failing them.

Optimized dedicated engines

Beyond stock llama.cpp, purpose-built inference engines are tuned for the gfx1151 iGPU and its large unified memory, selected automatically per model.

Appliance & kiosk

Boots straight into a full-screen local dashboard. Manage models, presets, and routing from the browser, with a local System Login escape hatch for shell access.

Fully offline

Everything runs on the box. No cloud dependency and no telemetry required — models and inference stay entirely local to your hardware.

Signed APT + reproducible ISO

Ships with a signed APT repository preconfigured, so updates install in place with standard apt tooling and GPG verification against the archive key.

Unified memory scale

Strix Halo's shared 128 GB GTT pool lets a single box hold large models that would not fit in a typical discrete-GPU VRAM budget.

Get started

Download the appliance image and flash it to a USB drive, or add the signed APT repository to an existing Ubuntu 24.04 install.