AMD Ryzen Llama Manager · Ryzen AI Max

Run open-source models on your own silicon.

Llama Manager turns a Strix Halo (gfx1151) box into a self-contained LLM server — a multi-model llama.cpp router with optimized engines — shipped as an Ubuntu 24.04 appliance ISO and a signed APT repository.

GPG-signed ISO & APTReproducible buildsFully offlineUbuntu 24.04.4 Desktop
Target hardware
AMD Ryzen AI Max / Strix Halo
GPU architecture
gfx1151 iGPU
Unified memory
Up to 128 GB shared GTT
Base OS
Ubuntu 24.04.4 Desktop
Recommended hardware

The boxes we build and test on.

Llama Manager runs on any Ryzen AI Max+ 395 (Strix Halo) system. These are the two machines the appliance image is developed and smoke-tested against — pick the 128 GB configuration so the whole unified-memory pool is available to models.

One BIOS setting, then you're done.

Flash the ISO, install it, and change exactly one firmware value: set the UMA Frame Buffer Size (dedicated VRAM) from the stock 64 MB to 1 GB. That is the whole manual setup. The iGPU keeps its small dedicated partition and reaches the rest of the 128 GB pool through GTT, which is where large models actually live — the appliance handles everything else on first boot.

These are affiliate links — buying through them supports the project at no extra cost to you. We are not affiliated with GMKtec and neither machine ships with Llama Manager preinstalled.

What's inside

One appliance, everything local.

Built specifically for the AMD Ryzen AI Max platform and its large unified-memory iGPU.

Multi-model llama.cpp router

Serve many GGUF models from one endpoint. The router loads and switches models on demand and queues incoming requests during a swap instead of failing them.

Optimized dedicated engines

Beyond stock llama.cpp, purpose-built inference engines are tuned for the gfx1151 iGPU and its large unified memory, selected automatically per model.

Appliance & kiosk

Boots straight into a full-screen local dashboard. Manage models, presets, and routing from the browser, with a local System Login escape hatch for shell access.

Fully offline

Everything runs on the box. No cloud dependency and no telemetry required — models and inference stay entirely local to your hardware.

Signed APT + reproducible ISO

Ships with a signed APT repository preconfigured, so updates install in place with standard apt tooling and GPG verification against the archive key.

Unified memory scale

Strix Halo's shared 128 GB GTT pool lets a single box hold large models that would not fit in a typical discrete-GPU VRAM budget.

Llama Manager Flasher

Flash it in one step.

Download the Flasher, plug in a USB drive or microSD, and it fetches the latest AMD Ryzen image — checksum-verified — and writes it for you. No installation: one portable file.

These builds are not yet code-signed. On macOS, right-click the app and choose Open on first launch (or run xattr -cr on the downloaded .dmg). On Windows, choose More info → Run anyway at the SmartScreen prompt.

Get started

Download the appliance image and flash it to a USB drive, or add the signed APT repository to an existing Ubuntu 24.04 install.