# llama.cpp (llama-server) In-cluster LLM inference via llama.cpp's `llama-server`, serving local models on the NUCBox APU (AMD Ryzen AI Max 395 / Strix Halo, Radeon 8060S, 128 GiB unified memory: 32 GiB RAM / 96 GiB VRAM). This replaces the bare-metal Ollama setup for models that benefit from always-loaded weights + tuned batching. LiteLLM (`litellm/`) points at these in-cluster Services instead of the external `10.88.20.12:11434` Ollama endpoint. ## Layout One Deployment + Service **per model**, all in namespace `llamacpp`, all pinned to the NUCBox (`nodeSelector: {kubernetes.io/arch: amd64, hardware: high-memory}`): | Alias | Model | GGUF | Service | |-------------------|----------------------------|-------------------------------------------------------|-------------------------------------------| | `qwen3.6-27b` | Qwen3.6-27B (dense) | unsloth/Qwen3.6-27B-MTP-GGUF (Q4_K_XL, ~16 GiB) | `llamacpp-qwen36-27b.llamacpp:80` | | `qwen3.6-35b-a3b` | Qwen3.6-35B-A3B (MoE/flash)| unsloth/Qwen3.6-35B-A3B-MTP-GGUF (Q4_K_XL, ~20 GiB) | `llamacpp-qwen36-35b-a3b.llamacpp:80` | The 35B-A3B is a Mixture-of-Experts model (3B active params per token), so inference is significantly faster than the dense 27B despite more total weights — hence the "flash" label. Use it for latency-sensitive workloads; use the 27B for deeper reasoning. Model files are downloaded idempotently by an initContainer into a shared hostPath PVC (`/data/llamacpp/models` on the NUCBox), so pods survive reboots without re-downloading. ## GPU / Vulkan The `server-vulkan` image (`ghcr.io/ggml-org/llama.cpp:server-vulkan`) bundles the Mesa/RADV Vulkan driver, which supports the Radeon 8060S (RDNA 3.5). Full layer offload (`-ngl 999`) puts model weights entirely in the 96 GiB VRAM pool. The container mounts `/dev/dri` and runs `privileged: true` — the simplest reliable way to give Vulkan access to the DRM render node on k3s without a device plugin. Tighten later with `supplementalGroups` (the host's `render` group GID) if desired. ### Verify the GPU is actually used ```bash kubectl logs -n llamacpp deploy/llamacpp-qwen36-27b | grep -iE 'vulkan|gpu|offload|device' kubectl logs -n llamacpp deploy/llamacpp-qwen36-35b-a3b | grep -iE 'vulkan|gpu|offload|device' ``` If only a CPU device shows up, the container can't see the GPU — check that `/dev/dri/renderD128` exists on the NUCBox and that the `amdgpu` module is loaded. ## VRAM budget (both models co-resident) Both models run simultaneously on the same 96 GiB VRAM pool. Approximate usage: | Model | Weights | KV cache (131k×4 for a3b, 32k×4 for 27b) | Subtotal | |-------------------|----------|-------------------------------------------|----------| | qwen3.6-27b | ~16 GiB | ~8 GiB (32k total, 8k/slot) | ~24 GiB | | qwen3.6-35b-a3b | ~20 GiB | ~9 GiB (131k total, 33k/slot) | ~29 GiB | | **Total** | | | **~53 GiB** | ~43 GiB headroom — very comfortable. The MoE's KV cache is tiny (~72 KiB/token vs the dense 27B's ~256 KiB/token), so large context is nearly free. ## Tuning The key knobs (in each `deployment-*.yaml`): - `-ngl 999` — offload all layers to GPU. Reduce only if VRAM is tight. - `-c 131072` (35b-a3b) / `32768` (27b) — total KV-cache context. With `-np 4` the 35b-a3b gets 32768 tokens per slot (enough for the full SOUL.md + prompt); the 27b gets 8192 per slot. The MoE's KV cache is ~72 KiB/token so large context is cheap; the dense 27B's is ~256 KiB/token. Raise the 27B's `-c` too if interactive sessions hit the context limit. - `-np 4` — parallel slots (concurrent requests). Each extra slot multiplies KV-cache VRAM usage. Bump higher on the flash model if you need more throughput (it has VRAM headroom). - `--threads 8` — CPU threads for sampling/overhead. Mostly irrelevant under full GPU offload; tune if CPU-bound. ## Memory accounting k8s sees only the ~32 GiB system RAM as allocatable (the 96 GiB VRAM is reserved by firmware and managed by `amdgpu`). The model weights and KV cache live in VRAM and are **not** counted against the container's cgroup memory limit — that limit only covers CPU-side overhead and the mmap'd GGUF pages during load. If the pod is OOM-killed during model load, raise the memory limit. ## Adding a model 1. Copy `deployment-qwen36-27b.yaml` → `deployment-.yaml`; change the `model:` label, GGUF URL/file, `--alias`, and Service name. 2. Point LiteLLM at it in `litellm/litellm.yaml`: ```yaml - model_name: litellm_params: model: openai/ api_base: http://.llamacpp/v1 api_key: "sk-no-auth" ``` 3. (No gen-apps.sh change needed — the `llamacpp` app already syncs the whole directory recursively.) 4. Check the VRAM budget table above — two large models may not coexist.