deploy qwen 3.8
This commit is contained in:
@@ -2,151 +2,58 @@
|
||||
|
||||
In-cluster LLM inference via llama.cpp's `llama-server`, serving a local model
|
||||
on the NUCBox APU (AMD Ryzen AI Max 395 / Strix Halo, Radeon 8060S, ~120 GiB
|
||||
unified memory: ~90 GiB VRAM / 30 GiB CPU RAM).
|
||||
unified memory).
|
||||
|
||||
LiteLLM (`litellm/`) points at these in-cluster Services instead of the
|
||||
external `10.88.20.12:11434` Ollama endpoint.
|
||||
LiteLLM (`litellm/`) points at the in-cluster Service instead of the external
|
||||
Ollama endpoint.
|
||||
|
||||
## Layout
|
||||
## Active Model
|
||||
|
||||
One Deployment + Service **per model**, all in namespace `llamacpp`, all pinned
|
||||
to the NUCBox (`nodeSelector: {kubernetes.io/arch: amd64, hardware: high-memory}`):
|
||||
| Alias | Model | Configuration | Service |
|
||||
|---|---|---|---|
|
||||
| `qwen3.8-27b` | Qwen3.8-27B with MTP | `Q4_K_M` primary, `Q4_0` draft, 196k context, q8_0 K/V cache | `llamacpp-qwen38-27b.llamacpp:80` |
|
||||
|
||||
| Alias | Model | GGUF | Service | Args ref |
|
||||
|--------------------------|--------------------------------|-------------------------------------------------------|------------------------------------------------------|----------|
|
||||
| `deepseek-v4-flash-0731` | DeepSeek-V4-Flash-0731 (MoE) | unsloth/DeepSeek-V4-Flash-0731-GGUF (UD-IQ1_M, ~87 GiB) | `llamacpp-deepseek-v4-flash-0731.llamacpp:80` | [args-deepseek-v4-flash-0731.md](args-deepseek-v4-flash-0731.md) |
|
||||
The active Deployment uses llama.cpp's Hugging Face downloader for both model
|
||||
repositories:
|
||||
|
||||
DeepSeek-V4-Flash-0731 is a Mixture-of-Experts model (256 experts, 6 active per
|
||||
token) with MLA attention, so only a small fraction of the weights is computed
|
||||
per token. The full ~87 GiB of IQ1_M weights is loaded into the unified memory
|
||||
pool.
|
||||
- Primary: `ggml-org/Qwen3.8-27B-GGUF:Q4_K_M`
|
||||
- Draft: `ggml-org/Qwen3.8-27B-GGUF:Q4_0`
|
||||
|
||||
> **Previously** the NUCBox ran two co-resident Qwen3.6 models (a 27B dense and
|
||||
> a 35B-A3B MoE "flash"). DeepSeek-V4-Flash-0731 (~87 GiB) nearly fills the
|
||||
> whole 90 GiB VRAM pool on its own, so both Qwen models were removed to make
|
||||
> room. Their GGUF files are deleted from the shared PVC by the new pod's
|
||||
> `fetch-model` initContainer on first boot. Their deployment arguments are
|
||||
> still documented for redeployment:
|
||||
> - [args-qwen36-27b.md](args-qwen36-27b.md) — dense 27B (deeper reasoning)
|
||||
> - [args-qwen36-35b-a3b.md](args-qwen36-35b-a3b.md) — MoE 35B-A3B "flash" (fast)
|
||||
The model cache is stored on the shared hostPath PVC at
|
||||
`/data/llamacpp/models` on the NUCBox. The server is configured with
|
||||
`--spec-default --spec-type draft-mtp`, `--reasoning-preserve`, `--fit off`,
|
||||
and `--agent`.
|
||||
|
||||
Model files are downloaded idempotently by an initContainer into a shared
|
||||
hostPath PVC (`/data/llamacpp/models` on the NUCBox), so pods survive reboots
|
||||
without re-downloading. The UD-IQ1_M GGUF is split across 3 shards
|
||||
(`-00001-of-00003` … `-00003-of-00003`); llama.cpp auto-loads all shards when
|
||||
pointed at the first one.
|
||||
## Retired Models
|
||||
|
||||
The following deployments are no longer active, but their argument references
|
||||
are retained for future redeployment:
|
||||
|
||||
- [DeepSeek-V4-Flash-0731](args-deepseek-v4-flash-0731.md)
|
||||
- [Qwen3.6-27B](args-qwen36-27b.md)
|
||||
- [Qwen3.6-35B-A3B](args-qwen36-35b-a3b.md)
|
||||
|
||||
## GPU / Vulkan
|
||||
|
||||
The `server-vulkan` image (`ghcr.io/ggml-org/llama.cpp:server-vulkan`) bundles
|
||||
the Mesa/RADV Vulkan driver, which supports the Radeon 8060S (RDNA 3.5). The
|
||||
Vulkan backend supports the `IQ1_M` matmul (including the MoE `matmul_id`
|
||||
variant), so the whole model runs on the GPU. `deepseek4` is a brand-new arch
|
||||
(2026-07), so the floating `server-vulkan` tag is used to pull a recent enough
|
||||
build — pin to a specific `server-vulkan-bXXXX` tag once a known-good one is
|
||||
verified.
|
||||
the Mesa/RADV Vulkan driver for the Radeon 8060S. The container mounts
|
||||
`/dev/dri` and runs privileged, which is the current way to provide Vulkan
|
||||
access on k3s without a device plugin.
|
||||
|
||||
The container mounts `/dev/dri` and runs `privileged: true` — the simplest
|
||||
reliable way to give Vulkan access to the DRM render node on k3s without a
|
||||
device plugin. Tighten later with `supplementalGroups` (the host's `render`
|
||||
group GID) if desired.
|
||||
|
||||
### Verify the GPU is actually used
|
||||
Verify GPU use with:
|
||||
|
||||
```bash
|
||||
kubectl logs -n llamacpp deploy/llamacpp-deepseek-v4-flash-0731 | grep -iE 'vulkan|gpu|offload|device'
|
||||
kubectl logs -n llamacpp deploy/llamacpp-qwen38-27b | grep -iE 'vulkan|gpu|offload|device'
|
||||
```
|
||||
|
||||
If only a CPU device shows up, the container can't see the GPU — check that
|
||||
`/dev/dri/renderD128` exists on the NUCBox and that the `amdgpu` module is loaded.
|
||||
If only a CPU device appears, check that `/dev/dri/renderD128` exists on the
|
||||
NUCBox and that the `amdgpu` module is loaded.
|
||||
|
||||
## VRAM budget (single model)
|
||||
## Adding Or Replacing A Model
|
||||
|
||||
The model (~87 GiB IQ1_M) is almost the size of the entire 90 GiB VRAM pool, so
|
||||
it **cannot be fully offloaded**: `-ngl 999` would try to put all 43 layers into
|
||||
VRAM and overflow once the KV cache + Vulkan compute buffers are added. Instead
|
||||
`-ngl 38` offloads 38 of 43 layers to the GPU and keeps 5 layers (~10 GiB) on
|
||||
CPU RAM, leaving ~5 GiB of VRAM headroom for the KV cache, compute buffers, and
|
||||
fragmentation.
|
||||
|
||||
**KV cache is f16, not q8_0** — the Vulkan backend has no Flash Attention for
|
||||
the `deepseek4` arch, and quantized V cache requires Flash Attention (llama.cpp
|
||||
hard-errors: *"quantized V cache was requested, but this requires Flash
|
||||
Attention"*). `deepseek4`/MLA models also require K and V cache types to be
|
||||
*identical*, so K cannot be quantized either. f16 MLA KV at 64k is ~5.7 GiB
|
||||
(576 K + 512 V elements/token/layer × 43 layers × 65536 tokens × 2 bytes) —
|
||||
larger than q8_0 would be, which is why `-ngl` is 38 rather than 40.
|
||||
|
||||
Approximate VRAM usage:
|
||||
|
||||
| Component | VRAM |
|
||||
|---------------------------------|-------------|
|
||||
| Weights (38 GPU layers) | ~77 GiB |
|
||||
| KV cache (f16, 64k, 1 slot) | ~5.7 GiB |
|
||||
| Vulkan compute buffers | ~2 GiB |
|
||||
| **Total in VRAM** | **~85 GiB** |
|
||||
| **Headroom (of 90 GiB)** | **~5 GiB** |
|
||||
|
||||
5 layers (~10 GiB) live in CPU RAM (counted against the pod's cgroup memory
|
||||
limit, not VRAM). VRAM is exclusive to this model; the other NUCBox pods only
|
||||
compete for the 30 GiB CPU RAM.
|
||||
|
||||
Several `deepseek4`-specific fused ops (Lightning Indexer, HC pre/comb/post)
|
||||
are not yet implemented in the Vulkan backend and fall back to CPU (logged as
|
||||
warnings, not fatal). Inference still works; it will speed up once those ops
|
||||
land in a future `server-vulkan` build.
|
||||
|
||||
## Tuning
|
||||
|
||||
The key knobs (in `deployment-deepseek-v4-flash-0731.yaml`):
|
||||
|
||||
- `-ngl 38` — offload 38 of 43 layers to GPU. The model (~87 GiB) is nearly the
|
||||
whole 90 GiB VRAM pool, so full offload would overflow once the f16 KV cache +
|
||||
compute buffers are added. 5 layers (~10 GiB) on CPU leaves ~5 GiB VRAM
|
||||
headroom. Raise toward 43 if VRAM allows; lower (e.g. 36) if the pod OOMs /
|
||||
Vulkan runs out of device memory.
|
||||
- `-c 65536` — total KV-cache context (64k, the required minimum). 1 slot gets
|
||||
the full 64k. f16 MLA KV at 64k is ~5.7 GiB; capped at the minimum to maximise
|
||||
VRAM headroom. Raise if headroom allows.
|
||||
- `-np 1` — 1 parallel slot (the full 64k goes to a single concurrent request).
|
||||
Extra slots multiply the f16 KV cost (~5.7 GiB/slot); 1 slot keeps headroom
|
||||
maximal.
|
||||
- `--cache-type-k f16 --cache-type-v f16` — **f16 KV cache (NOT quantized).**
|
||||
The Vulkan backend has no Flash Attention for `deepseek4`, and quantized V
|
||||
cache requires Flash Attention. `deepseek4`/MLA models also require K and V
|
||||
cache types to be identical, so K cannot be quantized either. This is the
|
||||
reason `-ngl` is 38 rather than 40.
|
||||
- `--temp 1.0 --top-p 0.95` — default sampling parameters (DeepSeek-V4
|
||||
recommendation). These are server defaults; clients can override per request
|
||||
via the OpenAI-compatible API.
|
||||
- `--threads 8` — CPU threads for sampling + the 5 CPU-resident layers.
|
||||
|
||||
## Memory accounting
|
||||
|
||||
k8s sees only the ~30 GiB system RAM as allocatable (the ~90 GiB VRAM is
|
||||
reserved by firmware and managed by `amdgpu`). The GPU-resident model weights
|
||||
and KV cache live in VRAM and are **not** counted against the container's cgroup
|
||||
memory limit — that limit only covers CPU-side overhead, the mmap'd GGUF pages
|
||||
for the 5 CPU-resident layers (~10 GiB, resident during inference), and
|
||||
reclaimable page cache during load. If the pod is OOM-killed during model load
|
||||
or inference, raise the memory limit (and/or lower `-ngl` to push more layers
|
||||
to VRAM).
|
||||
|
||||
## Adding / replacing a model
|
||||
|
||||
1. Copy `deployment-deepseek-v4-flash-0731.yaml` → `deployment-<new>.yaml`;
|
||||
change the `model:` label, GGUF URL/file(s), `--alias`, and Service name.
|
||||
For split GGUFs, point `-m` at the first shard and download all shards in
|
||||
the `fetch-model` initContainer.
|
||||
2. Point LiteLLM at it in `litellm/litellm.yaml`:
|
||||
```yaml
|
||||
- model_name: <alias>
|
||||
litellm_params:
|
||||
model: openai/<alias>
|
||||
api_base: http://<service>.llamacpp/v1
|
||||
api_key: "sk-no-auth"
|
||||
```
|
||||
3. (No gen-apps.sh change needed — the `llamacpp` app already syncs the whole
|
||||
directory recursively.)
|
||||
4. Check the VRAM budget table above — at ~87 GiB this model nearly fills the
|
||||
90 GiB pool on its own, so co-locating another large model is not possible.
|
||||
1. Copy `deployment-qwen38-27b.yaml` to `deployment-<new>.yaml` and change the
|
||||
model repositories, alias, labels, and Service name.
|
||||
2. Add the matching alias and Service URL to `litellm/litellm.yaml`.
|
||||
3. No `gen-apps.sh` change is needed because the `llamacpp` ArgoCD Application
|
||||
syncs the directory recursively.
|
||||
4. Check the model and KV-cache size against the NUCBox's available VRAM.
|
||||
|
||||
Reference in New Issue
Block a user