① Hardware Baseline and Stack Selection
| Hardware | Specs |
|---|---|
| GPU | RTX 5090 32 GB (Blackwell architecture, whole card 32,607 MB) |
| CUDA | 13.4 (driver 616.92, must be ≥ 12.x to use Blackwell NVFP4 acceleration) |
| CPU / RAM | AMD Ryzen 9 9950X3D 16C · 61.4 GB |
Why choose llama-server over Ollama / LM Studio:
| Option | Pros | Cons |
|---|---|---|
| llama-server | Lightweight, no dependencies · native GGUF · NVFP4 tensor-core acceleration · OpenAI-compatible API | Must compile yourself or download a binary |
| Ollama | Usable after install · easy model management | Fork is lagging · offload not thorough · firmly holds port 11434 and self-recovery is strong and cannot be terminated |
| LM Studio | GUI-friendly · auto download | Closed source · feature-limited · slow to start |
Core reason: it can use the full capacity of Blackwell's tensor cores, with no middle-layer loss. Ollama is only retained for standalone operation on 11434; the main stack switches to 11435 (llama-server) / 11436 (proxy).
② Model and Quantization Selection (two generations)
Generation 1: NVFP4-quantized model → Generation 2: ternary-quantized base; the selection logic is entirely different.
Generation 1: qwen3.8-27b-nvfp4 (deployed 9/13) — decide the parameter count first:
| Candidate | Parameter count | Quantization | Size | Verdict |
|---|---|---|---|---|
| qwen3.8-7b | 7B | NVFP4 | ~5 GB | Too small; not enough quality for multi-tool agent reasoning |
| qwen3.8-14b | 14B | NVFP4 | ~10 GB | Wastes half the VRAM |
| qwen3.8-27b | 27B | NVFP4 | 19.6 GB | Fits a 32 GB card exactly without blowing up |
| qwen3.8-57b | 57B | NVFP4 | ~40 GB | Not enough VRAM |
NVFP4 is a Blackwell-native 4-bit float format: in GGUF, tensors are stored directly with the NVFP4 dtype and after loading run with tcgen05 tensor-core acceleration; switching to integer formats like Q4_K_M can only be emulated in F16, roughly twice as slow. The model was taken as a pre-quantized build from ModelScope, saving the 54 GB raw weights + 3 hours of quantization time.
Generation 2: Ternary-Bonsai-2-27B-PQ2_0 (switched 9/22) — 1.71-bit ternary quantization; double the context in the same VRAM (ctx cap → 262,144). Changing the model is not just a download: the quantization format changed, so the KV quantization tier, the chat-template patch anchors, and the presence/absence of the MTP layer all change along with it (see the performance page and the troubleshooting page).
KV cache quantization three-tier evolution (all measured):
- f16 native: 66.3 KiB/token; 260K ctx does not fit;
- q8_0 / q5_1 (9/21): down to 30.0 KiB/token, minimal quality loss. K must stay q8_0 — K enters the attention scores and quantization error is amplified into every token; V is only a weighted sum and the error averages out;
- q8_0 / q4_0 is a known-problem config tier: the log reports
no FlashAttention vector kernel compiled for K/V types q8_0-q4_0; every attention converts to f16 on the fly — saving VRAM but paying speed, do not configure it this way; - q4_0 / q4_0 (10/4, with -fa on + Bonsai base): quality 7/7 with zero degradation, speed +13.6%, VRAM −4.35 GB — the nvfp4-era premise of "don't use q4_0" has changed.
③ Build and Launch
Built on 2026-09-12 with the VS2026 Developer Command Prompt (vcvarsall x64) so the CUDA SDK is found.
REM 1. 拉源码(带 CUDA 后端) git clone --recursive https://github.com/ggml-org/llama.cpp.git D:\_Qwen3.8-27b\_llama_src REM 2. CMake configure —— 指定 CUDA 后端 cd D:\_Qwen3.8-27b\_llama_src cmake -B build -G Ninja -DGGML_CUDA=ON ^ -DGGML_CUDA_CUDA_PATH="C:/Program Files/NVIDIA GPU Computing Toolkit/CUDA/v13.4" REM 3. 编译 + 产物复制 cmake --build build --config Release --parallel copy build\bin\Release\*.exe ..\llama_cpp\ copy build\bin\Release\*.dll ..\llama_cpp\
Key artifacts: llama-server-impl.dll (8.9 MB core implementation) + ggml-cuda.dll (144.9 MB CUDA acceleration backend). First load took 65 seconds — mmap has to page the 19.6 GB GGUF into VRAM. Later switched to the upstream prebuilt prism-b10754 (includes the MTP Hadamard fix PR #205 and prefill optimization), but the old directory is preserved untouched as a rollback point.
④ Launch Parameter Explanations
Current effective configuration (2026-10-04). Every parameter has a "why"; read this table before changing anything.
| Parameter | Why |
|---|---|
-c 327680 | 320K ctx: 260K is the training cap; extrapolated 1.25× to 320K (36-10). VRAM usage stays manageable as KV quantization drops (see ⑤), while leaving an absolute buffer for ZCode's estimation bias |
-fa on | Flash Attention — a prerequisite for KV quantization, and also faster |
--rope-scaling yarn | yarn extrapolation: scale RoPE frequencies by original length, push the 260K training line out by 1.25×, bypassing server-side active capping (36-10) |
--yarn-orig-ctx 262144 | Tells extrapolation the original training length (GGUF metadata qwen35.context_length); it is the reference for yarn extrapolation |
--override-kv qwen35.context_length=int:327680 | Overrides context_length in the GGUF metadata; without it the server caps at the training line and 300K-token requests are rejected (HTTP 400) |
-ctk q4_0 -ctv q4_0 | KV double q4_0: quality 7/7 with zero degradation + VRAM −4.35 GB. Prerequisite is -fa on; the q8_0/q4_0 mixed tier is a known problem, do not configure it |
-ngl 999 | Full GPU offload |
-np 1 | Agent requests are serial; multiple slots just waste KV |
-b 2048 -ub 512 | batch / ubatch. Values set in the cont-batching era; prefill 70%+ faster and saves 4 GB |
--no-reasoning-preserve | The template by default keeps all thinking from history and continuously consumes context |
--reasoning-budget 1000 | Suppresses overthinking: thinking −76%, body text 0 → 4,820 chars, accuracy unchanged (see the performance page) |
--mmproj ...gguf | Vision projector (0.6 GB): reads real screenshots word by word. The script has an if exist guard, so deleting the file automatically falls back to text only |
--sleep-idle-seconds 180 | Return VRAM to UE after 3 minutes unused; the port keeps listening, wakes in 7.6 seconds (see ⑦) |
--chat-template-file ... | Patched template: fixes the two 500 root causes of reasoning_effort + the system message mid-conversation |
-a "local,qwen3.8:27b,...offload" | Alias list: whatever the client model name is typed, it matches |
--metrics --log-file ... | Metrics + on-disk logs: troubleshooting and speed attribution rely entirely on it |
Tuning knobs (environment variables, temporary override, affect only this launch):
set LLM_CTX=131072 & start_local_llm.bat REM 降 ctx(同步把 contextWindow 改成 ~×0.75) set LLM_SLEEP=60 & start_local_llm.bat REM 1 分钟就还显存;-1 永不归还 set LLM_THINK=0 & start_local_llm.bat REM 关 thinking:约快一倍,推理变浅 set LLM_RBUDGET=500 & start_local_llm.bat REM 收紧思考预算
⑤ VRAM Calibration (measured formula)
VRAM baselines must be measured with "UE open" — baseline usage measured with the engine closed (2,086 MB) cannot be used for budget calculations.
Calibration for the first-generation model (nvfp4, q8_0 KV era):
| ctx | KV type | Actual usage (UE open) | Whole-card remainder | Judgment |
|---|---|---|---|---|
| 65536 | q8_0/q8_0 | ~18,900 MB | ~6,200 MB | Conservative, but ZCode tends to overshoot the cap |
| 81920 | q8_0/q8_0 | ~19,200 MB | ~5,300 MB | Measured overshoot requests did occur |
| 114688 | q8_0/q8_0 | ~20,608 MB | ~4,753 MB | Stable operation achieved |
| 131072 | q8_0/q5_1 | ~20,594 MB | ~4,750 MB | Same VRAM, context +14% |
Calibration for the second-generation model (Bonsai 2): net usage = whole-card wake − sleep = 20,702 − 2,828 ≈ 17.5 GiB; post-KV4 whole-card peak 20,964 MiB. The model budget (user-set ~21 GB) is always calculated with "UE on": whole card 32,607 − UE desktop + engine 7,489 ≈ 25 GB available, minus a margin for UE's later growth. Post ctx extrapolation actual usage ~22.5 GB (+2.6 GB), still ~10 GB left for UE.
⑥ ctx Tuning History (9 phases)
The full trajectory of changing -c again and again — every step was forced by a 400/performance problem.
32768
9/13 first launch. It runs, but is not enough for an agent with tools
8192
Shrunk for faster first token → 400: 46,911 > 8,192
131072
Enlarged to fix the 400 → VRAM filled at 30.7 GB, prefill slow
65536
9/16 halved + cont-batching: prefill 3.8 s → 1.1 s, saves 4 GB VRAM
81920→98304
9/17 ZCode era: raised by pairing on "estimation bias + compression overhead"
114688
9/18 raised again: trigger point 81,920 is safe under the 1.26 coefficient
262144
9/22 Bonsai 2 base swap: 260K fits under ternary quantization
327680
10/4 evening yarn extrapolation: discovered 262144 is the training cap → capping → override-kv bypass → three-needle verification passed
Among these 9 steps, only one rule is truly transferable: server ctx follows the client's estimation bias, not the VRAM cap — every tier increase happened "after overshooting the cap again" (see the performance page · token estimation bias).
⑦ Sleep Mechanism: Why "Clear" Must Be Sleep
The only correct implementation of "use when opened, return VRAM when idle" = process resident + idle sleep.
| Action | Port | VRAM | When the client uses it |
|---|---|---|---|
| Exit the process | Not listening | 0 | Connection refused → "reconnecting N/10" |
| Sleep | Still listening | ~2.4 GB | Request arrives → 7.6 seconds to wake → normal answer |
Measured (--sleep-idle-seconds 180): in use 27,430 MB (incl. UE) → after 180 s idle ~2,400 MB (log handle_sleep: server is entering sleeping state) → next request returns 200 in 7.2 seconds. Early on, proxy-level kill implemented the release, but it had a timing bug and reverted to old parameters on restart — all abandoned. llama.cpp natively supports it; one layer is enough.