① Hardware and Selection ② Model and Quantization ③ Build and Launch ④ Parameter Explanations ⑤ VRAM Calibration ⑥ ctx Tuning History ⑦ Sleep Mechanism

① Hardware Baseline and Stack Selection

HardwareSpecs
GPURTX 5090 32 GB (Blackwell architecture, whole card 32,607 MB)
CUDA13.4 (driver 616.92, must be ≥ 12.x to use Blackwell NVFP4 acceleration)
CPU / RAMAMD Ryzen 9 9950X3D 16C · 61.4 GB

Why choose llama-server over Ollama / LM Studio:

OptionProsCons
llama-serverLightweight, no dependencies · native GGUF · NVFP4 tensor-core acceleration · OpenAI-compatible APIMust compile yourself or download a binary
OllamaUsable after install · easy model managementFork is lagging · offload not thorough · firmly holds port 11434 and self-recovery is strong and cannot be terminated
LM StudioGUI-friendly · auto downloadClosed source · feature-limited · slow to start

Core reason: it can use the full capacity of Blackwell's tensor cores, with no middle-layer loss. Ollama is only retained for standalone operation on 11434; the main stack switches to 11435 (llama-server) / 11436 (proxy).

② Model and Quantization Selection (two generations)

Generation 1: NVFP4-quantized model → Generation 2: ternary-quantized base; the selection logic is entirely different.

Generation 1: qwen3.8-27b-nvfp4 (deployed 9/13) — decide the parameter count first:

CandidateParameter countQuantizationSizeVerdict
qwen3.8-7b7BNVFP4~5 GBToo small; not enough quality for multi-tool agent reasoning
qwen3.8-14b14BNVFP4~10 GBWastes half the VRAM
qwen3.8-27b27BNVFP419.6 GBFits a 32 GB card exactly without blowing up
qwen3.8-57b57BNVFP4~40 GBNot enough VRAM

NVFP4 is a Blackwell-native 4-bit float format: in GGUF, tensors are stored directly with the NVFP4 dtype and after loading run with tcgen05 tensor-core acceleration; switching to integer formats like Q4_K_M can only be emulated in F16, roughly twice as slow. The model was taken as a pre-quantized build from ModelScope, saving the 54 GB raw weights + 3 hours of quantization time.

Generation 2: Ternary-Bonsai-2-27B-PQ2_0 (switched 9/22) — 1.71-bit ternary quantization; double the context in the same VRAM (ctx cap → 262,144). Changing the model is not just a download: the quantization format changed, so the KV quantization tier, the chat-template patch anchors, and the presence/absence of the MTP layer all change along with it (see the performance page and the troubleshooting page).

KV cache quantization three-tier evolution (all measured):

  1. f16 native: 66.3 KiB/token; 260K ctx does not fit;
  2. q8_0 / q5_1 (9/21): down to 30.0 KiB/token, minimal quality loss. K must stay q8_0 — K enters the attention scores and quantization error is amplified into every token; V is only a weighted sum and the error averages out;
  3. q8_0 / q4_0 is a known-problem config tier: the log reports no FlashAttention vector kernel compiled for K/V types q8_0-q4_0; every attention converts to f16 on the fly — saving VRAM but paying speed, do not configure it this way;
  4. q4_0 / q4_0 (10/4, with -fa on + Bonsai base): quality 7/7 with zero degradation, speed +13.6%, VRAM −4.35 GB — the nvfp4-era premise of "don't use q4_0" has changed.

③ Build and Launch

Built on 2026-09-12 with the VS2026 Developer Command Prompt (vcvarsall x64) so the CUDA SDK is found.

REM 1. 拉源码(带 CUDA 后端)
git clone --recursive https://github.com/ggml-org/llama.cpp.git D:\_Qwen3.8-27b\_llama_src

REM 2. CMake configure —— 指定 CUDA 后端
cd D:\_Qwen3.8-27b\_llama_src
cmake -B build -G Ninja -DGGML_CUDA=ON ^
  -DGGML_CUDA_CUDA_PATH="C:/Program Files/NVIDIA GPU Computing Toolkit/CUDA/v13.4"

REM 3. 编译 + 产物复制
cmake --build build --config Release --parallel
copy build\bin\Release\*.exe ..\llama_cpp\
copy build\bin\Release\*.dll ..\llama_cpp\

Key artifacts: llama-server-impl.dll (8.9 MB core implementation) + ggml-cuda.dll (144.9 MB CUDA acceleration backend). First load took 65 seconds — mmap has to page the 19.6 GB GGUF into VRAM. Later switched to the upstream prebuilt prism-b10754 (includes the MTP Hadamard fix PR #205 and prefill optimization), but the old directory is preserved untouched as a rollback point.

④ Launch Parameter Explanations

Current effective configuration (2026-10-04). Every parameter has a "why"; read this table before changing anything.

ParameterWhy
-c 327680320K ctx: 260K is the training cap; extrapolated 1.25× to 320K (36-10). VRAM usage stays manageable as KV quantization drops (see ⑤), while leaving an absolute buffer for ZCode's estimation bias
-fa onFlash Attention — a prerequisite for KV quantization, and also faster
--rope-scaling yarnyarn extrapolation: scale RoPE frequencies by original length, push the 260K training line out by 1.25×, bypassing server-side active capping (36-10)
--yarn-orig-ctx 262144Tells extrapolation the original training length (GGUF metadata qwen35.context_length); it is the reference for yarn extrapolation
--override-kv qwen35.context_length=int:327680Overrides context_length in the GGUF metadata; without it the server caps at the training line and 300K-token requests are rejected (HTTP 400)
-ctk q4_0 -ctv q4_0KV double q4_0: quality 7/7 with zero degradation + VRAM −4.35 GB. Prerequisite is -fa on; the q8_0/q4_0 mixed tier is a known problem, do not configure it
-ngl 999Full GPU offload
-np 1Agent requests are serial; multiple slots just waste KV
-b 2048 -ub 512batch / ubatch. Values set in the cont-batching era; prefill 70%+ faster and saves 4 GB
--no-reasoning-preserveThe template by default keeps all thinking from history and continuously consumes context
--reasoning-budget 1000Suppresses overthinking: thinking −76%, body text 0 → 4,820 chars, accuracy unchanged (see the performance page)
--mmproj ...ggufVision projector (0.6 GB): reads real screenshots word by word. The script has an if exist guard, so deleting the file automatically falls back to text only
--sleep-idle-seconds 180Return VRAM to UE after 3 minutes unused; the port keeps listening, wakes in 7.6 seconds (see ⑦)
--chat-template-file ...Patched template: fixes the two 500 root causes of reasoning_effort + the system message mid-conversation
-a "local,qwen3.8:27b,...offload"Alias list: whatever the client model name is typed, it matches
--metrics --log-file ...Metrics + on-disk logs: troubleshooting and speed attribution rely entirely on it

Tuning knobs (environment variables, temporary override, affect only this launch):

set LLM_CTX=131072      & start_local_llm.bat   REM 降 ctx(同步把 contextWindow 改成 ~×0.75)
set LLM_SLEEP=60        & start_local_llm.bat   REM 1 分钟就还显存;-1 永不归还
set LLM_THINK=0         & start_local_llm.bat   REM 关 thinking:约快一倍,推理变浅
set LLM_RBUDGET=500     & start_local_llm.bat   REM 收紧思考预算

⑤ VRAM Calibration (measured formula)

VRAM baselines must be measured with "UE open" — baseline usage measured with the engine closed (2,086 MB) cannot be used for budget calculations.

Calibration for the first-generation model (nvfp4, q8_0 KV era):

KV per token = 16 layers × 4 heads × (256+256) dims × 1.0625 B(q8_0) ≈ 35.2 KiB / token Total usage ≈ 16.5 GiB (weights + compute buffer + CUDA context) + ctx × 35.2 KiB
ctxKV typeActual usage (UE open)Whole-card remainderJudgment
65536q8_0/q8_0~18,900 MB~6,200 MBConservative, but ZCode tends to overshoot the cap
81920q8_0/q8_0~19,200 MB~5,300 MBMeasured overshoot requests did occur
114688q8_0/q8_0~20,608 MB~4,753 MBStable operation achieved
131072q8_0/q5_1~20,594 MB~4,750 MBSame VRAM, context +14%

Calibration for the second-generation model (Bonsai 2): net usage = whole-card wake − sleep = 20,702 − 2,828 ≈ 17.5 GiB; post-KV4 whole-card peak 20,964 MiB. The model budget (user-set ~21 GB) is always calculated with "UE on": whole card 32,607 − UE desktop + engine 7,489 ≈ 25 GB available, minus a margin for UE's later growth. Post ctx extrapolation actual usage ~22.5 GB (+2.6 GB), still ~10 GB left for UE.

⑥ ctx Tuning History (9 phases)

The full trajectory of changing -c again and again — every step was forced by a 400/performance problem.

32768

9/13 first launch. It runs, but is not enough for an agent with tools

8192

Shrunk for faster first token → 400: 46,911 > 8,192

131072

Enlarged to fix the 400 → VRAM filled at 30.7 GB, prefill slow

65536

9/16 halved + cont-batching: prefill 3.8 s → 1.1 s, saves 4 GB VRAM

81920→98304

9/17 ZCode era: raised by pairing on "estimation bias + compression overhead"

114688

9/18 raised again: trigger point 81,920 is safe under the 1.26 coefficient

262144

9/22 Bonsai 2 base swap: 260K fits under ternary quantization

327680

10/4 evening yarn extrapolation: discovered 262144 is the training cap → capping → override-kv bypass → three-needle verification passed

Among these 9 steps, only one rule is truly transferable: server ctx follows the client's estimation bias, not the VRAM cap — every tier increase happened "after overshooting the cap again" (see the performance page · token estimation bias).

⑦ Sleep Mechanism: Why "Clear" Must Be Sleep

The only correct implementation of "use when opened, return VRAM when idle" = process resident + idle sleep.

ActionPortVRAMWhen the client uses it
Exit the processNot listening0Connection refused → "reconnecting N/10"
SleepStill listening~2.4 GBRequest arrives → 7.6 seconds to wake → normal answer

Measured (--sleep-idle-seconds 180): in use 27,430 MB (incl. UE) → after 180 s idle ~2,400 MB (log handle_sleep: server is entering sleeping state) → next request returns 200 in 7.2 seconds. Early on, proxy-level kill implemented the release, but it had a timing bug and reverted to old parameters on restart — all abandoned. llama.cpp natively supports it; one layer is enough.