① Three-Second Overview
The whole chain has only four roles, each with one decisive fact.
Hardware
RTX 5090 32 GB
Blackwell architecture · CUDA 13.4. Whole card 32,607 MB; with the UE editor open at the same time, about 7.5 GB is held by the desktop + engine basics, leaving ~21 GB budget for the model.
Model
Bonsai-2 27B PQ2_0
Ternary-quantization base (1.71-bit), qwen35 hybrid attention: of 65 layers only 16 participate in the KV cache; the rest are fixed-size linear attention layers — this is what makes a 260K context possible.
Inference stack
llama.cpp b10754
Started from a CMake compile, now using the upstream prebuilt prism-b10754. KV cache double q4_0 quantized, Flash Attention on by default, all parameters centralized in one place: start_local_llm.bat.
Client
ZCode direct to :11435
openai-chat-completions native protocol + native tool_calls, no proxy of any kind (fewer layers, fewer failure classes). VRAM returned to UE after 180 s idle; 7.6 s to wake on next use.
② Architecture Panorama
One mainline, one bypass. Key decision: ZCode has no proxy; it connects directly to llama-server.
Final system structure: ZCode connects directly to llama-server (OpenAI-compatible protocol, native tool_calls); the proxy exists only for CherryStudio's Ollama protocol and does not occupy VRAM.
Why not go through a proxy: ZCode uses exactly openai-chat-completions, natively supported by llama-server, and tool calls are native too (measured finish_reason: tool_calls + native JSON arguments + results fed back in a second round). The only reason a proxy exists is that CherryStudio's Ollama provider goes through /api/chat — fewer layers, fewer failure classes.
③ Key Numbers (all measured, not estimated)
Six numbers carry the whole solution's judgment.
327,680
Server ctx (-c)
256,000
ZCode contextWindow
≈ 78% of server ctx
68–74
Real-load decode speed t/s
(KV4 + no MTP) (100K depth)
7.6 s
Sleep → full wake
(VRAM 2.4 → 20.6 GB)
−4.35 GB
VRAM benefit of KV4 quantization
Quality 7/7, zero degradation
99.7%
Prompt cache hit rate
(18,385 / 18,445)
④ Final Configuration Quick Reference
Two configuration sites + one pairing formula, current effective version (2026-10-04).
llama-server launch parameters (maintained uniformly by start_local_llm.bat; the parameters live only here):
llama-server.exe -m Ternary-Bonsai-2-27B-PQ2_0.gguf ^ --host 127.0.0.1 --port 11435 ^ -c 327680 ^ -ngl 999 -np 1 ^ --no-webui ^ -fa on ^ --rope-scaling yarn --yarn-orig-ctx 262144 ^ --override-kv qwen35.context_length=int:327680 ^ -ctk q4_0 -ctv q4_0 ^ -b 2048 -ub 512 ^ --no-reasoning-preserve ^ --reasoning-budget 1000 ^ --mmproj Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf ^ --sleep-idle-seconds 180 ^ --chat-template-file D:\_Qwen3.8-27b\_zcode_setup\qwen35_chat_template.jinja ^ -a "local,qwen3.8:27b,qwen3.8:27b-nvfp4,qwen3.8:27b-offload" ^ --metrics --log-file D:\_Qwen3.8-27b\_zcode_setup\llama_server.log
ZCode provider configuration (.zcode\v2\provider_config.json, capability declaration part):
{
"api": { "type": "openai-chat-completions",
"baseUrl": "http://127.0.0.1:11435/v1" },
"model": { "properties": {
"contextWindow": 256000,
"supportsToolCall": true,
"supportsMidConversationSystem": false,
"inputFormat": { "supportsText": true, "supportsImage": true },
"outputFormat": { "supportsText": true }
}}
}
⑤ Issues and Fixes Timeline (23 days)
From compilation to a 1.71-bit base swap: every node is a real incident or a measured verdict; details on each subpage.
2026-09-12
Compiled llama.cpp with VS2026 + CUDA 13.4
2026-09-13
NVFP4 model download complete, llama-server first launch, 66 t/s
2026-09-15
CherryStudio issue chain: ctx overflow, dual-process VRAM contention
2026-09-16
ctx 65536 + cont-batching stable; prefill 70% faster
2026-09-17
ZCode direct to :11435; false ctx / 256 output cap / dual instance all fixed
2026-09-20
LF line-ending accident: 22 scripts silently failed, all converted to CRLF
2026-09-22
Switched to model Ternary Bonsai 2 (ternary quantization), cut straight to :11435
2026-09-26
--reasoning-budget 1000: thinking −76%, body 0 → 4,820 chars
2026-09-29
mmproj vision enabled: real screenshots read word by word
2026-10-02
27 t/s mystery = UE viewport competing for GPU (not model degradation)
2026-10-04
KV4 adopted (−4.35 GB); MTP measured negative yield, rolled back
2026-10-04 (evening)
yarn extrapolation crossed the training line: -c 327680, 256K window, three-needle verification 3/3
⑥ Eleven True Causes at a Glance
Diagnosis conclusions in brief — none of the eleven is "the model is no good"; it is always that "what is fed to it / how it is launched" is wrong. Each expanded on the troubleshooting page.
| # | Symptom | True cause | Fix |
|---|---|---|---|
| 1 | Only one sentence, then it stops | proxy injected max_tokens=256 | Delete the injection |
| 2 | Repeated invalid responses | contextWindow: 1000000 is fake data | Change to the real value |
| 3 | Compression failed for 12 minutes | 380K-token old cloud session switched over | Start a new session when using the local model |
| 4 | Reconnect 9/10 | reasoning_effort: high template rejection | Patch the template / LLM_THINK=0 |
| 5 | Cannot connect after startup | Startup script -c 16384 | Startup goes through the unified launch script |
| 6 | New session also fails compression | contextWindow headroom down to 1,100 tokens | Start at 49152 |
| 7 | Dropped after 3 compressions | Window too small + tool output large (fetching whole web pages) | Enlarge the window |
| 8 | context_exceeded 400 | ZCode estimate underestimates 26%, overshoots the server cap | Use server ctx as an absolute buffer |
| 9 | First connect after boot fails daily | .bat is LF line-ending, cmd.exe cannot execute | Convert uniformly to CRLF |
| 10 | Messages with screenshots reconnect | Model has no vision yet declares supportsImage: true | Set to false (switched back to true after mmproj enabled) |
| 11 | Still reconnects after image removed | Template rejects the system message mid-conversation | Template changed to render as a system turn |
⑥→ Deep Dive (Four Subpages)
Split by problem domain: how the stack is built, how the client is wired, how to fix it when broken, and why the speed is what it is.
LLAMA-SERVER stack
Model and Inference Service
Hardware baseline · Quantization selection · Item-by-item explanation of compile and launch parameters · VRAM calibration formula · ctx tuning history (8192 → 327680) · Sleep mechanism
ZCODE integration
Client Configuration and Boot Chain
Provider configuration · The two-digit pairing · The 65% compression-line arithmetic · Modification checklist · Dual-path boot auto-start · 6 UE sub-agents
Troubleshooting manual
Quick-Reference Table and Thirteen True Causes
What to run first when "it doesn't respond again" · Symptom → root cause → fix → verify · The three shapes of compression failure · Each row expandable
Performance and tuning
Measured Data and Full Optimization History
Speed–context curves · UE competing for GPU attribution · Token estimation bias · reasoning-budget · Full KV4 and MTP review
Capability boundaries (memorize before touching anything):
- Big web pages / big documents → use a cloud model. Measured: summarizing a WeChat article (3.5 MB page), 8 compressions occupied 17.6 minutes of a 43-minute task — a volume–window mismatch, not a configuration problem.
- When you see "auto-compressed", prepare a new session — compression summaries accumulate in the history (each compression permanently adds 2~3.5K tokens), the only fixed cost that cannot be cleared.
- Do not switch old cloud sessions to the local model — history easily runs into the hundreds of thousands of tokens and necessarily triggers a compression deadlock.