← Back to project technical overview
① Three-Second Overview ② Architecture Panorama ③ Final Configuration ④ Issues and Fixes Timeline ⑤ Eleven True Causes ⑥ Deep Dive Reading

① Three-Second Overview

The whole chain has only four roles, each with one decisive fact.

Hardware

RTX 5090 32 GB

Blackwell architecture · CUDA 13.4. Whole card 32,607 MB; with the UE editor open at the same time, about 7.5 GB is held by the desktop + engine basics, leaving ~21 GB budget for the model.

Model

Bonsai-2 27B PQ2_0

Ternary-quantization base (1.71-bit), qwen35 hybrid attention: of 65 layers only 16 participate in the KV cache; the rest are fixed-size linear attention layers — this is what makes a 260K context possible.

Inference stack

llama.cpp b10754

Started from a CMake compile, now using the upstream prebuilt prism-b10754. KV cache double q4_0 quantized, Flash Attention on by default, all parameters centralized in one place: start_local_llm.bat.

Client

ZCode direct to :11435

openai-chat-completions native protocol + native tool_calls, no proxy of any kind (fewer layers, fewer failure classes). VRAM returned to UE after 180 s idle; 7.6 s to wake on next use.

② Architecture Panorama

One mainline, one bypass. Key decision: ZCode has no proxy; it connects directly to llama-server.

ZCode desktop openai-chat-completions contextWindow 256000 llama-server 127.0.0.1:11435 -c 327680 · KV q4_0 · fa on RTX 5090 32 GB Ternary-Bonsai-2-27B.gguf Weights + KV ≈ 20.5 GiB Sleep state ~2.8 GB openai_proxy.py :11436 · for CherryStudio only CherryStudio (optional) Ollama protocol /api/chat /v1/chat/completions GGUF · ngl 999 full offload Idle 180 s → sleep (VRAM 20.6 → 2.4 GB, port keeps listening) · 7.6 s to wake on next request

Final system structure: ZCode connects directly to llama-server (OpenAI-compatible protocol, native tool_calls); the proxy exists only for CherryStudio's Ollama protocol and does not occupy VRAM.

Why not go through a proxy: ZCode uses exactly openai-chat-completions, natively supported by llama-server, and tool calls are native too (measured finish_reason: tool_calls + native JSON arguments + results fed back in a second round). The only reason a proxy exists is that CherryStudio's Ollama provider goes through /api/chat — fewer layers, fewer failure classes.

③ Key Numbers (all measured, not estimated)

Six numbers carry the whole solution's judgment.

327,680

Server ctx (-c)

256,000

ZCode contextWindow
≈ 78% of server ctx

68–74

Real-load decode speed t/s
(KV4 + no MTP) (100K depth)

7.6 s

Sleep → full wake
(VRAM 2.4 → 20.6 GB)

−4.35 GB

VRAM benefit of KV4 quantization
Quality 7/7, zero degradation

99.7%

Prompt cache hit rate
(18,385 / 18,445)

④ Final Configuration Quick Reference

Two configuration sites + one pairing formula, current effective version (2026-10-04).

llama-server launch parameters (maintained uniformly by start_local_llm.bat; the parameters live only here):

llama-server.exe -m Ternary-Bonsai-2-27B-PQ2_0.gguf ^
  --host 127.0.0.1 --port 11435 ^
  -c 327680 ^
  -ngl 999 -np 1 ^
  --no-webui ^
  -fa on ^
  --rope-scaling yarn --yarn-orig-ctx 262144 ^
  --override-kv qwen35.context_length=int:327680 ^
  -ctk q4_0 -ctv q4_0 ^
  -b 2048 -ub 512 ^
  --no-reasoning-preserve ^
  --reasoning-budget 1000 ^
  --mmproj Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf ^
  --sleep-idle-seconds 180 ^
  --chat-template-file D:\_Qwen3.8-27b\_zcode_setup\qwen35_chat_template.jinja ^
  -a "local,qwen3.8:27b,qwen3.8:27b-nvfp4,qwen3.8:27b-offload" ^
  --metrics --log-file D:\_Qwen3.8-27b\_zcode_setup\llama_server.log

ZCode provider configuration (.zcode\v2\provider_config.json, capability declaration part):

{
  "api":     { "type": "openai-chat-completions",
               "baseUrl": "http://127.0.0.1:11435/v1" },
  "model":   { "properties": {
    "contextWindow": 256000,
    "supportsToolCall": true,
    "supportsMidConversationSystem": false,
    "inputFormat":  { "supportsText": true, "supportsImage": true },
    "outputFormat": { "supportsText": true }
  }}
}
Pairing rule: contextWindow × 1.26 (worst-case underestimation factor) + 2,563 (compression instructions) ≤ server -c → 256,000 × 1.26 + 2,563 ≈ 325,123 ≤ 327,680 ✓ (headroom 2,557) Changing one side requires changing the other.

⑤ Issues and Fixes Timeline (23 days)

From compilation to a 1.71-bit base swap: every node is a real incident or a measured verdict; details on each subpage.

2026-09-12

Compiled llama.cpp with VS2026 + CUDA 13.4

2026-09-13

NVFP4 model download complete, llama-server first launch, 66 t/s

2026-09-15

CherryStudio issue chain: ctx overflow, dual-process VRAM contention

2026-09-16

ctx 65536 + cont-batching stable; prefill 70% faster

2026-09-17

ZCode direct to :11435; false ctx / 256 output cap / dual instance all fixed

2026-09-20

LF line-ending accident: 22 scripts silently failed, all converted to CRLF

2026-09-22

Switched to model Ternary Bonsai 2 (ternary quantization), cut straight to :11435

2026-09-26

--reasoning-budget 1000: thinking −76%, body 0 → 4,820 chars

2026-09-29

mmproj vision enabled: real screenshots read word by word

2026-10-02

27 t/s mystery = UE viewport competing for GPU (not model degradation)

2026-10-04

KV4 adopted (−4.35 GB); MTP measured negative yield, rolled back

2026-10-04 (evening)

yarn extrapolation crossed the training line: -c 327680, 256K window, three-needle verification 3/3

⑥ Eleven True Causes at a Glance

Diagnosis conclusions in brief — none of the eleven is "the model is no good"; it is always that "what is fed to it / how it is launched" is wrong. Each expanded on the troubleshooting page.

#SymptomTrue causeFix
1Only one sentence, then it stopsproxy injected max_tokens=256Delete the injection
2Repeated invalid responsescontextWindow: 1000000 is fake dataChange to the real value
3Compression failed for 12 minutes380K-token old cloud session switched overStart a new session when using the local model
4Reconnect 9/10reasoning_effort: high template rejectionPatch the template / LLM_THINK=0
5Cannot connect after startupStartup script -c 16384Startup goes through the unified launch script
6New session also fails compressioncontextWindow headroom down to 1,100 tokensStart at 49152
7Dropped after 3 compressionsWindow too small + tool output large (fetching whole web pages)Enlarge the window
8context_exceeded 400ZCode estimate underestimates 26%, overshoots the server capUse server ctx as an absolute buffer
9First connect after boot fails daily.bat is LF line-ending, cmd.exe cannot executeConvert uniformly to CRLF
10Messages with screenshots reconnectModel has no vision yet declares supportsImage: trueSet to false (switched back to true after mmproj enabled)
11Still reconnects after image removedTemplate rejects the system message mid-conversationTemplate changed to render as a system turn

⑥→ Deep Dive (Four Subpages)

Split by problem domain: how the stack is built, how the client is wired, how to fix it when broken, and why the speed is what it is.

LLAMA-SERVER stack

Model and Inference Service

Hardware baseline · Quantization selection · Item-by-item explanation of compile and launch parameters · VRAM calibration formula · ctx tuning history (8192 → 327680) · Sleep mechanism

ZCODE integration

Client Configuration and Boot Chain

Provider configuration · The two-digit pairing · The 65% compression-line arithmetic · Modification checklist · Dual-path boot auto-start · 6 UE sub-agents

Troubleshooting manual

Quick-Reference Table and Thirteen True Causes

What to run first when "it doesn't respond again" · Symptom → root cause → fix → verify · The three shapes of compression failure · Each row expandable

Performance and tuning

Measured Data and Full Optimization History

Speed–context curves · UE competing for GPU attribution · Token estimation bias · reasoning-budget · Full KV4 and MTP review

Capability boundaries (memorize before touching anything):

  1. Big web pages / big documents → use a cloud model. Measured: summarizing a WeChat article (3.5 MB page), 8 compressions occupied 17.6 minutes of a 43-minute task — a volume–window mismatch, not a configuration problem.
  2. When you see "auto-compressed", prepare a new session — compression summaries accumulate in the history (each compression permanently adds 2~3.5K tokens), the only fixed cost that cannot be cleared.
  3. Do not switch old cloud sessions to the local model — history easily runs into the hundreds of thousands of tokens and necessarily triggers a compression deadlock.