① Model Specs (Parsed from the GGUF Header, Not Estimated)
general.architecture = qwen35 (Qwen3.5 hybrid attention) total parameters = 27.32B (dense model, not MoE) block_count = 65 (blk.64 is the nextn/MTP layer, ignored by llama.cpp) attention.head_count = 24 attention.head_count_kv = 4 (KV has only 4 heads) key_length/value_length = 256 (256 dims per head) full_attention_interval = 4 (only 1 of every 4 layers is full attention) ssm.* (the remaining ~49 layers are linear-attention GatedDeltaNet) file size = 19,653,896,608 bytes (18.30 GiB, nvfp4 era)
Key inference: only 16 layers participate in the KV cache (layers where (i+1)%4==0); the other 49 are SSM linear attention — fixed-size state that does not grow with tokens. This is where a same-size traditional Transformer overflows VRAM, and it is why this model can run a 260K context. Prefill is also extremely fast: 64K-token prompt in 33.5 seconds (1,905 t/s), prompt cache hit rate 99.7%.
② Generation Speed Drops with Context Length (measured, not a fault)
On 2026-09-21 the eval time was read directly from llama_server.log. Every generated token attends over all cached KV — longer context, slower; it is arithmetic.
Overflow confirmed absent at the same time: shared VRAM only 505 MB (within budget), dedicated VRAM matches the calibration. System memory 86% is another matter — the 21 GB model resides via mmap; memory pressure only affects load speed, not generation speed (KV is all in VRAM).
Three levers if it feels slow (by cost-effectiveness): ① LLM_THINK=0 to turn off the thinking chain, ~2× faster (or use a better --reasoning-budget, see ⑤); ② shorten the context (cost: compression happens more often); ③ reduce tool noise so the same length carries more useful content. Self-check: findstr /C:"eval time" llama_server.log.
③ The 27 t/s Mystery = the UE Editor Competing for GPU
Speed flips between two tiers of 26~28 and 71 t/s (12h log: 1,780 fast / 74 slow), and it jumps even mid-task.
| Troubleshooting action | Conclusion |
|---|---|
| Same params, same model, the 71 t/s fast tier exists | Not quantization/model degradation — capability is still there |
Identify by process with Get-Counter '\GPU Engine(*)' | UE viewport real-time rendering resident at 42~65% GPU; the lighting control program MythCool takes another 9~10% |
| Probing request when UE is quiet | 70 t/s; 35.79 t/s when UE is rendering |
| VRAM check | No overflow — it is not a VRAM problem; compute is being siphoned off |
Countermeasures (while the model is generating): minimize UE / Ctrl+R to disable viewport real-time rendering / console t.MaxFPS 30 / check "Use Less CPU in Background" in editor preferences.
④ ZCode's Token Estimation Cannot Be Trusted (three measurements, unstable bias)
| Scenario | ZCode believes | Actually sent | Underestimate |
|---|---|---|---|
| Scenario 1 | ≤61,440 | 64,824 | 5.5% |
| Scenario 2 | ≤65,536 | 75,208 | 15% |
| Scenario 3 | ≤65,536 | 82,324 | 26% |
The third one directly 400'd the server: request (82324 tokens) exceeds the available context size (81920 tokens). Therefore: llama-server's --context-shift is off by default, overshoot gives a hard error, not a silent truncate; ZCode's contextWindow is not protection, only a "when to start compressing" trigger point — compression cannot stop it from overshooting on its own. The server ctx must be a large margin above the client cap as an absolute buffer (pairing formula on the integration page).
⑤ Curing Overthinking: --reasoning-budget 1000
Background: one UE task ran 2h36m; item-by-item decomposition found 85% was model generation, and the model often burns the entire output budget on thinking, with a single word in the body.
Probe measurement (the same UE reasoning problem, max_tokens=4000):
| Config | thinking | content | finish |
|---|---|---|---|
| As-is (no budget) | 10,037 chars | 0 chars | length |
--reasoning-budget 2500 | 2,197 chars | 5,368 chars | ✅ |
--reasoning-budget 1000 (shipped) | 2,430 chars | 4,820 chars | ✅ |
Root cause: the template injects a "please verify key assumptions, consider alternatives" system instruction on the xhigh tier — that sentence is literally teaching the model to overthink, and the high that ZCode sends every turn is forced to xhigh, so every turn gets this instruction.
Accuracy comparison (speed alone is not enough): three tests — sequence reasoning / instruction following / tool_calls — all matched the as-is answers at RB=1000 — what was cut was indeed useless content (the model itself spotted the numeric contradiction in the prompt within the body, so reasoning quality held).
Tried but reverted: forcing the template to medium — measured two slightly better, one slightly worse, one unstable, all within noise, and "unverified changes do not stay in production", so it was restored byte-for-byte. All three candidate routes (Swift-1.5+kvmem, NVFP4-Q8mix, Swift-Bonsai-2) were also measured or evaluated and rejected — Swift's claim of "58.5% less thinking" did not reproduce (on the same problem and budget it actually thought more); vendor claims cannot be trusted directly.
⑥ Four Optimizations in Full: KV4 Adopted, MTP Rolled Back
Executed on 2026-10-04 by "build the baseline first, optimize item by item, self-check every step". Stages: T0 = before optimization; T_KV4 / T_MTP = intermediate states.
| Optimization item | Result | Verdict |
|---|---|---|
| ① Sampling parameter check | Server defaults are exactly the officially recommended values (temp 1.0 / top_p 0.95 / top_k 20) | This optimization does not exist, zero changes |
| ② KV4 (-ctk/-ctv q4_0) | Quality 7/7 zero degradation · 27.5K+100K retrieval PASS · speed +13.6% · whole-card VRAM −4.35 GB | ✅ Adopted |
| ③ MTP speculative decoding | Three routes tried (community MTP model → Hadamard fix missing; kvmem package → missing sleep-idle; upgrade to prism-b10754 → worked, +23% on the server side) | ⚠️ Rollback (see below) |
| ④ Baseline test script | llm_bench.py: 7-question quality + long-text retrieval + speed, zero file writes | ✅ Settled |
Why MTP was rolled back — user's measured feedback: "feels slower: VRAM indeed dropped, but GPU utilization is only in the twenties". Same-time comparison at ~100K context:
| Config | Synthetic request (fresh prefill) | User session (prefix reuse) |
|---|---|---|
| Original (q8_0 / no MTP) | 67.7 t/s | ~68 t/s |
| KV4 + MTP | 125~142 t/s | 32~54 t/s ⚠️ |
| KV4 + no MTP (final) | 70.3 t/s | 68~74 t/s |
Root cause (upstream known defect llama.cpp #28049): speculative decoding's boundary states are bound to the exact last position of the cached prefix. Real agent traffic is "partial prefix hit" (this machine's log f_sim=0.997 / f_keep=0.993) — 0.3% off and cache reuse fails, every step recomputed. A synthetic fresh prefill or exact hit cannot expose this problem, so the user's "feels slower" is right and the machine data confirms it. During troubleshooting, four confounders were also ruled out one by one (all measured): the context cliff, VRAM overflow, tool schema, and UE using the GPU.
Final effective config: prism-b10754 + Bonsai-2-27B-PQ2_0 + -ctk q4_0 -ctv q4_0 + -c 327680 + yarn + override-kv + no --spec-type (MTP rollback conclusion unchanged) + all other params unchanged. ZCode side contextWindow 256000. MTP model files are retained (usable with llm_bench.py --reuse to retest once upstream fixes them). Rollback credentials are triple-layered (.bak script + old binary directory intact + bench raw data).
⑥ Verification of 1.25× Context Extrapolation (2026-10-04 evening)
yarn extrapolation is not just "turn up -c" — it must be verified that it holds up after crossing the training line. Verdict: all three needles green; the speed cost is only paid on super-long contexts.
Three-needle embedding method (source 36-10): three different passwords embedded in the same prompt, at 10% / 50% / 90% depth (testing full-range retrieval, not just the tail), plus a math reasoning problem:
| Target depth | Actual prompt | Needle hits | Reasoning question | Elapsed |
|---|---|---|---|---|
| 150K | 141,636 tok | 3/3 ✓ | ✓ | 79 s |
| 260K | 245,748 tok | 3/3 ✓ | ✓ | 175 s |
| 310K (1.12× training line) | 293,088 tok | 3/3 ✓ | ✓ | 230 s |
Speed: decode ≈ 27.8 t/s at 300K depth, prefill ≈ 1,000 t/s — ~100-130K daily session speed unchanged (the length-correlated decode decay is gradual; the cost is only paid on long contexts).
A/B decision: Option A (conservative -c 262144 + window ≈206,000, not adopted) vs Option B (yarn extrapolation -c 327680 + window 256,000, adopted). Fallback path = Option A (drop the 3 new params, -c back to 262144, window 206000); A/B are two independent knobs — falling back to A does not require giving up the KV4 and b10754 gains.
Limitations (recorded honestly): each depth ran once, 3 needles, reasoning verified on only 1 problem — "can retrieve, can reason" holds, but no controlled comparison with 100K-depth inference of equivalent quality. For critical decisions in super-long sessions, important conclusions are still recommended to be reread and checked.
VRAM cost: whole-card peak 22.5 GB after extrapolation (+2.6 GB), still ~10 GB left for UE.
⑦ Methodology Lessons (Biggest Gain This Round)
Speed-type verification must test both request shapes at the same time: ① fresh prefill; ② prefix reuse (send the same prompt twice in a row, simulating real agent per-turn accumulation). The conclusions can be opposite — local MTP was 110~142 t/s for the first, 32~54 t/s for the second. When synthetic tests fail to reproduce real traffic shapes, a negative yield gets verified as positive. The baseline phase had the same kind of problem: test questions given too-small max_tokens (160~320), the thinking budget didn't burn out before the body got its turn, three empty answers — that was a script problem, not a model problem.
Another engineering detail: probe/benchmark scripts must be zero file writes (results to stdout, summaries to stderr) — this avoids path-traversal interception by security scans and leaves no garbage behind.