① Troubleshooting Quick-Reference Table
The first action is always status_local_llm.bat, then locate the issue against the table below.
| Sign | Meaning | Action |
|---|---|---|
llama-server 0 listeners | Model not loaded | start_local_llm.bat |
llama-server 2+ listeners | Two instances contending for VRAM (most dangerous) | stop then start |
shared > 1024 MB | Spilling into shared memory, throughput will collapse | Close VRAM-hogging programs; or lower ctx (sync contextWindow) |
Log tail has no listening on | Loading interrupted (something is killing the process) | Check proxy log for "launch attempt/cooling" |
| Reconnecting… N/10 | Server is returning 500 | Check llama_server.log for send_error / Jinja Exception |
| Reconnecting and message carries screenshots | Vision declaration mismatches real capability → image requests 500 | No local image paste / confirm mmproj is enabled |
Reconnecting and log has System message must be at the beginning | Old template rejects system messages mid-conversation | Restart the service with the patched template |
| Context compression failed | History exceeds contextWindow | Confirm the pairing relationship (see the integration page) |
in = above 97,000 | Prompt exceeds the client limit | Confirm contextWindow; or tools got more numerous again |
out = 256 | proxy's 256 output cap injection took effect again | Check whether openai_proxy.py was reverted to an old version |
Oversize request HTTP 400, log has capping | Server capped ctx at the model's training limit (262144) | yarn extrapolation + override-kv (36-10); or fallback Option A: -c 262144 + window 206000 |
What did ZCode actually send (fastest):
python D:\_Qwen3.8-27b\_zcode_setup\last_usage.py 10 REM in = 是本次请求的真实提示词大小;基础底噪应在 4.6 万左右 REM 如果还是 10 万+,说明工具开关没生效(需要重启 ZCode)
② Thirteen True Causes (click to see forensics)
Ordered by discovery date. Each is a closed loop of "symptom → root cause → fix → verify".
5.1Has a reply but stops after one sentencehigh
Root cause: proxy lines 609-610 if body.get("max_tokens", 0) < 16: body["max_tokens"] = 256 — when the client does not send max_tokens, 0 is received, the condition always holds, every turn gets forced to 256; thinking eats part first, leaving one sentence.
Evidence: 35 records in the ZCode database with output_tokens all 256.
Fix: delete the injection. Verify: out returns to normal length.
5.2Repeated invalid responseshigh
Root cause: provider config contextWindow: 1000000 is fake data (real 65536). ZCode believed there were 1 million, never compressed; measured requests of 130,170 and 382,703 tokens were sent → 400 → default 10-round retry.
Fix: fill the real value and link it with the server -c per the pairing formula. Verify: in/out data match real usage.
5.3Compression failed, 12 minuteshigh
Root cause: an old cloud session nurtured on a 1-million-context (2,128 messages / 25.8 MiB / history ~380K tokens) was switched to the local model — compression itself cannot fix it: cannot compress because the history must be sent, cannot send because it is too big. 380K tokens ≈ 13.4 GiB of pure KV, plus 29.9 GiB of weights, fills the whole card even without UE.
Fix: start a new session when using the local model; it is discipline, not a suggestion. Run session_sizes.py before switching.
5.4Reconnecting… 9/10 (reasoning_effort)
Root cause: llama-server log spammed with Jinja Exception: Unexpected reasoning effort high. — ZCode sends high by default; the template only accepts xhigh/medium/low and throws a 500.
Went the wrong way: neither --reasoning-effort xhigh nor --chat-template-kwargs helped — the value carried in the request has top priority.
Fix: print_patched_template.py reads back the current template from the running server and swaps raise_exception for "fall back to xhigh when unknown", with not one whitespace changed. Verify: all effort values return 200, thinking preserved.
5.5Direct reconnect right after boot
Root cause: a hand-crafted-parameter start_autostart.bat shortcut (from 9/14) in the startup folder: -c 16384 — smaller than ZCode's base prompt for every request (~38,000 tokens); every post-boot request 400s, and it lacks -fa / KV quantization / patched template.
Fix: boot through the unified launch script (idempotent + hidden). Verify: after boot, check the running command-line args with Get-CimInstance.
5.6"Compression failed" even in a new session (thin headroom)
Root cause: under contextWindow 61,440 a real 64,824 request went through (estimation bias ~5.5% low); the compression request added 2,563 instructions = 67,387 > 65,536. Only 1,100 token headroom was left; the real need is "estimation error + compression overhead" ≥ 5,900.
Fix: start with contextWindow lowered to 49,152 (headroom 11,114), then raise tier by tier per the pairing formula.
5.7Seven attached problems (fixed at once)multiple
① Two llama-server instances contending for VRAM: two instances measured listening on :11435 simultaneously, together demanding ~39 GiB and overflowing — the launch script now confirms "0 processes and port free" before starting; ② HARD_MSG_CAP dropping messages (log 丢 1 条, 29027→221 tokens); ③ _trim_messages can wipe messages (while other_msgs deletes everything when a single message exceeds budget, model receives an empty conversation); ④ launch-script false warning ('' is not an escape, health check always fails); ⑤ concurrent auto-restart livelock — "kill leftover" kills the previous still-loading process so it never finishes loading; fix: lock + 180 s cooling + LLM_SKIP_KILL=1 (regression: 6 concurrent → only 1 launch); ⑥ in VBS, cmd /c ""路径"" does not execute silently; ⑦ redirection locks the log handle — status line switched to Add-Content.
5.8Companion: bring the 38,000-token noise down
Noise = tool definitions + system prompts. Turn off two unused tool sources: the cloudbase-skills plugin (~40 tools), the chrome-devtools MCP (~25), keeping UE/Blender/Houdini workflow-related tools. Expected noise 38,000 → 18,000~20,000; usable conversation space 11,000 → ~30,000.
Note: plugin switches take effect only after a ZCode restart. Later measurement found the rider MCP alone occupies 57,596 tokens (see 5.13).
5.9Dropped after 3 compressions (rapid_refill breaker)
Root cause: ZCode breaker compact_rapid_refill_breaker — refilled in fewer than 3 tool turns after compression, 3 consecutive times and it gives up. The decision threshold is hard-coded in the program, not exposed as a setting. Two stacking causes: usable space too small + tool output especially large (fetching a full HTML page costs 10K~20K tokens per fetch).
Wrong direction: increasing compression count only postpones failure (each compression sends the whole history ≈ 25 s prefill; measured 8 compressions cost 17.6 minutes).
Fix: enlarge the window + do one thing at a time + read in chunks ("extract only the first 2000 chars of the body") + switch to the cloud when it genuinely cannot be handled.
5.10Cannot connect on every first boot of the daysubtle
Symptom signature: first boot of the day fails to connect, works after a while or after manual run — easily misdiagnosed as "wrong boot trigger timing".
Forensics: pulled the autostart.log timeline — after the .bat was rewritten with the Write tool on 09-18 12:10, the boot log was empty every time; manual execution output garbled fragments ('01' is not an internal or external command).
True cause: files written by the Write tool have LF line endings; cmd.exe reads batch files by byte offset and LF-only splits lines at the wrong place — the whole script runs as if it never ran, and leaves no trace. 22 scripts were affected (12 .bat / 8 .ps1 / 2 .vbs).
Fix: fix_crlf.ps1 converts to CRLF (UTF-8 no BOM) + redundant boot entry self-heals line endings on every boot + the rule is written into AGENTS.md (run manually once, on the spot, after any change). Do not judge line endings with Git Bash grep -c (it will deceive you); use LC_ALL=C tr -cd '\r' < 文件 | wc -c.
5.11Service is normal but "reconnecting" (message with screenshots)
Forensics: /health normal, process normal, but one message stalled for two minutes; the llama_server.log tail has 7 identical errors image input is not supported - you may need to provide the mmproj.
Root cause: that message carried 2 screenshots; the GGUF has no vision projector, and llama-server 500s every request with images; ZCode treated it as a transient fault and retried to 10. Meanwhile the provider config wrongly declared supportsImage: true.
Fix: set false + restart ZCode (switched back to true only after mmproj was enabled on 9/29).
Capability boundary: do not bypass with an "image-dropping" middle layer — the model cannot see the images yet answers anyway and will fabricate confidently, which is worse than direct failure.
5.12Still "reconnecting" after removing the images
Forensics: log Jinja Exception: System message must be at the beginning. — reproduced exactly with curl (system/user/system/user four segments); the error matches byte for byte what ZCode received.
Root cause: ZCode inserts a system reminder mid-conversation, while the official template hard-codes "system must be first", otherwise it throws — the client dares to send, the server refuses.
Fix: template patch two (mid-conversation system rendered as a system turn), the generation script made idempotent + dual-patch, hard error and exit if an anchor is not found. Verify: all three request shapes return 200.
Attached issue: PowerShell 5.1 reads a .ps1 without BOM as GBK (Chinese comments garbled), while a .bat with BOM invalidates the first line — fix_crlf.ps1 preserves the original BOM state.
5.13Single-turn request fills the entire context window
Measured (context_breakdown.py): first-turn prompt of a new session 103,762 tokens; compression request in the same session 14,151 — the difference ≈ 89,600 is all tool definitions + system prompts. Tokenized per source: rider MCP 102 tools = 57,596 tokens, monolith 28 = 11,887, unrealmcp only 323.
Conclusion: these directories are resent every request — "single turn fills the window" is unrelated to session duration. Not UnrealMCP's fault (runtime on-demand tool discovery).
Fix: remove rider from mcp.servers in config.json; after restart, noise 104K → 46K, freeing 35K of usable workspace. If it must be retained: check per tool group in the JetBrains MCP plugin settings, turning off xdebug/dotTrace/databases/Unity Profiler.
③ Three Shapes of Context Compression Failure
| Shape | Trigger | Way out |
|---|---|---|
| Deadlock (382K old session) | History too big: cannot compress because history must be sent, cannot send because too big | No fix — start a new session; never switch an old session |
| Thin headroom (new sessions explode too) | contextWindow differs from real usage by only ~1,100 tokens; the 2,563 compression instruction overflows it | Recompute and raise per the pairing formula |
| Breaker (rapid_refill) | Refilled in <3 tool turns after compression ×3 times; typical is fetching full HTML pages | Enlarge the window + split the task fine; the count is hard-coded and not adjustable |
Common background: compression summaries accumulate permanently in the history (+2~3.5K each time); the more you compress, the bigger the noise — so when you see "auto-compressed", consider a new session.
④ VRAM Contention (UE and model both open)
The standard answer is three layers:
- Prevention layer: the budget is measured with "UE open" (base occupancy 7,489 MB); ctx tiers are calibrated accordingly — do not use engine-off measurements for the budget;
- Automation layer:
--sleep-idle-seconds 180returns ~18 GiB to UE when idle; wake takes only 7.6 seconds; - Manual layer:
stop_local_llm.batreleases immediately. Judge overflow by shared memory:shared > 1024 MBmeans already shuttling over PCIe (measured ttft=104 seconds).
For the attribution method of "the model slowed by UE", see the performance page · 27 t/s mystery.
⑤ Two Points Easy to Misdiagnose
1. Process count ≠ instance count. This machine's Python has a launcher shell, so one proxy is two python.exe in the process table (parent-child). Judge duplicated instances by the number of listeners on the port — llama-server is a native exe; two processes really are two instances.
2. "Works at start, breaks later" ≠ a config problem. When a config that worked at first suddenly reconnects, check three things first: are there images in the messages (5.11), mid-conversation system messages (5.12), or a larger tool list (5.13) — all three are failures triggered only by specific input shapes.