① Quick-Reference Table ② Thirteen True Causes ③ Three Shapes of Compression Failure ④ VRAM Contention ⑤ Easy Misdiagnosis Points

① Troubleshooting Quick-Reference Table

The first action is always status_local_llm.bat, then locate the issue against the table below.

SignMeaningAction
llama-server 0 listenersModel not loadedstart_local_llm.bat
llama-server 2+ listenersTwo instances contending for VRAM (most dangerous)stop then start
shared > 1024 MBSpilling into shared memory, throughput will collapseClose VRAM-hogging programs; or lower ctx (sync contextWindow)
Log tail has no listening onLoading interrupted (something is killing the process)Check proxy log for "launch attempt/cooling"
Reconnecting… N/10Server is returning 500Check llama_server.log for send_error / Jinja Exception
Reconnecting and message carries screenshotsVision declaration mismatches real capability → image requests 500No local image paste / confirm mmproj is enabled
Reconnecting and log has System message must be at the beginningOld template rejects system messages mid-conversationRestart the service with the patched template
Context compression failedHistory exceeds contextWindowConfirm the pairing relationship (see the integration page)
in = above 97,000Prompt exceeds the client limitConfirm contextWindow; or tools got more numerous again
out = 256proxy's 256 output cap injection took effect againCheck whether openai_proxy.py was reverted to an old version
Oversize request HTTP 400, log has cappingServer capped ctx at the model's training limit (262144)yarn extrapolation + override-kv (36-10); or fallback Option A: -c 262144 + window 206000

What did ZCode actually send (fastest):

python D:\_Qwen3.8-27b\_zcode_setup\last_usage.py 10
REM in = 是本次请求的真实提示词大小;基础底噪应在 4.6 万左右
REM 如果还是 10 万+,说明工具开关没生效(需要重启 ZCode)

② Thirteen True Causes (click to see forensics)

Ordered by discovery date. Each is a closed loop of "symptom → root cause → fix → verify".

5.1Has a reply but stops after one sentencehigh

Root cause: proxy lines 609-610 if body.get("max_tokens", 0) < 16: body["max_tokens"] = 256 — when the client does not send max_tokens, 0 is received, the condition always holds, every turn gets forced to 256; thinking eats part first, leaving one sentence.

Evidence: 35 records in the ZCode database with output_tokens all 256.

Fix: delete the injection. Verify: out returns to normal length.

5.2Repeated invalid responseshigh

Root cause: provider config contextWindow: 1000000 is fake data (real 65536). ZCode believed there were 1 million, never compressed; measured requests of 130,170 and 382,703 tokens were sent → 400 → default 10-round retry.

Fix: fill the real value and link it with the server -c per the pairing formula. Verify: in/out data match real usage.

5.3Compression failed, 12 minuteshigh

Root cause: an old cloud session nurtured on a 1-million-context (2,128 messages / 25.8 MiB / history ~380K tokens) was switched to the local model — compression itself cannot fix it: cannot compress because the history must be sent, cannot send because it is too big. 380K tokens ≈ 13.4 GiB of pure KV, plus 29.9 GiB of weights, fills the whole card even without UE.

Fix: start a new session when using the local model; it is discipline, not a suggestion. Run session_sizes.py before switching.

5.4Reconnecting… 9/10 (reasoning_effort)

Root cause: llama-server log spammed with Jinja Exception: Unexpected reasoning effort high. — ZCode sends high by default; the template only accepts xhigh/medium/low and throws a 500.

Went the wrong way: neither --reasoning-effort xhigh nor --chat-template-kwargs helped — the value carried in the request has top priority.

Fix: print_patched_template.py reads back the current template from the running server and swaps raise_exception for "fall back to xhigh when unknown", with not one whitespace changed. Verify: all effort values return 200, thinking preserved.

5.5Direct reconnect right after boot

Root cause: a hand-crafted-parameter start_autostart.bat shortcut (from 9/14) in the startup folder: -c 16384 — smaller than ZCode's base prompt for every request (~38,000 tokens); every post-boot request 400s, and it lacks -fa / KV quantization / patched template.

Fix: boot through the unified launch script (idempotent + hidden). Verify: after boot, check the running command-line args with Get-CimInstance.

5.6"Compression failed" even in a new session (thin headroom)

Root cause: under contextWindow 61,440 a real 64,824 request went through (estimation bias ~5.5% low); the compression request added 2,563 instructions = 67,387 > 65,536. Only 1,100 token headroom was left; the real need is "estimation error + compression overhead" ≥ 5,900.

Fix: start with contextWindow lowered to 49,152 (headroom 11,114), then raise tier by tier per the pairing formula.

5.7Seven attached problems (fixed at once)multiple

① Two llama-server instances contending for VRAM: two instances measured listening on :11435 simultaneously, together demanding ~39 GiB and overflowing — the launch script now confirms "0 processes and port free" before starting; ② HARD_MSG_CAP dropping messages (log 丢 1 条, 29027→221 tokens); ③ _trim_messages can wipe messages (while other_msgs deletes everything when a single message exceeds budget, model receives an empty conversation); ④ launch-script false warning ('' is not an escape, health check always fails); ⑤ concurrent auto-restart livelock — "kill leftover" kills the previous still-loading process so it never finishes loading; fix: lock + 180 s cooling + LLM_SKIP_KILL=1 (regression: 6 concurrent → only 1 launch); ⑥ in VBS, cmd /c ""路径"" does not execute silently; ⑦ redirection locks the log handle — status line switched to Add-Content.

5.8Companion: bring the 38,000-token noise down

Noise = tool definitions + system prompts. Turn off two unused tool sources: the cloudbase-skills plugin (~40 tools), the chrome-devtools MCP (~25), keeping UE/Blender/Houdini workflow-related tools. Expected noise 38,000 → 18,000~20,000; usable conversation space 11,000 → ~30,000.

Note: plugin switches take effect only after a ZCode restart. Later measurement found the rider MCP alone occupies 57,596 tokens (see 5.13).

5.9Dropped after 3 compressions (rapid_refill breaker)

Root cause: ZCode breaker compact_rapid_refill_breaker — refilled in fewer than 3 tool turns after compression, 3 consecutive times and it gives up. The decision threshold is hard-coded in the program, not exposed as a setting. Two stacking causes: usable space too small + tool output especially large (fetching a full HTML page costs 10K~20K tokens per fetch).

Wrong direction: increasing compression count only postpones failure (each compression sends the whole history ≈ 25 s prefill; measured 8 compressions cost 17.6 minutes).

Fix: enlarge the window + do one thing at a time + read in chunks ("extract only the first 2000 chars of the body") + switch to the cloud when it genuinely cannot be handled.

5.10Cannot connect on every first boot of the daysubtle

Symptom signature: first boot of the day fails to connect, works after a while or after manual run — easily misdiagnosed as "wrong boot trigger timing".

Forensics: pulled the autostart.log timeline — after the .bat was rewritten with the Write tool on 09-18 12:10, the boot log was empty every time; manual execution output garbled fragments ('01' is not an internal or external command).

True cause: files written by the Write tool have LF line endings; cmd.exe reads batch files by byte offset and LF-only splits lines at the wrong place — the whole script runs as if it never ran, and leaves no trace. 22 scripts were affected (12 .bat / 8 .ps1 / 2 .vbs).

Fix: fix_crlf.ps1 converts to CRLF (UTF-8 no BOM) + redundant boot entry self-heals line endings on every boot + the rule is written into AGENTS.md (run manually once, on the spot, after any change). Do not judge line endings with Git Bash grep -c (it will deceive you); use LC_ALL=C tr -cd '\r' < 文件 | wc -c.

5.11Service is normal but "reconnecting" (message with screenshots)

Forensics: /health normal, process normal, but one message stalled for two minutes; the llama_server.log tail has 7 identical errors image input is not supported - you may need to provide the mmproj.

Root cause: that message carried 2 screenshots; the GGUF has no vision projector, and llama-server 500s every request with images; ZCode treated it as a transient fault and retried to 10. Meanwhile the provider config wrongly declared supportsImage: true.

Fix: set false + restart ZCode (switched back to true only after mmproj was enabled on 9/29).

Capability boundary: do not bypass with an "image-dropping" middle layer — the model cannot see the images yet answers anyway and will fabricate confidently, which is worse than direct failure.

5.12Still "reconnecting" after removing the images

Forensics: log Jinja Exception: System message must be at the beginning. — reproduced exactly with curl (system/user/system/user four segments); the error matches byte for byte what ZCode received.

Root cause: ZCode inserts a system reminder mid-conversation, while the official template hard-codes "system must be first", otherwise it throws — the client dares to send, the server refuses.

Fix: template patch two (mid-conversation system rendered as a system turn), the generation script made idempotent + dual-patch, hard error and exit if an anchor is not found. Verify: all three request shapes return 200.

Attached issue: PowerShell 5.1 reads a .ps1 without BOM as GBK (Chinese comments garbled), while a .bat with BOM invalidates the first line — fix_crlf.ps1 preserves the original BOM state.

5.13Single-turn request fills the entire context window

Measured (context_breakdown.py): first-turn prompt of a new session 103,762 tokens; compression request in the same session 14,151 — the difference ≈ 89,600 is all tool definitions + system prompts. Tokenized per source: rider MCP 102 tools = 57,596 tokens, monolith 28 = 11,887, unrealmcp only 323.

Conclusion: these directories are resent every request — "single turn fills the window" is unrelated to session duration. Not UnrealMCP's fault (runtime on-demand tool discovery).

Fix: remove rider from mcp.servers in config.json; after restart, noise 104K → 46K, freeing 35K of usable workspace. If it must be retained: check per tool group in the JetBrains MCP plugin settings, turning off xdebug/dotTrace/databases/Unity Profiler.

③ Three Shapes of Context Compression Failure

ShapeTriggerWay out
Deadlock (382K old session)History too big: cannot compress because history must be sent, cannot send because too bigNo fix — start a new session; never switch an old session
Thin headroom (new sessions explode too)contextWindow differs from real usage by only ~1,100 tokens; the 2,563 compression instruction overflows itRecompute and raise per the pairing formula
Breaker (rapid_refill)Refilled in <3 tool turns after compression ×3 times; typical is fetching full HTML pagesEnlarge the window + split the task fine; the count is hard-coded and not adjustable

Common background: compression summaries accumulate permanently in the history (+2~3.5K each time); the more you compress, the bigger the noise — so when you see "auto-compressed", consider a new session.

④ VRAM Contention (UE and model both open)

The standard answer is three layers:

  1. Prevention layer: the budget is measured with "UE open" (base occupancy 7,489 MB); ctx tiers are calibrated accordingly — do not use engine-off measurements for the budget;
  2. Automation layer: --sleep-idle-seconds 180 returns ~18 GiB to UE when idle; wake takes only 7.6 seconds;
  3. Manual layer: stop_local_llm.bat releases immediately. Judge overflow by shared memory: shared > 1024 MB means already shuttling over PCIe (measured ttft=104 seconds).

For the attribution method of "the model slowed by UE", see the performance page · 27 t/s mystery.

⑤ Two Points Easy to Misdiagnose

1. Process count ≠ instance count. This machine's Python has a launcher shell, so one proxy is two python.exe in the process table (parent-child). Judge duplicated instances by the number of listeners on the port — llama-server is a native exe; two processes really are two instances.

2. "Works at start, breaks later" ≠ a config problem. When a config that worked at first suddenly reconnects, check three things first: are there images in the messages (5.11), mid-conversation system messages (5.12), or a larger tool list (5.13) — all three are failures triggered only by specific input shapes.