① Provider Configuration (Copy-Paste Ready)
File .zcode\v2\provider_config.json. Values are the current effective version; evolution on ②③.
// Provider body
{
"providerId": "88ff30fa-af55-4443-a1ff-ba8a7ff681c6",
"providerName": "Local Ollama",
"config": {
"group": "standard-personal",
"access": { "type": "api-key", "apiKey": "ollama-local-no-key" },
"api": { "type": "openai-chat-completions",
"baseUrl": "http://127.0.0.1:11435/v1" },
"personalModelIds": ["qwen3.8:27b-offload"],
"modelOrder": ["qwen3.8:27b-offload"]
}
}
// Model rules (capability declarations must match the model's real capabilities)
{
"modelId": "qwen3.8:27b-offload",
"config": { "properties": {
"contextWindow": 256000,
"requiresMfjsToolSchema": false,
"supportsToolCall": true,
"supportsJsonSchemaOutput": false,
"supportsMidConversationSystem": false,
"inputFormat": { "supportsText": true, "supportsImage": true },
"outputFormat": { "supportsText": true }
}}
}
Wrong capability-declaration fields create faults directly: on 2026-09-20 supportsImage was declared true while the model had no vision projector → image requests 500 in bulk → "reconnecting N/10" retry storm. After enabling mmproj (9/29) it was truly true. The contextWindow change 190000 → 256000 landed in two provider_config.json copies (G/E), and provider config changes require a ZCode restart — it reads it only at startup.
② Two Numbers Must Be Maintained in Pair
Server -c and client contextWindow — adjusting one side only guarantees the fault recurs.
| Server -c | contextWindow | Actual usage (UE open) | Judgment |
|---|---|---|---|
| 65536 | 49152 | ~18,900 MB | Conservative, but ZCode tends to overshoot |
| 81920 | 65536 | ~19,200 MB | Measured overshoot requests did occur (82,324 > 81,920) |
| 114688 | 81920 | ~20,608 MB | Stable operation achieved |
| 262144 | 190000 | ~20.5 GiB | Superseded by the 327680 tier (10-04 evening) |
| 327680 | 256000 | ~22.5 GB | ✅ Current (yarn extrapolation 1.25×, headroom 2,557) |
Where the 1.26 coefficient comes from: three measurements in the same environment, ZCode's token estimate underestimating 5.5% → 15% → 26% (unstable bias). Compression requests also add ~2,563 token instructions on top of the history. So the check formula must take the worst observed value and then keep a margin — do not configure "just enough". Counterexample: 98,304 × 1.26 + 2,563 ≈ 126,426 > 114,688, so the old-era trigger point cannot be raised to 98,304.
③ Why the Compression Line Is 65%, Not 85%
ZCode's auto compression line is not "a percentage of the window", it is "the window first fixed-deducts 34,000 tokens".
| contextWindow | Actual trigger line | Trigger percentage |
|---|---|---|
| 49,152 | 15,152 | 30.8% |
| 65,536 | 32,536 | 48.1% |
| 98,304 → 190,000 reuses the same formula | 64,304 (at 98,304) | 65.4% |
| 256,000 | 222,000 | 86.7% |
| 1,000,000 (cloud) | 966,000 | 96.6% |
Both reserves are hard-coded constants — 185 compression events, across 6 window tiers, including the cloud 1M model, the values never changed. So the "compress at 85%" intuition holds only at large windows. Two practical conclusions:
- The only effective way to gain usable space is to enlarge the window: 81,920 → 98,304 moved the trigger line from 47,920 → 64,304, usable tokens +34%;
- The denominator of the panel's "used X/limit (P%)" is the window, not the trigger line — seeing compression at 70% means actual tokens have already crossed that line.
Reproduction command: python D:\_Qwen3.8-27b\_zcode_setup\_compact_probe.py (parses ZCode's jsonl logs).
④ Change List (new / modified / backup)
The full picture of files created and modified during integration — in troubleshooting, first think "which file does this feature live in".
New files (core items):
| File | Role |
|---|---|
start_local_llm.bat | The single launch entry: kill leftovers first and confirm the port is truly free, then launch; wait for health; print VRAM |
stop_local_llm.bat | Release VRAM immediately. Kills only llama-server, not random python.exe |
status_local_llm.bat | One-click health check (process/listeners/health/VRAM/shared memory/logs/request records) |
autostart_local_llm.bat + .vbs | Idempotent boot starter + hidden window wrapper |
_zcode_setup\qwen35_chat_template.jinja | Patched chat template (fixes the two 500 root causes) |
_zcode_setup\is_ready.ps1 / wait_ready.ps1 | Silent health probe / wait-for-load (replacing unreliable quote escaping in batch) |
_zcode_setup\last_usage.py | Reads ZCode's sqlite, prints in/out/ttft of recent requests — fastest forensics of "what did ZCode actually send" |
_zcode_setup\session_sizes.py | Lists each session's real max prompt, judges whether a local model can run it |
_zcode_setup\api_test.py / zcode_sim.py | 6 API self-checks / simulate ZCode request shapes (96 tools + 18K prompt) |
Modified files (each maps to a true cause):
| File | Change | Reason |
|---|---|---|
| provider_config.json | baseUrl 11436 → 11435/v1; contextWindow 1000000 → real value (now 190000); capability declarations completed | Direct connect; fake ctx never compresses; wrong declarations = 500 storm |
| openai_proxy.py | Delete max_tokens injection / delete HARD_MSG_CAP second trim / fix _trim_messages wiping all / add lock + cooling | 256 output cap injection; whole messages dropped; empty conversation; concurrent livelock |
| Startup folder .lnk | Point to running the .vbs via wscript | Boot hidden window + idempotent start |
| ZCode config.json | Disable cloudbase-skills plugin, remove chrome-devtools / rider MCP | Tool schemas are a window killer (noise 38K → 46K) |
All changes have timestamped backups (*.bak_date_time); rollback is copy-overwrite.
⑤ Boot Chain (Two, Independent of Each Other)
Design goal: model auto-loads 20 seconds after boot; each path is idempotent (probe /health first), and simultaneous triggers do not interrupt a loading model.
Main path (Startup folder):
NVFP4_llama-server.lnk
→ wscript.exe autostart_local_llm.vbs (fully hidden window)
→ autostart_local_llm.bat (wait 20s for GPU driver)
→ is_ready.ps1 probes /health
├─ already 200 → log one line, do nothing (idempotent)
└─ not serving → start_local_llm.bat to launch
→ also raises CherryStudio's proxy (if 11436 is free)
→ raises watchdog (60s cycle; restarts if the service disappears)
Redundant path (Registry Run key LocalLLM_ensure):
powershell.exe -WindowStyle Hidden -File ensure_model.ps1
→ First self-heal: convert LF line endings of .bat/.cmd/.vbs/.ps1 to CRLF in place
→ Then call the same autostart_local_llm.bat
Why the redundant path: the main path once failed silently for two days, with no log trace at all (the LF line-ending accident). The redundant path is written in PowerShell (unaffected by that trap) and self-heals line endings every boot. The Task Scheduler version is stronger but was rejected — Register-ScheduledTask reports 0x80070005 access denied in an unprivileged environment.
Troubleshooting order for "not usable after boot": status_local_llm.bat → by [1] Processes (0 = model is gone) → [9] Watchdog (running = wait 1 minute for auto-raise; NOT running = manual cscript //nologo //B watchdog.vbs) → [8] Logon autostart (whether it ran at boot) → [5] log tail (should have listening on).
⑥ Six UE Sub-Agents
The motive is saving tokens: the full tool schema eats ~46K, while each sub-agent carries only 6 basic tools (Read/Write/Edit/Glob/Grep/Bash), dropping the schema to ~7K and freeing 39K for the task itself.
| # | File | Positioning | Loop cap |
|---|---|---|---|
| 1 | ue-local-runner.md | Primary executor: run UE tasks for the local model | 2 |
| 2 | ue-anim-debugger.md | Animation tracing: evidence chain for "why did it play X" | 3 |
| 3 | ue-pcg-builder.md | PCG scene generation: node chains + empty-generation troubleshooting | 4 |
| 4 | ue-asset-explorer.md | Read-only investigator: Blueprint variables / material params / CDO | None (read-only) |
| 5 | ue-build-runner.md | Build + verify: Live Coding / UBT full | None (fixed flow) |
| 6 | project-rule-guard.md | Compliance review: read rule files to verify changes | None |
Loop design philosophy (a second-iteration conclusion): the true state of UE assets can only be known by reading the CDO at runtime, so exploratory loops must be retained (read-back verification after each Bash step); what was cut was the useless overhead in the loop — Monolith goes through the Python client (not the MCP schema, which would eat 46K), default_timeout=120 (a 15-second timeout false-fails when the editor is busy), a hard 2-4 cap per step to prevent repeated retries, output ≤ 200 lines. Three strategies compared: uncapped loop 71 minutes (25% repeated operations), single script falsely fast, second iteration 10-15 minutes with adaptivity preserved.
UE Memory CDO Iron Law (distilled after handling the AM_CL_Attack_01 asset-pollution incident): any run_python that may change assets must first back up the .uasset; load_asset() returns the in-memory CDO, so it must be cloned before modification; reimport / refresh do not overwrite the loaded dirty CDO; total failure → report "editor restart required" and stop the task, writing the bak path into the reply.
A 3-day silent fault triggered by a storage migration
Symptom: from 2026-10-02 all UE sub-agents vanished — log evidence: ue-local-runner dispatches were 42 times on 10-01, 0 from 10-02; session titles became "You are an execution agent (ue-local-runner degraded)", i.e. the caller noticed the agent type was missing and hand-wrote a prompt to substitute, unnoticed for 3 straight days (until the user asked "why can't I see the sub-agents").
Root cause: user-level sub-agents are read from <ZCODE_STORAGE_DIR>/agents. The 10-01 storage migration moved ZCODE_STORAGE_DIR to G:\_ZCodeData but did not move agents/ — ZCode looked for agents in the new storage root, which was empty. (Skills are HOME-based so unaffected, creating the confusion "skills present, agents gone".)
Fix: copied agents/ (7) and commands/ (10, which follow the storage root the same way) to the new storage root G:\_ZCodeData; client restart takes effect.
Lesson: when migrating ZCode data, agents/ and commands/ must move with ZCODE_STORAGE_DIR; when an agent type is unavailable, silent degradation is forbidden — state it in the reply.
Data-driven agent tuning (last week's UE sub-agent check-up: 466 sessions / 20,868 tool calls / error rate 1.3%):
| Illness found | Prescription (written into the agent definitions) |
|---|---|
| One step per script (longest 228 minutes / 276 Bash; another case 79 minutes writing 104 scripts) | ue-local-runner gains "batch first": one target = one script; same target ≤3 attempts (trying new styles also counts) |
| Guessing file paths (31% of 264 errors were "File does not exist") | Glob/ls before read/write; two consecutive "not found" stops and reports |
Script library unused (audit_assets.py has 0 references in three executor agents) | Find assets / baselines / cleanup prefer the _AgentTools/ script library |
| Silent degradation (agents failed 3 days without a report) | New item in project AGENTS.md: a dispatch failure may not silently degrade — state it in the reply |
Also note: a bare : in the agent frontmatter description fails YAML parsing — ue-local-runner.md / project-rule-guard.md were changed to the folded block scalar description: >-. Do not use : in descriptions when writing agent/skill definitions.
⑦ Four Daily-Use Disciplines
Normally
Do nothing
Auto-loads 20 seconds after boot; returns VRAM to UE automatically after 3 minutes unused; 7 seconds to wake on reuse.
When using
Start a new session, pick the local model up front
Never switch an old cloud session over — hundreds of thousands of tokens of history necessarily hits the compression deadlock. When unsure, run session_sizes.py.
Boundaries
Big web pages / big docs: not for local
Fetching one full HTML page eats 10K~20K tokens. Changing code, reading files, running tools is the home turf.
Signal
See "auto-compressed" and prepare a new session
Compression summaries accumulate permanently; the only cleanup is a new session with the conclusions pasted in.
Common commands:
D:\_Qwen3.8-27b\start_local_llm.bat REM manual launch D:\_Qwen3.8-27b\stop_local_llm.bat REM release VRAM immediately D:\_Qwen3.8-27b\status_local_llm.bat REM one-click health check (run this first on trouble)