① Provider Configuration ② Number Pairing ③ 65% Compression Line ④ Change List ⑤ Boot Chain ⑥ Sub-Agents ⑦ Daily Use

① Provider Configuration (Copy-Paste Ready)

File .zcode\v2\provider_config.json. Values are the current effective version; evolution on ②③.

// Provider body
{
  "providerId": "88ff30fa-af55-4443-a1ff-ba8a7ff681c6",
  "providerName": "Local Ollama",
  "config": {
    "group": "standard-personal",
    "access": { "type": "api-key", "apiKey": "ollama-local-no-key" },
    "api": { "type": "openai-chat-completions",
             "baseUrl": "http://127.0.0.1:11435/v1" },
    "personalModelIds": ["qwen3.8:27b-offload"],
    "modelOrder": ["qwen3.8:27b-offload"]
  }
}

// Model rules (capability declarations must match the model's real capabilities)
{
  "modelId": "qwen3.8:27b-offload",
  "config": { "properties": {
    "contextWindow": 256000,
    "requiresMfjsToolSchema": false,
    "supportsToolCall": true,
    "supportsJsonSchemaOutput": false,
    "supportsMidConversationSystem": false,
    "inputFormat":  { "supportsText": true, "supportsImage": true },
    "outputFormat": { "supportsText": true }
  }}
}

Wrong capability-declaration fields create faults directly: on 2026-09-20 supportsImage was declared true while the model had no vision projector → image requests 500 in bulk → "reconnecting N/10" retry storm. After enabling mmproj (9/29) it was truly true. The contextWindow change 190000 → 256000 landed in two provider_config.json copies (G/E), and provider config changes require a ZCode restart — it reads it only at startup.

② Two Numbers Must Be Maintained in Pair

Server -c and client contextWindow — adjusting one side only guarantees the fault recurs.

contextWindow × 1.26 (measured worst-case underestimate factor) + 2,563 (compression instruction overhead) ≤ server -c
Server -ccontextWindowActual usage (UE open)Judgment
6553649152~18,900 MBConservative, but ZCode tends to overshoot
8192065536~19,200 MBMeasured overshoot requests did occur (82,324 > 81,920)
11468881920~20,608 MBStable operation achieved
262144190000~20.5 GiBSuperseded by the 327680 tier (10-04 evening)
327680256000~22.5 GB✅ Current (yarn extrapolation 1.25×, headroom 2,557)

Where the 1.26 coefficient comes from: three measurements in the same environment, ZCode's token estimate underestimating 5.5% → 15% → 26% (unstable bias). Compression requests also add ~2,563 token instructions on top of the history. So the check formula must take the worst observed value and then keep a margin — do not configure "just enough". Counterexample: 98,304 × 1.26 + 2,563 ≈ 126,426 > 114,688, so the old-era trigger point cannot be raised to 98,304.

③ Why the Compression Line Is 65%, Not 85%

ZCode's auto compression line is not "a percentage of the window", it is "the window first fixed-deducts 34,000 tokens".

Trigger line = contextWindow − 21,000 (output reserve) − 13,000 (safety margin)
contextWindowActual trigger lineTrigger percentage
49,15215,15230.8%
65,53632,53648.1%
98,304 → 190,000 reuses the same formula64,304 (at 98,304)65.4%
256,000222,00086.7%
1,000,000 (cloud)966,00096.6%

Both reserves are hard-coded constants — 185 compression events, across 6 window tiers, including the cloud 1M model, the values never changed. So the "compress at 85%" intuition holds only at large windows. Two practical conclusions:

  1. The only effective way to gain usable space is to enlarge the window: 81,920 → 98,304 moved the trigger line from 47,920 → 64,304, usable tokens +34%;
  2. The denominator of the panel's "used X/limit (P%)" is the window, not the trigger line — seeing compression at 70% means actual tokens have already crossed that line.

Reproduction command: python D:\_Qwen3.8-27b\_zcode_setup\_compact_probe.py (parses ZCode's jsonl logs).

④ Change List (new / modified / backup)

The full picture of files created and modified during integration — in troubleshooting, first think "which file does this feature live in".

New files (core items):

FileRole
start_local_llm.batThe single launch entry: kill leftovers first and confirm the port is truly free, then launch; wait for health; print VRAM
stop_local_llm.batRelease VRAM immediately. Kills only llama-server, not random python.exe
status_local_llm.batOne-click health check (process/listeners/health/VRAM/shared memory/logs/request records)
autostart_local_llm.bat + .vbsIdempotent boot starter + hidden window wrapper
_zcode_setup\qwen35_chat_template.jinjaPatched chat template (fixes the two 500 root causes)
_zcode_setup\is_ready.ps1 / wait_ready.ps1Silent health probe / wait-for-load (replacing unreliable quote escaping in batch)
_zcode_setup\last_usage.pyReads ZCode's sqlite, prints in/out/ttft of recent requests — fastest forensics of "what did ZCode actually send"
_zcode_setup\session_sizes.pyLists each session's real max prompt, judges whether a local model can run it
_zcode_setup\api_test.py / zcode_sim.py6 API self-checks / simulate ZCode request shapes (96 tools + 18K prompt)

Modified files (each maps to a true cause):

FileChangeReason
provider_config.jsonbaseUrl 11436 → 11435/v1; contextWindow 1000000 → real value (now 190000); capability declarations completedDirect connect; fake ctx never compresses; wrong declarations = 500 storm
openai_proxy.pyDelete max_tokens injection / delete HARD_MSG_CAP second trim / fix _trim_messages wiping all / add lock + cooling256 output cap injection; whole messages dropped; empty conversation; concurrent livelock
Startup folder .lnkPoint to running the .vbs via wscriptBoot hidden window + idempotent start
ZCode config.jsonDisable cloudbase-skills plugin, remove chrome-devtools / rider MCPTool schemas are a window killer (noise 38K → 46K)

All changes have timestamped backups (*.bak_date_time); rollback is copy-overwrite.

⑤ Boot Chain (Two, Independent of Each Other)

Design goal: model auto-loads 20 seconds after boot; each path is idempotent (probe /health first), and simultaneous triggers do not interrupt a loading model.

Main path (Startup folder):
NVFP4_llama-server.lnk
  → wscript.exe autostart_local_llm.vbs        (fully hidden window)
    → autostart_local_llm.bat                  (wait 20s for GPU driver)
      → is_ready.ps1 probes /health
         ├─ already 200 → log one line, do nothing (idempotent)
         └─ not serving → start_local_llm.bat to launch
      → also raises CherryStudio's proxy (if 11436 is free)
      → raises watchdog (60s cycle; restarts if the service disappears)

Redundant path (Registry Run key LocalLLM_ensure):
powershell.exe -WindowStyle Hidden -File ensure_model.ps1
  → First self-heal: convert LF line endings of .bat/.cmd/.vbs/.ps1 to CRLF in place
  → Then call the same autostart_local_llm.bat

Why the redundant path: the main path once failed silently for two days, with no log trace at all (the LF line-ending accident). The redundant path is written in PowerShell (unaffected by that trap) and self-heals line endings every boot. The Task Scheduler version is stronger but was rejected — Register-ScheduledTask reports 0x80070005 access denied in an unprivileged environment.

Troubleshooting order for "not usable after boot": status_local_llm.bat → by [1] Processes (0 = model is gone) → [9] Watchdog (running = wait 1 minute for auto-raise; NOT running = manual cscript //nologo //B watchdog.vbs) → [8] Logon autostart (whether it ran at boot) → [5] log tail (should have listening on).

⑥ Six UE Sub-Agents

The motive is saving tokens: the full tool schema eats ~46K, while each sub-agent carries only 6 basic tools (Read/Write/Edit/Glob/Grep/Bash), dropping the schema to ~7K and freeing 39K for the task itself.

#FilePositioningLoop cap
1ue-local-runner.mdPrimary executor: run UE tasks for the local model2
2ue-anim-debugger.mdAnimation tracing: evidence chain for "why did it play X"3
3ue-pcg-builder.mdPCG scene generation: node chains + empty-generation troubleshooting4
4ue-asset-explorer.mdRead-only investigator: Blueprint variables / material params / CDONone (read-only)
5ue-build-runner.mdBuild + verify: Live Coding / UBT fullNone (fixed flow)
6project-rule-guard.mdCompliance review: read rule files to verify changesNone

Loop design philosophy (a second-iteration conclusion): the true state of UE assets can only be known by reading the CDO at runtime, so exploratory loops must be retained (read-back verification after each Bash step); what was cut was the useless overhead in the loop — Monolith goes through the Python client (not the MCP schema, which would eat 46K), default_timeout=120 (a 15-second timeout false-fails when the editor is busy), a hard 2-4 cap per step to prevent repeated retries, output ≤ 200 lines. Three strategies compared: uncapped loop 71 minutes (25% repeated operations), single script falsely fast, second iteration 10-15 minutes with adaptivity preserved.

UE Memory CDO Iron Law (distilled after handling the AM_CL_Attack_01 asset-pollution incident): any run_python that may change assets must first back up the .uasset; load_asset() returns the in-memory CDO, so it must be cloned before modification; reimport / refresh do not overwrite the loaded dirty CDO; total failure → report "editor restart required" and stop the task, writing the bak path into the reply.

A 3-day silent fault triggered by a storage migration

Symptom: from 2026-10-02 all UE sub-agents vanished — log evidence: ue-local-runner dispatches were 42 times on 10-01, 0 from 10-02; session titles became "You are an execution agent (ue-local-runner degraded)", i.e. the caller noticed the agent type was missing and hand-wrote a prompt to substitute, unnoticed for 3 straight days (until the user asked "why can't I see the sub-agents").

Root cause: user-level sub-agents are read from <ZCODE_STORAGE_DIR>/agents. The 10-01 storage migration moved ZCODE_STORAGE_DIR to G:\_ZCodeData but did not move agents/ — ZCode looked for agents in the new storage root, which was empty. (Skills are HOME-based so unaffected, creating the confusion "skills present, agents gone".)

Fix: copied agents/ (7) and commands/ (10, which follow the storage root the same way) to the new storage root G:\_ZCodeData; client restart takes effect.

Lesson: when migrating ZCode data, agents/ and commands/ must move with ZCODE_STORAGE_DIR; when an agent type is unavailable, silent degradation is forbidden — state it in the reply.

Data-driven agent tuning (last week's UE sub-agent check-up: 466 sessions / 20,868 tool calls / error rate 1.3%):

Illness foundPrescription (written into the agent definitions)
One step per script (longest 228 minutes / 276 Bash; another case 79 minutes writing 104 scripts)ue-local-runner gains "batch first": one target = one script; same target ≤3 attempts (trying new styles also counts)
Guessing file paths (31% of 264 errors were "File does not exist")Glob/ls before read/write; two consecutive "not found" stops and reports
Script library unused (audit_assets.py has 0 references in three executor agents)Find assets / baselines / cleanup prefer the _AgentTools/ script library
Silent degradation (agents failed 3 days without a report)New item in project AGENTS.md: a dispatch failure may not silently degrade — state it in the reply

Also note: a bare : in the agent frontmatter description fails YAML parsing — ue-local-runner.md / project-rule-guard.md were changed to the folded block scalar description: >-. Do not use : in descriptions when writing agent/skill definitions.

⑦ Four Daily-Use Disciplines

Normally

Do nothing

Auto-loads 20 seconds after boot; returns VRAM to UE automatically after 3 minutes unused; 7 seconds to wake on reuse.

When using

Start a new session, pick the local model up front

Never switch an old cloud session over — hundreds of thousands of tokens of history necessarily hits the compression deadlock. When unsure, run session_sizes.py.

Boundaries

Big web pages / big docs: not for local

Fetching one full HTML page eats 10K~20K tokens. Changing code, reading files, running tools is the home turf.

Signal

See "auto-compressed" and prepare a new session

Compression summaries accumulate permanently; the only cleanup is a new session with the conclusions pasted in.

Common commands:

D:\_Qwen3.8-27b\start_local_llm.bat     REM manual launch
D:\_Qwen3.8-27b\stop_local_llm.bat      REM release VRAM immediately
D:\_Qwen3.8-27b\status_local_llm.bat    REM one-click health check (run this first on trouble)