- Fix ENV VARS header: document PILOT=0-3, SMOOTH, CONF_LIMIT; remove stale REBAL entry
- Fix per-layer EMA: apply routing momentum to all layers (not just layer 0) with correct offset
- Fix in-flight slot race in expert_get: LRU eviction now skips slots with eid==-1 (being loaded)
- Fix in-flight slot race in pilot_realload: same fix, prevents concurrent writes into active slot
- Fix idx[] buffer overflow: clamp max_cand to 128 before E in pilot_prefetch
The dashboard's wall-time stack assumed ewait = felt I/O stall and
edisk = overlapped read service, but the engine measured neither:
t_ewait was declared and reported yet never incremented (the blue
'I/O wait' segment always read 0.00s), while t_edisk accumulated on
the COMPUTE thread — blocking OMP loads, PIPE dispatch, pipe_wait
spins — i.e. exactly the felt stall. The UI excluded it from the
stack as 'overlapped', so the whole disk stall landed in 'Other'
(e.g. 46.9s Other vs 45.3s disk on a 54.1s turn).
Now the semantics match the contract:
- t_ewait accumulates the compute-thread stall (blocking loads,
PIPE dispatch, pipe_wait spins) and feeds the in-stack I/O wait.
- Disk service is real: expert_load is timed on whichever thread
runs the read (PIPE workers, OMP loaders, pilot) into an atomic
ns counter (g_edisk_ns) — thread-seconds, overlapped with compute.
- prof_report's I/O share and verdict use the felt wait only, and
the expert I/O line prints 'read service / felt wait' so the
wait:service gap shows how well PIPE is hiding the reads.
PROF wire format, /profile JSON and the web UI are unchanged —
the dashboard's existing assumptions are simply true now.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dd4hHD2Dsvgr7VAwEVFqWK
The engine already tracks where each turn's wall time goes (expert-disk
service, I/O wait, expert matmul, attention, lm_head) — it just only spoke
at exit or under PROF=1. Stream it instead:
- glm.c: mux serve emits a per-turn "PROF" protocol line next to TIERS/HITS
(window deltas per request, same convention as the STAT hit%); the phase
window base is now always captured (a few loads per request).
- openai_server.py: parses PROF into a 120-turn rolling window and serves it
at /profile (read-only, same trust level as /health).
- web: new Profiling tab — stat tiles (tok/s, wall, tokens/forward, disk
service), wall-time composition bars for the last turn and the window,
per-turn throughput and stacked phase columns with hover readouts, and a
table of recent turns. Disk service is shown apart from the stack: it
overlaps with compute, so only the I/O wait the compute thread felt counts
inside wall time. Phase colours are a CVD-validated set with gaps + legend
+ table so identity never rides on colour alone.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WhTmF8yvEBgSkUKSVfZF7P
run_serve_mux (SERVE_BATCH, used by openai_server.py / coli web) completes
requests in mux_done, so the run_serve per-turn hook never fired there.
Snapshot the window where hits0 is taken, record batched-forward latency
around step_decode_batch, report on stderr at DONE. With KV_SLOTS>1 the
window shares batched forwards across slots — same convention as the
existing STAT hit%.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TCBNCciBaHea41QmidLUMn
Answers 'where does it slow down on THIS machine with THIS config' so users
can tune RAM_GB/PIPE/DIRECT/PIN for their hardware without folklore:
- startup header: CPU/cores/RAM/backend + effective knobs (cache cap, pin,
DRAFT/PIPE/DIRECT/MMAP/IDOT/DSA/PILOT/CACHE_ROUTE) — every saved log is
self-describing when comparing runs across configs or machines
- per-forward decode latency ring (32k) -> p50/p90/p99/max, plus a tail
diagnosis when p99 >> p50 (cold-cache expert loads)
- expert I/O accounting at the pread/mmap-touch sites: GB fetched, MB/token,
GB/s, hit rate, loads/token, pinned/LRU tier fill
- phase shares of wall time and a plain-language verdict naming the knob
most likely to move tok/s (I/O-bound vs compute-bound vs attention-bound)
- reports after REPLAY / PROMPT / oracle runs (stdout) and per turn in serve
mode (stderr; stdout stays the framed protocol)
Additive only: with PROF unset every mode's output is byte-identical.
hwinfo_emit's /proc probe is factored into hw_probe() and shared.
Validated end-to-end on a tiny-random unquantized fixture (REPLAY, PROF on/
off, RAM_GB squeeze flips hit 96.9%->26.6% and the verdict follows);
make check clean, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TCBNCciBaHea41QmidLUMn
Both sentinels in c/coli (READY and END) end with \n. Under CRT text
mode on Windows, printf() translates \n to \r\n, so the engine emits
\x01\x01READY\x01\x01\r\n. The Python coli wrapper checks
endswith(SENTINEL) which expects bare \n — the match never fires and
chat hangs forever.
_setmode(fileno(stdout), O_BINARY) at engine startup switches stdout
to binary mode so \n passes through unchanged. This matches the
belt-and-braces reasoning already in compat.h:88 (O_BINARY for file
I/O). The fix is defense-in-depth: on the documented w64devkit/MinGW
build path, CRT text mode may or may not apply depending on how the
terminal is attached, so this ensures correctness regardless.
Note: I was unable to compile-test this on the full codebase because
glm.c uses POSIX-only symbols (mmap, madvise, select, fd_set, struct
stat/fstat) that are unavailable on Windows without platform guards.
The codebase compiles on Linux (Ubuntu 24.04, GCC 13) and macOS
(Apple clang). Windows compilation requires wrapping those POSIX calls
in #if defined(__APPLE__) || defined(__linux__) || defined(__FreeBSD__)
guards — I opened this as a separate concern for the maintainer.
@bokiko benchmarked the feature on three hosts and found every setting is
either no faster or no longer coherent. Reproduced here on a 25 GB box, and
the numbers are worse than "a quality/speed tradeoff":
- budget=8 -> hellaswag 30% vs 90% with it off (25% is the chance floor)
- budget=4 -> decode is literal noise ("The **1...: s2151:")
- MTP acceptance 0%. Which experts survive the cap depends on cache
residency at the moment of the forward, so draft and verify do not
compute the same function -- the invariant #294 just established for
#163, violated through cache state instead of kernel dispatch. SPEC_PIN
cannot fix that and should not try: it is feature semantics, not
dispatch.
- 0.13 tok/s vs 0.30 baseline, while loading 14.66 experts per layer
against a topk=8 baseline. The cap does MORE disk I/O than not using it.
"~335 GB I/O saved" counts dropped experts, not bytes not read.
EXPERT_BUDGET>0 is now ignored unless EXPERT_BUDGET_EXPERIMENTAL=1, with the
measurements printed. The code stays compiled and developable: MoE-Spec
(arXiv 2602.16052) is not a wrong idea, this implementation just has no
point where it is both faster and correct. Re-enabling it by default needs a
quality number next to every speed number.
Also gates the cap to decode (S<=4), @woolcoxm's fix from #292/#298: during
prefill the batch union is 30-100+ experts, and capping to 4-8 drops most of
them, corrupting the hidden state and therefore the KV cache. Necessary but
not sufficient -- the run above already includes it.
Removes issue_budget.md, issue_diskio.md and issue_grouped_quant.md: design
notes in the repo root, unlinked from any README, whose only user-facing
content was "EXPERT_BUDGET=6-8 -- good speedup, minimal quality loss" and
"+83% decode at budget=4". Measurement says otherwise, and they were the
only thing on main telling anyone to switch this on. They stay in history.
Reported-by: bokiko <#303>
Co-Authored-By: woolcoxm <#292>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two silent failures that compound into "[engine terminated]" with no cause.
cap_for_ram() floors the expert cache at cap=1. When resident+slack have
already blown the budget, avail is negative and capmax would be 0 -- "I do
not fit in your budget". Flooring it to 1 and carrying on turns "I do not
fit" into "I overshoot", which is precisely the mid-generation OOM-kill the
function exists to prevent: it printed "projected peak 25.1 GB" against a
22 GB budget and started anyway. Now it says so, names PIN_GB when that is
what inflated the resident set, and refuses to start when the peak also
exceeds the memory actually available on the machine (COLI_RAM_OVERCOMMIT=1
overrides).
The kernel kills with SIGKILL: no error, no log, stdout just closes. coli
read that EOF and printed "[engine terminated]" without ever reaping the
child, so an OOM-kill was indistinguishable from a clean exit -- the report
in #305 (Debian 12, dies mid-generation, no message). engine_diag() now
reports the signal or exit code, names the OOM-killer when it was SIGKILL,
and shows the tail of the engine's stderr.
Reported-by: Ne00n <#305>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
One HF 504 killed the whole bench. Now: load_dataset retries with
exponential backoff (hf_hub resumes partial downloads from cache); a task
that still fails is skipped instead of killing the rest; JSONLs are written
atomically (coli only checks existence, so a truncated file from an
interrupted run would block re-download forever); coli bench drops
still-missing tasks with a warning and refuses to run eval with none.
Co-Authored-By: Claude <noreply@anthropic.com>
Apple clang 16 (clang-1600.0.26.6) defaults objective-c++ to a pre-C++11
dialect, so the raw string literal holding the Metal shader in
backend_metal.mm fails to parse:
backend_metal.mm:12:29: error: use of undeclared identifier 'R'
backend_metal.mm:13:10: fatal error: 'metal_stdlib' file not found
Pinning gnu++17 on METALXX fixes 'make glm METAL=1' and 'make metal-test'.
Verified on macOS 15 / M4 Max: both targets build and all metal backend
tests pass.
PIN=auto resolves to <model>/.coli_usage, the history that serve mode
appends after every turn, so each restart's hot-store placement follows the
accumulated REAL workload instead of a frozen one-shot profile; stats.txt is
the fallback for a virgin model dir, and with neither present the run simply
starts unpinned. Same magic-value convention as PIN_GB=all (#80); explicit
paths and the AUTOPIN flow are untouched (PIN set skips AUTOPIN as before).
This lifts deployment entrypoints' "prefer .coli_usage over stats.txt" shell
plumbing into the engine, plus the ENVIRONMENT.md row.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Select a portable architecture from the compiler target instead of forcing x86-64-v3 on every platform. On macOS, only enable Homebrew OpenMP when its header and library actually exist, preserving the dependency-free fallback.
MTP acceptance collapses when the draft (S=1) and verify (S>=2) forwards do
not compute the same function. Three switches were S-dependent: the int4
IDOT gate (S>=g_i4s, asymmetric exactly on ISAs where g_i4s>1), the
S==1-only fused gate+up pair, and the Metal GEMM row threshold.
SPEC_PIN=1 (default) pins every forward issued while model drafts are live
to the platform S=1 kernel family, so draft and verify agree by
construction. Prefill and non-speculative decode keep their gates.
SPEC_PIN=0 restores the old behavior for A/B.
The 24-proposal <10% guard now pauses drafting for 256 tokens and re-arms
on a fresh window instead of latching off for the whole session.
Co-Authored-By: Claude <noreply@anthropic.com>
Two issues found by the diagnostic sweep in #292:
1. Profile double-count in pipe_layer_sparse: the outer t_emm span wrapped
both the shared-expert GPU dispatch AND the moe() call, but moe()
self-times its own t_emm internally (6 accumulation sites). Routed-expert
matmul was counted twice -> accounted > elapsed -> 'other' went negative,
and expert-matmul showed an inflated ~120s. Fix: split into two narrow
spans around only the GPU work moe() does NOT cover (shared-expert
dispatch, routed upload+add), letting moe() self-time as all other
callers already do. Verified: 'other' now positive (14-17s),
expert-matmul realistic (5-6s).
2. MTP blanket-disabled under CUDA (g_draft=0 when g_cuda_enabled). This
was a conservative guard from #163 before the root cause (cold-expert
fused-pair + IDOT kernel divergence) was fully diagnosed. GPU-resident
experts have no divergence; the cold subset still does but achieves
30-50% acceptance anyway (matching non-CUDA builds). Add COLI_CUDA_MTP=1
opt-in so users can test speculation under CUDA. Default unchanged.
Verified: COLI_CUDA_MTP=1 -> MTP ACTIVE (draft=3), 44% acceptance,
3 forwards for 8 tokens (vs 7 without MTP).
Refs #292#163
pin_wire() only locked the legacy private-slab path (s->slab / s->fslab),
which stay NULL under COLI_MMAP -- so MLOCK=1 on an mmap build silently
wired 0 bytes and the 'pinned' set was just warm page cache, evictable
under memory pressure (measured: hundreds of MB/s of re-reads from disk
mid-generation on a 503 GB host once the cache tightened).
Three pieces:
- qt_wire_mmap(): mlock each pinned expert's weight + scale ranges inside
the file mappings. Skips cuda_eligible QTs: VRAM-tier experts compute
from device memory, and expert_host_release() early-returns for mmap
experts (no slab) WITHOUT nulling q8/q4, so a host-pointer check alone
wires the whole VRAM tier too (~137 GB of never-touched locked pages on
a 6x32GB-class rig -- enough to starve the kernel into thrashing).
- qt_unwire_mmap(): REPIN gpu_swap promotions drop the promoted expert's
host lock instead of leaking it on every swap.
- expert_load() deliberately does NOT wire (it also runs for the transient
VRAM-staging pass); wiring happens once in pin_wire() on the final set.
Measured on GLM-5.2 int4, 4x RTX 5090 + 1x RTX 4090, 503 GB RAM,
PIN_GB=all: wired goes 0 -> 226 GB (exactly the RAM tier), decode-time
disk reads drop to zero once warm.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Windows cmd.exe always exports PROMPT (its prompt template, default "$P$G")
into a child's environment. The engine's `getenv("PROMPT")` picked that up, so
running `glm.exe` from cmd for the oracle self-test instead entered text-
generation mode — loading a tokenizer the tiny model doesn't ship and failing
with "tokenizer.json: No such file". PowerShell has no PROMPT env var, so it
worked there (galmok's 32/32) but not in cmd — same command, different shell.
New coli_user_prompt(): honors COLI_PROMPT everywhere, and on Windows ignores a
PROMPT that carries cmd's $-metacodes ($P,$G,...) — a real prompt has none. cmd
users can still pass a prompt via COLI_PROMPT or a non-$ PROMPT. Verified: $P$G
-> oracle mode, "Explain recursion" -> honored, COLI_PROMPT -> honored.
Third native-Windows bug surfaced by a from-scratch main download (after the
-lpsapi link fix and the bilingual oracle message).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Running a real model with no PROMPT lands in oracle self-test mode, which
compares against ref_glm.json — the TINY model's oracle. The guard that
detects the vocab mismatch and points the user at PROMPT=/coli chat was
Italian-only, which read as a crash to English users (#271, galmok). Lead
with English, keep an IT footer, and add the `coli chat` hint.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
compat.h's rss_gb() calls GetProcessMemoryInfo and links psapi via
#pragma comment(lib,"psapi.lib") — an MSVC-ism. MinGW gcc ignores that pragma
(emits -Wunknown-pragmas), and the Windows LDFLAGS never linked psapi, so on a
gcc that doesn't honor the pragma (e.g. 16.1.0 UCRT) the build fails with
`undefined reference to GetProcessMemoryInfo`. Add -lpsapi to the Windows
LDFLAGS; harmless on toolchains where the pragma also resolves it. Found while
building on native Windows 11 with winlibs GCC 16.1.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Opt-in (default off): prints the generated token ids to stderr after
generation, for exact comparison across decode paths (e.g. the S=1
resident pipeline vs the CPU path). Used to verify the pipe_layer_sparse
gate relaxation is token-exact vs the CPU decode path (see #273).
The grouped-int4 (fmt=4) loader hard-coded gs=128 at all three detection
sites (resident weights + the two expert-load paths: mmap and pread-slab).
A checkpoint converted with --group-size 64 (or any non-128 size) wrote a
valid file that the engine then misdetected as plain per-row int4 (fmt=2)
and read the wrong number of scales -> garbage output, silently.
The conversion path (convert_fp8_to_int4.py --group-size) already emits
arbitrary group sizes, and the compute kernel (matmul_i4_grouped) is
fully gs-generic (only constraint: gs multiple of 16, the AVX2 width).
The loader was the sole gap.
Replace the three hardcoded checks with one shared helper, detect_group_size(),
that derives gs from the scale-array byte count by probing candidate sizes
{16,32,48,64,96,128,192,256} finest-first. Data-driven: any listed size
just works; per-row int4 (ns == O*4) correctly returns gs=0 -> stays fmt=2.
Verified against real GLM-5.2 expert dims (gate/up O=2048,I=6144 and down
O=6144,I=2048): g64/g128/g256 all detect correctly, and per-row is never
misdetected as grouped. g128 (the only previously-supported size) is
unchanged -> no regression for existing checkpoints.
This unblocks the g64 lane that ZacharyZcR's #225 ablation identified as
beating shipped per-row int4 (-7.5pp vs -9.3pp) at ~19% fewer bits: the
converter can produce it, and now the engine can load it.
Refs #225
Per ZacharyZcR's 6x5090 A/B on #273: the S=1 resident-pipeline relaxation is
+49% on a single GPU (5070 Ti) but a wash on multi-GPU. With layers sharded
across N devices, each resident forward at S=1 crosses P2P per layer group and
those small hops don't amortize — the same term that killed pipe x head-shard
in #111. Multi-GPU decode walls on disk service, which pipe2 can't touch.
Make the threshold device-count-dependent:
- single GPU (g_cuda_ndev<=1): S>=1 (the breakthrough path)
- multi GPU (g_cuda_ndev> 1): S>=8 (the original prefill-only gate)
with COLI_CUDA_PIPE_S_MIN as an env override for anyone who wants to measure.
The two calibration points bracket the design space:
1x 5070 Ti, modest CPU, "other"-bound decode -> S=1 (+49%)
6x 5090, sharded, disk-service-bound -> S=1 a wash, keep S>=8
Refs #273 (comment)