Windows cmd.exe always exports PROMPT (its prompt template, default "$P$G")
into a child's environment. The engine's `getenv("PROMPT")` picked that up, so
running `glm.exe` from cmd for the oracle self-test instead entered text-
generation mode — loading a tokenizer the tiny model doesn't ship and failing
with "tokenizer.json: No such file". PowerShell has no PROMPT env var, so it
worked there (galmok's 32/32) but not in cmd — same command, different shell.
New coli_user_prompt(): honors COLI_PROMPT everywhere, and on Windows ignores a
PROMPT that carries cmd's $-metacodes ($P,$G,...) — a real prompt has none. cmd
users can still pass a prompt via COLI_PROMPT or a non-$ PROMPT. Verified: $P$G
-> oracle mode, "Explain recursion" -> honored, COLI_PROMPT -> honored.
Third native-Windows bug surfaced by a from-scratch main download (after the
-lpsapi link fix and the bilingual oracle message).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Running a real model with no PROMPT lands in oracle self-test mode, which
compares against ref_glm.json — the TINY model's oracle. The guard that
detects the vocab mismatch and points the user at PROMPT=/coli chat was
Italian-only, which read as a crash to English users (#271, galmok). Lead
with English, keep an IT footer, and add the `coli chat` hint.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The grouped-int4 (fmt=4) loader hard-coded gs=128 at all three detection
sites (resident weights + the two expert-load paths: mmap and pread-slab).
A checkpoint converted with --group-size 64 (or any non-128 size) wrote a
valid file that the engine then misdetected as plain per-row int4 (fmt=2)
and read the wrong number of scales -> garbage output, silently.
The conversion path (convert_fp8_to_int4.py --group-size) already emits
arbitrary group sizes, and the compute kernel (matmul_i4_grouped) is
fully gs-generic (only constraint: gs multiple of 16, the AVX2 width).
The loader was the sole gap.
Replace the three hardcoded checks with one shared helper, detect_group_size(),
that derives gs from the scale-array byte count by probing candidate sizes
{16,32,48,64,96,128,192,256} finest-first. Data-driven: any listed size
just works; per-row int4 (ns == O*4) correctly returns gs=0 -> stays fmt=2.
Verified against real GLM-5.2 expert dims (gate/up O=2048,I=6144 and down
O=6144,I=2048): g64/g128/g256 all detect correctly, and per-row is never
misdetected as grouped. g128 (the only previously-supported size) is
unchanged -> no regression for existing checkpoints.
This unblocks the g64 lane that ZacharyZcR's #225 ablation identified as
beating shipped per-row int4 (-7.5pp vs -9.3pp) at ~19% fewer bits: the
converter can produce it, and now the engine can load it.
Refs #225
The HWINFO snapshot read /proc/cpuinfo, /proc/meminfo and
sysconf(_SC_NPROCESSORS_ONLN) — all absent on native Windows — so the
dashboard runtime panel showed "0 GB RAM / 0 cores" while the CUDA branch
populated VRAM fine. Now: CPUID brand string (0x80000002..4, no new
deps), GetSystemInfo for logical cores, and the existing compat_meminfo
(GlobalMemoryStatusEx) for RAM. Verified live: panel reads "AMD Ryzen 9
9950X3D 16-Core Processor / 66 GB RAM 24 GB free / 32 cores".
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
vmmlaq_s32 computes a 2x2 int32 tile (2 weight rows x 2 activation
rows) per instruction on 8-deep segments. Tile o and s in pairs,
halving weight traffic and doubling per-instruction work at S>=2.
Four independent accumulators over a 64-deep unroll keep the loop
throughput-bound (a single chained accumulator measures no better
than SDOT: latency-bound). S=1 and all tails (odd o, odd s, I not a
multiple of 16/32) keep the existing SDOT/scalar code, and scales
apply in the same order, so results are bit-identical.
Compile-time gated on __ARM_FEATURE_MATMUL_INT8. The default Darwin
build passes no -mcpu and is byte-identical (still SDOT, IDOT_KERNEL
"neon"). Opt in with ARCH=native (new Darwin Makefile knob, appends
-mcpu=<arch>), which reports IDOT_KERNEL "neon-i8mm". The same gate
lights up on any aarch64 with i8mm (Graviton3+, Grace).
test_idot grows a driver-level exactness check through matmul_qt_ex:
fmt 1 and 2, S in {2,3,4,5,8}, O in {1,2,3,64,65}, I in {16,17,100,
1408}, bitwise float equality against a plain-C reference. Green on
both build flavors.
Measured on an M5 Pro (18 threads, matmul_qt_ex microbenchmark at
GLM-5.2 expert shapes, best of 3 process runs, vs the SDOT baseline):
gateup int4 S=8 499.7 -> 1076.6 GF/s (+115%)
gateup int8 S=8 512.9 -> 1166.3 GF/s (+127%)
down int4 S=8 696.9 -> 1186.0 GF/s (+70%)
S=1 decode rows unchanged (SDOT path untouched)
The decode/prefill PROFILE line splits expert-disk time into service
(overlapped async dispatch) and wait (blocking stalls), but every
accumulation site wrote t_edisk and nothing ever wrote t_ewait, so the
wait column always printed 0.000s. Worse, the accounted sum used the
dead t_ewait instead of t_edisk, so the entire disk-read stall was
excluded from accounted and silently fell into the other bucket.
On a disk-streaming MoE the effect is large: other reads as ~60% of
decode when it is really the expert-load stall. Route the three
blocking sites (non-PIPE parallel load, the Metal drain barrier, and
the per-expert pipe_wait in the CPU matmul loop) to t_ewait, keep the
async dispatch in t_edisk, and include both in accounted.
Profiling-only; no behavioural change. Before, on a 168-expert model:
expert-disk 25.077s service / 0.000s wait | ... | other 28.401s
After:
expert-disk 0.111s service / 26.948s wait | ... | other 3.390s
other now holds only the genuinely-unbucketed work (router, norms).
In serve/chat mode, Ctrl-C during generation killed the whole engine, losing
the loaded model and forcing a full reload. Now a SIGINT handler (armed only
in run_serve / run_serve_mux) sets a flag that spec_decode's token loop treats
exactly like hitting the NGEN cap: the turn ends through the normal path, so
the END sentinel, STAT line, usage_save and KV append all run — and :more can
continue the interrupted answer. The mux loop closes in-flight requests via the
same mux_done path. One-shot ./glm runs and Windows keep default SIGINT (die).
coli: stream_turn survives the first KeyboardInterrupt, forwards SIGINT to the
engine (covers non-TTY), drains to the turn boundary, and reports it. A second
Ctrl-C quits. Help line and per-turn footer updated.
POSIX only (sigaction); no behaviour change on Windows. Verified end-to-end on
Apple M4 + Metal: interrupt mid-decode, engine stays up, next prompt answers,
:q exits 0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
dot_i8i8 and dot_i4i8 accumulated the whole SDOT reduction into a single
int32x4_t. SDOT has ~3-4 cycle latency, so the serial dependency on `acc`
capped each core at ~26 GB/s (int8) / ~12 GB/s (int4) of weight throughput
regardless of memory bandwidth. Split into 4 independent accumulators (64
values/iter) so the loads become the bottleneck instead of the reduction
chain; the original single-acc loop is kept as the tail handler.
Measured on an Apple M4 (isolated microbench, expert-shaped 2048x6144):
int8*int8 26.0 -> 63.2 GB/s/core (2.4x)
int4*int8 12.4 -> 29.9 GB/s/core (2.4x)
Output is bit-identical to the previous kernels (verified over random
inputs). Non-DOTPROD NEON, AVX2/AVX-512/VNNI and VSX paths are untouched;
only the __ARM_FEATURE_DOTPROD branch changed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Cut Windows decode disk I/O from 2.06s/tok to 1.70s/tok (budget=4), meeting the
<=2s/tok target. Two changes on the pread expert-load path:
1. compat.h: replace the posix_fadvise no-op with a real WILLNEED cache-warmer
(overlapped ReadFile into a scratch buffer -> populates the standby page cache
so the later synchronous pread faults from RAM). Re-arms the existing
expert_prefetch/PILOT/next-block prefetch chain on Windows. DONTNEED stays a
no-op (matches macOS; Windows standby-list trimming self-regulates).
Measured: hit rate 16.4% -> 27.6%.
2. glm.c: flip PIPE (async expert-load thread pool) from default OFF to default ON
on Windows. Dispatches expert pread onto worker threads so loads overlap the
matmul, instead of blocking serial load-then-compute. PIPE=0 opts out.
Measured: expert-disk 65.9s -> 54.3s (-18%).
Also adds compat_fadvise assertions to tests/test_compat_direct.c (data integrity
after cache-warmer, safe no-op on bad fd / non-WILLNEED).
mmap (CreateFileMapping/MapViewOfFile) was implemented and tested at length but
reverted: it regressed on Windows (RSS bloat from touched mapped pages collapsed
the expert cache via ullAvailPhys — a fundamental Windows-vs-Linux difference).
Full findings + the dead-end analysis recorded in issue_diskio.md.
With METAL=1 + COLI_METAL=1, run_serve_mux (SERVE_BATCH=1) truncated
every completion to exactly 1 token: step_decode_batch passes per-row
kvs[]/positions[] with pos_base=0, but the two Metal decode fast paths
(attention_rows and the FULL-LAYER CB in layer_forward_rows) ignored
them and dispatched coli_metal_attn_decode/coli_metal_layer_decode with
the model-bound Lc/Rc and the hardcoded pos_base. The kernels' contract
is one sequence, row s at pos_base+s: ragged rows got roped at position
0 and attended over a T=1 window of the wrong cache, so greedy decode
hit a stop token on the first batched step (DONE ... STAT 1).
Gate both fast paths on !kvs. Ragged mux rows now take the CPU absorb
path, which already reads kvs[s]/positions[s]/ks->kv_start per row;
plain serve, chat/run, prefill and MTP verification (kvs==NULL) keep
the fused GPU kernels unchanged.
Verified on GLM-5.2 int4 (M5 Max): mux with COLI_METAL=1 went from
1-token DONE to full 16/16-token greedy completions, byte-identical to
the plain-serve comparator on both test prompts; CPU-only mux was
already correct (bisection); test-c and metal-test pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The original budget dropped experts blindly — even cached ones that cost
zero disk I/O. The miss-aware version pre-scans pin/ecache residency
before applying the budget:
- ALL cache hits are kept (free to compute, no disk I/O)
- Only misses compete for the remaining budget slots
- Miss budget = EXPERT_BUDGET - nhits (min 0)
Results (budget=4, same prompt/config as before):
Original Miss-aware
tok/s 0.33 0.36 (+9%)
hit rate 16.4% 38.2% (2.3x)
prefill 8.9s 5.6s (1.6x faster)
decode 97.4s 88.9s (9% faster)
The hit rate doubling is the key quality signal: the model now gets the
full contribution from all resident experts plus the top-4 new loads,
instead of losing some hits to the budget.
Add EXPERT_BUDGET env var that caps the number of distinct experts loaded
per layer across the batch-union. When the union exceeds the budget, keeps
only the highest-aggregate-gate-weight experts and drops the rest from
idxs[] so they're never loaded from disk.
Complementary to TOPP (per-position) — this trims the cross-position union
that multiplies under MTP/prefill. Based on MoE-Spec (arXiv 2602.16052):
'top 32 of 64 experts capture 93% of routing weight.'
Measurements (GLM-5.2 744B, 24GB RAM, cap=2, MTP=0, 32 tokens):
Baseline (budget=0): 0.18 tok/s, 9.3% hit, 176s decode, 39s prefill
EXPERT_BUDGET=12: 0.19 tok/s, 14.0% hit, 171s decode, 12s prefill
EXPERT_BUDGET=6: 0.26 tok/s, 21.0% hit, 123s decode, 7s prefill
EXPERT_BUDGET=4: 0.33 tok/s, 16.4% hit, 97s decode, 9s prefill
Budget=4 nearly doubles decode speed (+83%) and 4x's prefill speed by
halving disk reads per layer. Default OFF (EXPERT_BUDGET=0).
Two independent fixes validated end-to-end on fresh fixtures:
1. KV cache disk I/O (issue_diskio.md opportunities #1 + #4):
- kv_disk_append: fopen/fclose every turn -> persistent FILE* kept open
for the engine lifetime, lazy open on first append, closed in
serve_ctx_free. Eliminates per-turn handle creation overhead.
- kv_disk_append: ~157 small fwrites per position -> one contiguous
record memcpy'd into a staging buffer then a single fwrite per
position. The staging buffer grows on demand via realloc.
- kv_disk_truncate: closes the persistent handle before truncating
so the file actually shrinks on disc, then reopens lazily.
- KVState gains disk_fp, disk_buf, disk_buf_cap fields.
- Verified: serve-mode round-trip, write 11 tokens then reload and
resume with no re-prefill, then append 8 more and reload to 19.
2. Expert weight unfusing in test-model generators:
- The real GLM-5.2-FP8 checkpoint stores routed experts UNFUSED as
per-expert 2-D tensors, each with its own _scale_inv. HF fuses
gate+up into a single 3-D gate_up_proj for compute efficiency.
- The converter and C engine both expect the unfused layout. The
fused 3-D tensors were silently skipped by the converter, and the
engine crashed with missing-tensor errors.
- New unfuse_experts in glm_fp8_emit.py splits gate_up_proj and
down_proj into per-expert 2-D tensors. Called after reference
generation but before saving, in both generators, both FP8 and bf16.
- Also fixed: make_glm_oracle.py FP8 round-trip guard used p.dim()<2
which let 3-D fused experts through and crashed fp8_block_quantize.
Changed to p.dim()!=2 to match the converter ndim!=2 guard.
Validated full chain on fresh fixtures:
generator --fp8 -> 570 e4m3 tensors + 629 scale_inv, was 90 when fused
converter --group-size 0 -> per-row int4 fmt=2, engine loads clean
converter --group-size 128 -> grouped int4 fmt=4, 8-16x more scales,
engine loads clean, fmt=4 auto-detected in both mmap and slab paths
dequant error: grouped 1.14-1.22x lower than per-row vs FP8 source
Two independent fixes validated end-to-end on fresh fixtures:
1. KV cache disk I/O (issue_diskio.md opportunities #1 + #4):
- kv_disk_append: fopen/fclose every turn -> persistent FILE* kept open
for the engine lifetime, lazy open on first append, closed in
serve_ctx_free. Eliminates per-turn handle creation overhead.
- kv_disk_append: ~157 small fwrites per position -> one contiguous
record memcpy'd into a staging buffer then a single fwrite per
position. The staging buffer grows on demand via realloc.
- kv_disk_truncate: closes the persistent handle before truncating
so the file actually shrinks on disc, then reopens lazily.
- KVState gains disk_fp, disk_buf, disk_buf_cap fields.
- Verified: serve-mode round-trip, write 11 tokens then reload and
resume with no re-prefill, then append 8 more and reload to 19.
2. Expert weight unfusing in test-model generators:
- The real GLM-5.2-FP8 checkpoint stores routed experts UNFUSED as
per-expert 2-D tensors, each with its own _scale_inv. HF fuses
gate+up into a single 3-D gate_up_proj for compute efficiency.
- The converter and C engine both expect the unfused layout. The
fused 3-D tensors were silently skipped by the converter, and the
engine crashed with missing-tensor errors.
- New unfuse_experts in glm_fp8_emit.py splits gate_up_proj and
down_proj into per-expert 2-D tensors. Called after reference
generation but before saving, in both generators, both FP8 and bf16.
- Also fixed: make_glm_oracle.py FP8 round-trip guard used p.dim()<2
which let 3-D fused experts through and crashed fp8_block_quantize.
Changed to p.dim()!=2 to match the converter ndim!=2 guard.
Validated full chain on fresh fixtures:
generator --fp8 -> 570 e4m3 tensors + 629 scale_inv, was 90 when fused
converter --group-size 0 -> per-row int4 fmt=2, engine loads clean
converter --group-size 128 -> grouped int4 fmt=4, 8-16x more scales,
engine loads clean, fmt=4 auto-detected in both mmap and slab paths
dequant error: grouped 1.14-1.22x lower than per-row vs FP8 source
Root cause of gibberish output: the int4 quantization uses one F32 scale per
output row (2048 scales for a 2048x6144 matrix). The FP8 source has 128x128
block scales — 48x finer. This destroys reasoning while keeping surface fluency.
Changes:
- Converter: quant_int4_grouped() with --group-size 128 arg. Same nibble
packing, but one scale per group of 128 elements along the input dim.
- Engine QT struct: added 'gs' field (group size, 0=per-row backward compat)
- Engine qt_from_disk: auto-detects fmt=4 when scale array is O*ceil(I/128)
elements instead of O. Old per-row models (fmt=2) work unchanged.
- Engine matmul_i4_grouped(): AVX2 kernel that applies per-group scales.
Accumulator resets at each group boundary: dot(x[grp],w[grp]) * scale[grp].
- Engine matmul_qt_ex: dispatches to grouped kernel for fmt=4 (always exact,
no IDOT approximation since the point is quality)
- Engine expert_load: both mmap and slab+pread paths detect fmt=4 from
scale array size and set gs=128
- qt_bytes: fmt=4 reports correct memory including group scales
Backward compatible: existing per-row int4 models work unchanged.
The fused gate+up pair path (matmul_i4_pair) falls back to separate
matmul_qt calls for fmt=4 — minor perf cost, correctness preserved.
Two small correctness fixes surfaced by community reports on non-Linux hosts:
#236 — expert_load's buffered pread paths (slab + qs scales) used
perror("pread expert") on a short read. Since pread returned a short count
(not -1), errno stays 0 and perror prints "Success" — a confusing message
right before exit(1) in the score/bench path. New pread_full() helper loops
over short reads and EINTR and reports actual/expected bytes and offset, so a
truncated shard reads as such instead of "Success".
#219 — Linux pulls -pthread in via -fopenmp; the *BSDs do not, so pthread_*
fail to link there. Added -pthread to the generic (Linux/*BSD x86-64) and
aarch64 CFLAGS/LDFLAGS. No-op on Linux, required on FreeBSD (complements #206).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
PIN_GB=all passed gb=-1.0 to pin_load, which took npin=n (all experts) and
ignored --ram entirely — the OOM-kill regression from #80. A 92 GB host with
--ram 78 was killed mid-generation (anon-rss ~89 GB). Now gb<0 clamps npin to
expert_avail() — how many experts fit the RAM budget, same accounting AUTOPIN
uses. pin_load already adds the pinned bytes to resident_bytes, so the later
cap_for_ram narrows the LRU accordingly with no double count.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
kv_alloc had two consecutive if(k->Lc) free blocks. The first freed every
k->Lc[i]/k->Rc[i] and both arrays without unregistering from Metal and
without nulling k->Lc; the second then re-tested the dangling pointer,
called coli_metal_unregister on freed pointers, and freed everything a
second time. Safe only when k->Lc is NULL (first call); any re-allocation
on the same KVState aborts in the allocator.
The first block is the pre-Metal version of the free path: 3716e40 (Metal
backend) replaced it with the Metal-aware block, and ec89136 (GPU resident
pipeline) re-added it above during the merge. Delete it so the Metal-aware
block is the only free path.
Caught by tests/test_kv_alloc from the previous commit:
before: malloc: *** error for object 0x7: pointer being freed was not allocated (exit 134)
after: OK kv_alloc re-allocation (exit 0)
Paper-style cache-aware MoE selection (arXiv:2412.00099 max-rank):
keep true top-J always; fill remaining K slots preferring experts already
resident in pin∪LRU within top-M. Default OFF so stock full top-K is
unchanged.
Env: CACHE_ROUTE, ROUTE_J/M/P/ALPHA, ROUTE_AGREE (auto-on with CACHE_ROUTE).
Telemetry: swap%/route_swaps/route_slots, route_agree, route_kl on footer
and serve STAT. Complementary to PILOT (prefetch vs selection change).
Routing-only PR for clean A/B vs PILOT / #119; no CUDA/fuse stack.
See docs/CACHE_ROUTE.md. Closes nothing; for #161 discussion.
Co-authored-by: Vincent Marquez <vincentmarquez405@gmail.com>
Wires the two-step prediction (kind==2 from experiment/two-step-predict)
into the real pilot_prefetch() path behind PILOT_TWO=1 env var.
Changes:
- la_predict kind==2: computes shared expert (resident, no disk I/O) on
L's post_ln-normalized state, adds output to residual, then runs L+1's
router on the corrected state. Guards n_shared==0 and mloe_inter<=0.
- pilot_prefetch(): when PILOT_TWO=1, computes the same shared-expert
correction before running the router. Workspace allocated once per
call (not per position) to avoid malloc churn.
- LOOKA measurement harness expanded to 4 slots (prev, skip-attn,
PILOT stale, two-step) with updated reporting at both exit points.
- PILOT_TWO env var wired into main().
Measurements (GLM-5.2 744B, 24GB RAM, cap=2):
LOOKA recall: PILOT stale 73.6% -> two-step 76.7% (+3.1%)
End-to-end tok/s: no change (0.16 tok/s) — cache too small (cap=2)
for prediction quality to matter; disk bandwidth saturated regardless.
On higher-RAM hosts (cap>=32) the +3.1% recall would translate to
measurably fewer disk misses.
Prior art: 'Speculating Experts' (arXiv:2603.19289) independently
developed the same idea as a 'quasi-hidden state' using a static
default vector. Our approach uses the actual computed shared expert,
which is input-dependent and more accurate.
Co-authored-by: woolcoxm <13604288+woolcoxm@users.noreply.github.com>
* Fuse CUDA expert MLP execution
* Group CUDA expert transfers by device
* Instrument grouped CUDA expert execution
* Bound grouped CUDA decode scratch
* Execute expert groups across GPUs in parallel
* Release host backing for multi-GPU experts
* Define quality-preserving memory policies
* Overlap cold expert loading with resident compute
* Adapt expert placement with session LFRU
* Fuse q4 expert gate and up dispatch
* Plan CPU work on physical cores
* Batch grouped expert CUDA kernels
* Separate VRAM and RAM expert placement
* Add ragged multi-sequence decode forward
* feat(runtime): add continuous decode scheduler
* Route concurrent API requests through batch scheduler
* Harden multiplex request lifecycle and framing
* Cancel disconnected multiplex requests
* Bind API port before starting the engine
* fix automatic KV slot allocation
* add native int4 Tensor Core grouped GEMM
* add Tensor Core throughput benchmark
* optimize packed int4 low-row kernels
* add asynchronous CUDA staging streams
* document validated six-GPU dense acceleration
* tune six-GPU expert hot set
* raise validated expert hot-set target
* add CUDA MLA absorption core
* fuse grouped expert gate and up projections
* Warn for explicit lossy routing flags
* Add full-resident expert placement mode
* Adapt VRAM expert slots to live routes
* Accelerate int4 matvec on AVX-512
* Reduce AVX-512 and RoPE decode overhead
* Seed every GPU expert layer after prefill
* Limit live GPU swaps during decode
* CUDA batch MLA attention, kv_b head-sharding, fused o_proj, expert-group dispatch, W4A16 kernels
Lab-qualified on the 6x RTX 5090 machine (914-token request benchmark):
- batch MLA absorption kernel (COLI_CUDA_ATTN=1): whole-batch attention on
device, 154.8s -> 102.4s
- attention -> o_proj fusion on the layer device: -> 97.4s
- kv_b head-sharding across cards (COLI_CUDA_ATTN_SHARD=1), no weight
duplication: -> 94.05s
- per-device expert-group dispatch with pinned-buffer async transfers,
W4A16 tensor-core kernels for the shared expert, OMP hot-thread tuning
Negative results (reverted, kept out): GPU-side weighted scatter-add
(atomics + per-layer D2H lose 43.8%), shared-expert fused small-batch
kernel (-38.8%), W4A4 grouped tensor cores (int4 activations corrupt
output). Details in the lab research log.
* GPU resident pipeline: device-resident prefill attention chain, GPU expert groups in prefill, batched router, W4A16 mixed dispatch
COLI_CUDA_PIPE=1 keeps the prefill data plane on the layer home device;
control flow (routing, cache/pin management) stays on CPU. Any CUDA
failure falls back to the unchanged CPU path.
- Device primitives + unit tests (tests/test_pipe_cuda.cu): rmsnorm
(strided), interleaved RoPE, silu-mul, residual add, fixed-order row
merge (no atomics), device-input GEMM, persistent per-device scratch.
All verified against the engine's CPU math on SM120 (worst 1.2e-5).
- attn_pipe_prefill: q_a -> norm -> q_b -> rope -> kv_a -> norm -> rope ->
batch attention -> o_proj in one device chain (q_a/q_b/kv_a colocated
with kv_b); only the final [S,D] and the new KV rows return to host.
Attention 41.2s -> 30.8s on the 1571-token benchmark.
- Prefill batch-union now uses the GPU expert groups (previously gated to
S<=64, leaving all VRAM-resident experts idle during prefill - measured
21ms of GPU expert time in a 148s prefill). Expert phase 78.9s -> 69.0s.
- Router computed as one batched matmul instead of S sequential rows
(bit-identical math).
- W4A16 tensor-core path for expert groups (COLI_CUDA_TC_W4A16=1) with
row-count mixed dispatch: >=16 rows per expert use tensor cores, smaller
batches keep the naive kernel (tensor cores measured negative below
~16 rows). Expert phase 69.0s -> 64.3s, decode unaffected.
Net on the 1571-token prefill benchmark: 148.8s -> 114.3-126.8s
(component timings stable across runs; wall drifts +-3-5s because
.coli_usage placement learning shifts the expert tiers between runs).
PROFILO now also prints the prefill-phase breakdown.
* Skip OMP hot-thread tuning when CUDA is enabled
The active-spin worker team measured 66.9s->20.9s on the CPU-only Zen5
build, but on the six-GPU full-residency workload the spinning workers
contend with the CUDA dispatch threads: ~4x slower prefill with the
process stuck near 1.8 cores. Gate the tuning on COLI_CUDA so each
configuration keeps the behavior it was measured to prefer.
* Inc.2a: sparse layers fully resident on the layer device, residual hops cards at layer boundaries
COLI_CUDA_PIPE=2 keeps the residual stream on the layer home device for
consecutive sparse layers (cudaMemcpyPeer at boundaries): in/post norms,
attention chain, both residual adds and the shared-expert MLP run on
device. Per layer only the post-norm activations (router + CPU-tier
experts + group gather), the new KV rows and, on DSA indexer layers, the
pre-attention norm leave the card. Per-layer transfers drop from ~130MB
to ~70MB. A device-side snapshot at layer entry makes any mid-layer CUDA
failure fall back to the unchanged CPU path idempotently.
1571-token prefill: 127.1s (PIPE=1 control) -> 117.6/118.9s, components
attention 30.8->26.1, other 31.8->22.5-24.5; output verified coherent
against the control.
* Head-sharded attention inside the pipe: negative on PCIe star topology, gated opt-in
Slicing q per card from the home device and collecting ctx back
serializes ~95MB/layer through the home card's PCIe link: attention
26.1s -> 41.4/44.4s on the 1571-token benchmark (two repeats), wall
117.6 -> 135-138s. The standalone host-path sharding won because six
cards uploaded from host RAM in parallel; a home-device star has no
such parallelism without NVLink. Kept behind COLI_CUDA_PIPE_SHARD=1
for interconnects where peer bandwidth does not share one root port.
* Inc.3: device-resident KV shadow for decode attention
Decode re-uploaded the whole latent+rope window per layer per token
(~300MB/token at 1571 context). Each layer now keeps a device shadow of
the compressed KV on its kv_b card, bulk-synced when behind and appended
incrementally; the host cache stays canonical. Invalidation on kv_bind
(slot switch), kv_alloc (resize) and on any overwrite of mirrored rows,
with the legacy full-upload path as fallback.
Measured (COLI_CUDA_PIPE gate): short-context decode 5.48 -> 5.59/5.87
tok/s, 1571-context decode 4.14 -> 4.22 tok/s. Decode remains CPU-expert
bound; the shadow removes the transfer tax, not the compute.
* tools: unified user-experience benchmark (bench_ux.sh)
Two fixed scenarios (short chat, long-document QA), TTFT + decode tok/s
+ first-line drift check, TEMP=0 DRAFT=0 enforced, medians over REPS
runs. Encodes the measurement discipline from the lab record: same
binary per comparison, judge medians because .coli_usage placement
learning drifts wall times between runs.
* tools: bench_ux.sh executable bit
* gitignore compiled test binaries
* tools: expert_atlas.py — measure per-expert topic affinity (#175)
Diffs .coli_usage across 10 themed probe batches (code/math/chinese/
prose/science/law/poetry/structured/translation/casual, 3 prompts each)
driven through a running API server — one engine load total. Every
touched expert gets a topic-affinity vector, entropy, and a specialist/
generalist label; output experts.json feeds the Brain page hover.
* serve: persist .coli_usage after every turn in mux mode, not only at exit
run_serve_mux saved the learning cache once at shutdown; a crash lost
the whole session's routing history, and live consumers of the file
(expert_atlas.py diffs it between probe batches) saw a frozen snapshot.
Now saved per turn like the interactive path (165KB write, negligible).
* web: Brain hover shows measured expert atlas when published
If /experts.json (from tools/expert_atlas.py, #175) is served next to
the app, the tooltip upgrades from the depth heuristic to measured
data: specialist/generalist label, entropy, and the top-3 topic
affinities. Row index maps to real layer (row+3, last row = MTP 78).
Falls back to the heuristic when no atlas is published.
---------
Co-authored-by: JustVugg <JustVugg@users.noreply.github.com>
Two Windows-only bugs in run_serve_mux left the gateway hanging after READY:
1. No _setmode(_O_BINARY): the CRT collapsed CRLF inside fread() payloads (waits
forever for missing bytes) and expanded LF in the READY/STAT sentinels.
2. WaitForSingleObject on an anonymous pipe is undefined (always-signaled or
WAIT_FAILED) and PeekNamedPipe fails on file/console handles, so the dispatch
gate never opened. New rule: idle -> block in getline (POSIX select(NULL)
semantics); active -> PeekNamedPipe poll.
Reproduced and verified on real Windows via MinGW cross-compile + WSL interop:
old binary writes READY (with CRLF corruption) then hangs forever on a crafted
SUBMIT frame; fixed binary answers DONE + STAT and exits cleanly. Linux path
untouched (0 warnings, oracle-exact).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
#152 carried leftover debug scaffolding from the lm_head measurement posted on that
PR: a `g_lmhead_exact` global, an undocumented `LMHEAD_EXACT` getenv, and 7 lm_head
call sites routed through matmul_qt_ex(..., !g_lmhead_exact). None of it was in the
PR description; it rode in because `git checkout -B` carries uncommitted working-tree
edits forward.
It should not stay:
- it does nothing worth having. Pinning lm_head off IDOT measures -0.03% perplexity
across 5 corpora (in-distribution, isolated) — nothing.
- LMHEAD_EXACT=1 would silently change lm_head numerics, undocumented and untested.
- lm_head is fmt=1 (int8), so the IDOT gate applies to it unconditionally on every
platform anyway; there is no threshold story here to expose.
The flag defaulted to 0 and `!g_lmhead_exact == 1 == matmul_qt`'s own allow_idot, so
this removal is a pure no-op. Verified rather than assumed — summed log-lik over 1023
tokens, in-distribution, CPU (deterministic), dev-with-probe vs dev-minus-probe:
prose -2295.245383 == -2295.245383
markdown -3272.146403 == -3272.146403
Bit-identical. make check green (6 C suites + 60 python tests), no new warnings.
The actual #152 changes are untouched: the three attention input projections and the
DSA indexer's ix_wk stay batched and pinned to the exact int4 kernel via
matmul_qt_ex(..., 0).
SCORE mode scored raw token streams: GLM sees [gMASK]<sop> at the start of
every training sequence, so unprefixed requests run the model out-of-
distribution and silently distort logprobs (#108). The eval harness got the
text-level fix in #194; contributors driving SCORE directly (e.g. the
perplexity work in #153) were still exposed.
Same detection rule as tools/eval_glm.py: config.json model_type contains
"glm" (case-insensitive). The two ids are looked up in the snapshot's
tokenizer.json ([gMASK], <sop>) rather than hardcoded. Requests that already
carry the prefix pass through untouched — the patched eval_glm.py sends
prefixed streams, so no double-prefixing. SCORE_PREFIX=0 restores stock
behavior. One [SCORE] stderr notice when active.
Validated on GLM-5.2 744B (M4 Pro, Metal): default run flips the '2 + 2 ='
smoke question back to ' 4' (-1.006 vs ' 5' -2.950); SCORE_PREFIX=0
reproduces stock output; pre-prefixed requests give bit-identical logprobs
to auto-prefixed ones.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Splits the expert DISK LOADS (LRU misses -> expert_load) two ways in the
final stats line of 'run':
- by context: loads issued while drafting (mtp_draft) vs absorbing
verified tokens into the MTP KV (mtp_absorb) vs verify/main forwards
- by layer kind: the MTP layer (int8 experts, ~2x bytes) vs the 78 main
layers (int4), with exact bytes read (weights + scales)
Gated behind DISK_SPLIT=1 (default OFF, parsed once at startup like the
other g_* flags): when unset no atomic is ever touched and the stats
output is byte-identical to stock. Measurement only - no effect on
routing, verification or output.
Measured on M4 Pro 48GB (GLM-5.2 int4, METAL, --ram 38): the draft path
is ~1.5% of misses and the MTP layer ~4.5% of disk bytes.
Requested in #182.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* SCHEMA=<file.json>: JSON-Schema -> GBNF compiler for grammar-forced drafts (#48/#70 follow-up)
schema_gbnf.h compiles a practical JSON-Schema subset (strict objects, string/
number/integer/boolean/null, enum/const, arrays with items, nesting) into the
byte-level GBNF subset grammar.h parses, so structured-output workloads get
grammar-forced drafts without hand-writing GBNF. Unsupported keywords fail soft:
the engine runs without a grammar and output is unchanged (drafts are verified,
never constraints - a wrong compile can only cost acceptance, not correctness).
grammar_setup: GRAMMAR= (raw GBNF) keeps precedence; SCHEMA= feeds the compiler
into the same gr_parse path. 8 test groups in tests/test_schema_gbnf.c walk
compiled grammars end-to-end through the PDA (forced spans, enum disambiguation,
nested instances, escapes, leading-zero rejection, fail-closed fallbacks).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* schema_gbnf: whitespace-tolerant emission (jws at separators)
Measured on GLM-5.2 current main (#146): the greedy continuation writes sloppy
JSON (spaces after colons, fences, long free text) and a compact-only grammar
desyncs at the first stray space, forfeiting every span after it. jws points are
not forced themselves (two legal bytes) but the multi-byte spans around them
keep drafting and the walker survives non-compact output - strictly
acceptance-positive for a verified draft source. Tests re-derived for the new
span boundaries + a sloppy-instance walk.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: JustVugg <JustVugg@users.noreply.github.com>
* cuda-dll: fix Windows build — MSVC host flags, CUDA_PATH default, POSIX setenv shim in the kernel test (#157)
First hardware validation of the #131 CUDA_DLL path (RTX PRO 6000 Blackwell
sm_120, MSVC 14.44 + CUDA 13.2, MSYS2 UCRT64 host build) found three blockers
that made 'make cuda-dll' unbuildable as shipped:
- NVCCFLAGS passed GCC-style -Xcompiler=-Wall,-Wextra to the MSVC host
compiler (hard error D8021). Use -Xcompiler=-W3 on Windows — dash form,
since MSYS make mangles /W3 into a filesystem path.
- NVCC defaulted to $(CUDA_HOME)/bin/nvcc with CUDA_HOME=/usr/local/cuda;
on Windows default CUDA_HOME from the installer's CUDA_PATH and NVCC to
plain 'nvcc' from PATH (CUDA_PATH contains spaces, which the unquoted
recipe checks cannot survive; an MSVC PATH environment is already required).
- tests/test_backend_cuda.cu used POSIX setenv/unsetenv (undefined under
MSVC); add a two-line _putenv_s shim.
Also corrects the stale '11 API symbols' comment (the header exports 15 and
backend_loader.c resolves all 15).
Validated: make cuda-dll (stock flags) + make glm CUDA_DLL=1 ARCH=native →
[CUDA] device init on sm_120, tiny oracle TF 32/32 + greedy 20/20, kernel
suite 'q8/q4/q2/f32 correctness ok', graceful no-dll fallback, plain build
byte-identical CPU behavior.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* cuda: fix heap corruption in expert_host_release on Windows (CUDA_RELEASE_HOST=1)
expert_host_release() freed the expert slab with plain free(), but the slab
is posix_memalign'd — which compat.h maps to _aligned_malloc on Windows, so
free() corrupts the CRT heap: instant 0xC0000374 crash on the first released
expert. This is the exact pattern the compat.h audit fixed at the original
expert_load site ("l'unico sito che libera memoria aligned e' free(s->slab)");
this call site was added later and reintroduced it. compat_aligned_free is
plain free on POSIX, so non-Windows behavior is unchanged. fslab stays plain
free (malloc/falloc on the CPU path).
Found running the VRAM expert tier on real hardware (80 GB resident on an
RTX PRO 6000, CUDA_RELEASE_HOST=1 to avoid 80 GB of host double-residency —
reproducible crash before, clean generation after).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: lEWFkRAD <186512915+lEWFkRAD@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: JustVugg <JustVugg@users.noreply.github.com>
* prefill: batch the attention input projections (q_a/q_b/kv_a) over the prompt
During prefill the three attention input projections were called one row at a time
(matmul_qt(..., 1)) inside the loop over the S prompt tokens, while o_proj right
below already runs batched and moe() already does a batch-union. Batching them over
all S rows lets each weight row be read once for the whole prompt instead of once
per token.
Measured on Zen5 + RTX 5090, GLM-5.2 int4, 78 layers, DSA on, OMP tuning on,
prefill run to completion (engine's own timings, upstream/dev @ 6d3ed7e):
prompt projection/RoPE total prefill
256 22.7s -> 17.7s (-22%) 120.9s -> 115.4s (-4.5%)
512 45.5s -> 36.3s (-20%) 231.4s -> 222.7s (-3.8%)
Output is BIT-IDENTICAL to dev: summed log-lik of a fixed 1023-token passage is
-4987.458316 before and after (SCORE mode, CPU, deterministic). Decode (S=1) is
untouched — same calls, same kernel.
matmul_qt_ex(..., allow_idot=0) keeps the projections on the EXACT int4 kernel.
This matters: batching alone would push S past the S>=g_i4s gate (2 on x86) and
silently move them onto IDOT's int8-quantized activations, which costs
+0.169 nats/token on markdown/code and +0.084 on prose (perplexity +18.4% / +8.7%,
measured) for only 1.4% more speed now that #95 made the exact kernel fast. The
gate is a decode-era speed threshold; it should not be crossed by a batching change.
Also: matmul_qt documents the IDOT threshold as "configurable via I4S", but the
getenv was missing and the knob did nothing. Wire it up — it is what isolates the
kernel switch from the batching win.
* prefill: batch the DSA indexer's ix_wk projection too
Same per-token bug as the attention projections, one block below: ix_wk was called
one row at a time inside the prefill loop, so its weights were re-read for every
token on every DSA layer.
matmul_qt_ex(..., 0) keeps it on the exact int4 kernel for the same reason as the
attention projections: batching alone would push S past the S>=g_i4s gate and
silently change the activation quantization.
Output stays BIT-IDENTICAL (summed log-lik over 1023 tokens, in-distribution):
prose -2295.245383 (unchanged)
markdown -3272.146403 (unchanged)
Speed, 256-token prompt, 3 alternating reps each (Zen5 + RTX 5090, dev + the
attention-projection commit as the baseline):
projection/RoPE 17.54s (sd 0.20) -> 16.96s (sd 0.15) -3.3%
total prefill 115.34s (sd 0.82) -> 113.69s (sd 0.35) -1.4%
Both separated beyond 2 sd. Small, but it is free and bit-identical.
ix_wq/ix_wp are left alone on purpose: they sit inside an `omp parallel for` and
are skipped entirely below index_topk context, so batching them is a different
change with a different risk profile.
---------
Co-authored-by: JustVugg <JustVugg@users.noreply.github.com>