The grouped-int4 (fmt=4) loader hard-coded gs=128 at all three detection
sites (resident weights + the two expert-load paths: mmap and pread-slab).
A checkpoint converted with --group-size 64 (or any non-128 size) wrote a
valid file that the engine then misdetected as plain per-row int4 (fmt=2)
and read the wrong number of scales -> garbage output, silently.
The conversion path (convert_fp8_to_int4.py --group-size) already emits
arbitrary group sizes, and the compute kernel (matmul_i4_grouped) is
fully gs-generic (only constraint: gs multiple of 16, the AVX2 width).
The loader was the sole gap.
Replace the three hardcoded checks with one shared helper, detect_group_size(),
that derives gs from the scale-array byte count by probing candidate sizes
{16,32,48,64,96,128,192,256} finest-first. Data-driven: any listed size
just works; per-row int4 (ns == O*4) correctly returns gs=0 -> stays fmt=2.
Verified against real GLM-5.2 expert dims (gate/up O=2048,I=6144 and down
O=6144,I=2048): g64/g128/g256 all detect correctly, and per-row is never
misdetected as grouped. g128 (the only previously-supported size) is
unchanged -> no regression for existing checkpoints.
This unblocks the g64 lane that ZacharyZcR's #225 ablation identified as
beating shipped per-row int4 (-7.5pp vs -9.3pp) at ~19% fewer bits: the
converter can produce it, and now the engine can load it.
Refs #225
vmmlaq_s32 computes a 2x2 int32 tile (2 weight rows x 2 activation
rows) per instruction on 8-deep segments. Tile o and s in pairs,
halving weight traffic and doubling per-instruction work at S>=2.
Four independent accumulators over a 64-deep unroll keep the loop
throughput-bound (a single chained accumulator measures no better
than SDOT: latency-bound). S=1 and all tails (odd o, odd s, I not a
multiple of 16/32) keep the existing SDOT/scalar code, and scales
apply in the same order, so results are bit-identical.
Compile-time gated on __ARM_FEATURE_MATMUL_INT8. The default Darwin
build passes no -mcpu and is byte-identical (still SDOT, IDOT_KERNEL
"neon"). Opt in with ARCH=native (new Darwin Makefile knob, appends
-mcpu=<arch>), which reports IDOT_KERNEL "neon-i8mm". The same gate
lights up on any aarch64 with i8mm (Graviton3+, Grace).
test_idot grows a driver-level exactness check through matmul_qt_ex:
fmt 1 and 2, S in {2,3,4,5,8}, O in {1,2,3,64,65}, I in {16,17,100,
1408}, bitwise float equality against a plain-C reference. Green on
both build flavors.
Measured on an M5 Pro (18 threads, matmul_qt_ex microbenchmark at
GLM-5.2 expert shapes, best of 3 process runs, vs the SDOT baseline):
gateup int4 S=8 499.7 -> 1076.6 GF/s (+115%)
gateup int8 S=8 512.9 -> 1166.3 GF/s (+127%)
down int4 S=8 696.9 -> 1186.0 GF/s (+70%)
S=1 decode rows unchanged (SDOT path untouched)
The decode/prefill PROFILE line splits expert-disk time into service
(overlapped async dispatch) and wait (blocking stalls), but every
accumulation site wrote t_edisk and nothing ever wrote t_ewait, so the
wait column always printed 0.000s. Worse, the accounted sum used the
dead t_ewait instead of t_edisk, so the entire disk-read stall was
excluded from accounted and silently fell into the other bucket.
On a disk-streaming MoE the effect is large: other reads as ~60% of
decode when it is really the expert-load stall. Route the three
blocking sites (non-PIPE parallel load, the Metal drain barrier, and
the per-expert pipe_wait in the CPU matmul loop) to t_ewait, keep the
async dispatch in t_edisk, and include both in accounted.
Profiling-only; no behavioural change. Before, on a 168-expert model:
expert-disk 25.077s service / 0.000s wait | ... | other 28.401s
After:
expert-disk 0.111s service / 26.948s wait | ... | other 3.390s
other now holds only the genuinely-unbucketed work (router, norms).
In serve/chat mode, Ctrl-C during generation killed the whole engine, losing
the loaded model and forcing a full reload. Now a SIGINT handler (armed only
in run_serve / run_serve_mux) sets a flag that spec_decode's token loop treats
exactly like hitting the NGEN cap: the turn ends through the normal path, so
the END sentinel, STAT line, usage_save and KV append all run — and :more can
continue the interrupted answer. The mux loop closes in-flight requests via the
same mux_done path. One-shot ./glm runs and Windows keep default SIGINT (die).
coli: stream_turn survives the first KeyboardInterrupt, forwards SIGINT to the
engine (covers non-TTY), drains to the turn boundary, and reports it. A second
Ctrl-C quits. Help line and per-turn footer updated.
POSIX only (sigaction); no behaviour change on Windows. Verified end-to-end on
Apple M4 + Metal: interrupt mid-decode, engine stays up, next prompt answers,
:q exits 0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
dot_i8i8 and dot_i4i8 accumulated the whole SDOT reduction into a single
int32x4_t. SDOT has ~3-4 cycle latency, so the serial dependency on `acc`
capped each core at ~26 GB/s (int8) / ~12 GB/s (int4) of weight throughput
regardless of memory bandwidth. Split into 4 independent accumulators (64
values/iter) so the loads become the bottleneck instead of the reduction
chain; the original single-acc loop is kept as the tail handler.
Measured on an Apple M4 (isolated microbench, expert-shaped 2048x6144):
int8*int8 26.0 -> 63.2 GB/s/core (2.4x)
int4*int8 12.4 -> 29.9 GB/s/core (2.4x)
Output is bit-identical to the previous kernels (verified over random
inputs). Non-DOTPROD NEON, AVX2/AVX-512/VNNI and VSX paths are untouched;
only the __ARM_FEATURE_DOTPROD branch changed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Cut Windows decode disk I/O from 2.06s/tok to 1.70s/tok (budget=4), meeting the
<=2s/tok target. Two changes on the pread expert-load path:
1. compat.h: replace the posix_fadvise no-op with a real WILLNEED cache-warmer
(overlapped ReadFile into a scratch buffer -> populates the standby page cache
so the later synchronous pread faults from RAM). Re-arms the existing
expert_prefetch/PILOT/next-block prefetch chain on Windows. DONTNEED stays a
no-op (matches macOS; Windows standby-list trimming self-regulates).
Measured: hit rate 16.4% -> 27.6%.
2. glm.c: flip PIPE (async expert-load thread pool) from default OFF to default ON
on Windows. Dispatches expert pread onto worker threads so loads overlap the
matmul, instead of blocking serial load-then-compute. PIPE=0 opts out.
Measured: expert-disk 65.9s -> 54.3s (-18%).
Also adds compat_fadvise assertions to tests/test_compat_direct.c (data integrity
after cache-warmer, safe no-op on bad fd / non-WILLNEED).
mmap (CreateFileMapping/MapViewOfFile) was implemented and tested at length but
reverted: it regressed on Windows (RSS bloat from touched mapped pages collapsed
the expert cache via ullAvailPhys — a fundamental Windows-vs-Linux difference).
Full findings + the dead-end analysis recorded in issue_diskio.md.
With METAL=1 + COLI_METAL=1, run_serve_mux (SERVE_BATCH=1) truncated
every completion to exactly 1 token: step_decode_batch passes per-row
kvs[]/positions[] with pos_base=0, but the two Metal decode fast paths
(attention_rows and the FULL-LAYER CB in layer_forward_rows) ignored
them and dispatched coli_metal_attn_decode/coli_metal_layer_decode with
the model-bound Lc/Rc and the hardcoded pos_base. The kernels' contract
is one sequence, row s at pos_base+s: ragged rows got roped at position
0 and attended over a T=1 window of the wrong cache, so greedy decode
hit a stop token on the first batched step (DONE ... STAT 1).
Gate both fast paths on !kvs. Ragged mux rows now take the CPU absorb
path, which already reads kvs[s]/positions[s]/ks->kv_start per row;
plain serve, chat/run, prefill and MTP verification (kvs==NULL) keep
the fused GPU kernels unchanged.
Verified on GLM-5.2 int4 (M5 Max): mux with COLI_METAL=1 went from
1-token DONE to full 16/16-token greedy completions, byte-identical to
the plain-serve comparator on both test prompts; CPU-only mux was
already correct (bisection); test-c and metal-test pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The original budget dropped experts blindly — even cached ones that cost
zero disk I/O. The miss-aware version pre-scans pin/ecache residency
before applying the budget:
- ALL cache hits are kept (free to compute, no disk I/O)
- Only misses compete for the remaining budget slots
- Miss budget = EXPERT_BUDGET - nhits (min 0)
Results (budget=4, same prompt/config as before):
Original Miss-aware
tok/s 0.33 0.36 (+9%)
hit rate 16.4% 38.2% (2.3x)
prefill 8.9s 5.6s (1.6x faster)
decode 97.4s 88.9s (9% faster)
The hit rate doubling is the key quality signal: the model now gets the
full contribution from all resident experts plus the top-4 new loads,
instead of losing some hits to the budget.
Add EXPERT_BUDGET env var that caps the number of distinct experts loaded
per layer across the batch-union. When the union exceeds the budget, keeps
only the highest-aggregate-gate-weight experts and drops the rest from
idxs[] so they're never loaded from disk.
Complementary to TOPP (per-position) — this trims the cross-position union
that multiplies under MTP/prefill. Based on MoE-Spec (arXiv 2602.16052):
'top 32 of 64 experts capture 93% of routing weight.'
Measurements (GLM-5.2 744B, 24GB RAM, cap=2, MTP=0, 32 tokens):
Baseline (budget=0): 0.18 tok/s, 9.3% hit, 176s decode, 39s prefill
EXPERT_BUDGET=12: 0.19 tok/s, 14.0% hit, 171s decode, 12s prefill
EXPERT_BUDGET=6: 0.26 tok/s, 21.0% hit, 123s decode, 7s prefill
EXPERT_BUDGET=4: 0.33 tok/s, 16.4% hit, 97s decode, 9s prefill
Budget=4 nearly doubles decode speed (+83%) and 4x's prefill speed by
halving disk reads per layer. Default OFF (EXPERT_BUDGET=0).
Two independent fixes validated end-to-end on fresh fixtures:
1. KV cache disk I/O (issue_diskio.md opportunities #1 + #4):
- kv_disk_append: fopen/fclose every turn -> persistent FILE* kept open
for the engine lifetime, lazy open on first append, closed in
serve_ctx_free. Eliminates per-turn handle creation overhead.
- kv_disk_append: ~157 small fwrites per position -> one contiguous
record memcpy'd into a staging buffer then a single fwrite per
position. The staging buffer grows on demand via realloc.
- kv_disk_truncate: closes the persistent handle before truncating
so the file actually shrinks on disc, then reopens lazily.
- KVState gains disk_fp, disk_buf, disk_buf_cap fields.
- Verified: serve-mode round-trip, write 11 tokens then reload and
resume with no re-prefill, then append 8 more and reload to 19.
2. Expert weight unfusing in test-model generators:
- The real GLM-5.2-FP8 checkpoint stores routed experts UNFUSED as
per-expert 2-D tensors, each with its own _scale_inv. HF fuses
gate+up into a single 3-D gate_up_proj for compute efficiency.
- The converter and C engine both expect the unfused layout. The
fused 3-D tensors were silently skipped by the converter, and the
engine crashed with missing-tensor errors.
- New unfuse_experts in glm_fp8_emit.py splits gate_up_proj and
down_proj into per-expert 2-D tensors. Called after reference
generation but before saving, in both generators, both FP8 and bf16.
- Also fixed: make_glm_oracle.py FP8 round-trip guard used p.dim()<2
which let 3-D fused experts through and crashed fp8_block_quantize.
Changed to p.dim()!=2 to match the converter ndim!=2 guard.
Validated full chain on fresh fixtures:
generator --fp8 -> 570 e4m3 tensors + 629 scale_inv, was 90 when fused
converter --group-size 0 -> per-row int4 fmt=2, engine loads clean
converter --group-size 128 -> grouped int4 fmt=4, 8-16x more scales,
engine loads clean, fmt=4 auto-detected in both mmap and slab paths
dequant error: grouped 1.14-1.22x lower than per-row vs FP8 source
Two independent fixes validated end-to-end on fresh fixtures:
1. KV cache disk I/O (issue_diskio.md opportunities #1 + #4):
- kv_disk_append: fopen/fclose every turn -> persistent FILE* kept open
for the engine lifetime, lazy open on first append, closed in
serve_ctx_free. Eliminates per-turn handle creation overhead.
- kv_disk_append: ~157 small fwrites per position -> one contiguous
record memcpy'd into a staging buffer then a single fwrite per
position. The staging buffer grows on demand via realloc.
- kv_disk_truncate: closes the persistent handle before truncating
so the file actually shrinks on disc, then reopens lazily.
- KVState gains disk_fp, disk_buf, disk_buf_cap fields.
- Verified: serve-mode round-trip, write 11 tokens then reload and
resume with no re-prefill, then append 8 more and reload to 19.
2. Expert weight unfusing in test-model generators:
- The real GLM-5.2-FP8 checkpoint stores routed experts UNFUSED as
per-expert 2-D tensors, each with its own _scale_inv. HF fuses
gate+up into a single 3-D gate_up_proj for compute efficiency.
- The converter and C engine both expect the unfused layout. The
fused 3-D tensors were silently skipped by the converter, and the
engine crashed with missing-tensor errors.
- New unfuse_experts in glm_fp8_emit.py splits gate_up_proj and
down_proj into per-expert 2-D tensors. Called after reference
generation but before saving, in both generators, both FP8 and bf16.
- Also fixed: make_glm_oracle.py FP8 round-trip guard used p.dim()<2
which let 3-D fused experts through and crashed fp8_block_quantize.
Changed to p.dim()!=2 to match the converter ndim!=2 guard.
Validated full chain on fresh fixtures:
generator --fp8 -> 570 e4m3 tensors + 629 scale_inv, was 90 when fused
converter --group-size 0 -> per-row int4 fmt=2, engine loads clean
converter --group-size 128 -> grouped int4 fmt=4, 8-16x more scales,
engine loads clean, fmt=4 auto-detected in both mmap and slab paths
dequant error: grouped 1.14-1.22x lower than per-row vs FP8 source
Root cause of gibberish output: the int4 quantization uses one F32 scale per
output row (2048 scales for a 2048x6144 matrix). The FP8 source has 128x128
block scales — 48x finer. This destroys reasoning while keeping surface fluency.
Changes:
- Converter: quant_int4_grouped() with --group-size 128 arg. Same nibble
packing, but one scale per group of 128 elements along the input dim.
- Engine QT struct: added 'gs' field (group size, 0=per-row backward compat)
- Engine qt_from_disk: auto-detects fmt=4 when scale array is O*ceil(I/128)
elements instead of O. Old per-row models (fmt=2) work unchanged.
- Engine matmul_i4_grouped(): AVX2 kernel that applies per-group scales.
Accumulator resets at each group boundary: dot(x[grp],w[grp]) * scale[grp].
- Engine matmul_qt_ex: dispatches to grouped kernel for fmt=4 (always exact,
no IDOT approximation since the point is quality)
- Engine expert_load: both mmap and slab+pread paths detect fmt=4 from
scale array size and set gs=128
- qt_bytes: fmt=4 reports correct memory including group scales
Backward compatible: existing per-row int4 models work unchanged.
The fused gate+up pair path (matmul_i4_pair) falls back to separate
matmul_qt calls for fmt=4 — minor perf cost, correctness preserved.
Two small correctness fixes surfaced by community reports on non-Linux hosts:
#236 — expert_load's buffered pread paths (slab + qs scales) used
perror("pread expert") on a short read. Since pread returned a short count
(not -1), errno stays 0 and perror prints "Success" — a confusing message
right before exit(1) in the score/bench path. New pread_full() helper loops
over short reads and EINTR and reports actual/expected bytes and offset, so a
truncated shard reads as such instead of "Success".
#219 — Linux pulls -pthread in via -fopenmp; the *BSDs do not, so pthread_*
fail to link there. Added -pthread to the generic (Linux/*BSD x86-64) and
aarch64 CFLAGS/LDFLAGS. No-op on Linux, required on FreeBSD (complements #206).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
PIN_GB=all passed gb=-1.0 to pin_load, which took npin=n (all experts) and
ignored --ram entirely — the OOM-kill regression from #80. A 92 GB host with
--ram 78 was killed mid-generation (anon-rss ~89 GB). Now gb<0 clamps npin to
expert_avail() — how many experts fit the RAM budget, same accounting AUTOPIN
uses. pin_load already adds the pinned bytes to resident_bytes, so the later
cap_for_ram narrows the LRU accordingly with no double count.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
kv_alloc had two consecutive if(k->Lc) free blocks. The first freed every
k->Lc[i]/k->Rc[i] and both arrays without unregistering from Metal and
without nulling k->Lc; the second then re-tested the dangling pointer,
called coli_metal_unregister on freed pointers, and freed everything a
second time. Safe only when k->Lc is NULL (first call); any re-allocation
on the same KVState aborts in the allocator.
The first block is the pre-Metal version of the free path: 3716e40 (Metal
backend) replaced it with the Metal-aware block, and ec89136 (GPU resident
pipeline) re-added it above during the merge. Delete it so the Metal-aware
block is the only free path.
Caught by tests/test_kv_alloc from the previous commit:
before: malloc: *** error for object 0x7: pointer being freed was not allocated (exit 134)
after: OK kv_alloc re-allocation (exit 0)
Paper-style cache-aware MoE selection (arXiv:2412.00099 max-rank):
keep true top-J always; fill remaining K slots preferring experts already
resident in pin∪LRU within top-M. Default OFF so stock full top-K is
unchanged.
Env: CACHE_ROUTE, ROUTE_J/M/P/ALPHA, ROUTE_AGREE (auto-on with CACHE_ROUTE).
Telemetry: swap%/route_swaps/route_slots, route_agree, route_kl on footer
and serve STAT. Complementary to PILOT (prefetch vs selection change).
Routing-only PR for clean A/B vs PILOT / #119; no CUDA/fuse stack.
See docs/CACHE_ROUTE.md. Closes nothing; for #161 discussion.
Co-authored-by: Vincent Marquez <vincentmarquez405@gmail.com>
Wires the two-step prediction (kind==2 from experiment/two-step-predict)
into the real pilot_prefetch() path behind PILOT_TWO=1 env var.
Changes:
- la_predict kind==2: computes shared expert (resident, no disk I/O) on
L's post_ln-normalized state, adds output to residual, then runs L+1's
router on the corrected state. Guards n_shared==0 and mloe_inter<=0.
- pilot_prefetch(): when PILOT_TWO=1, computes the same shared-expert
correction before running the router. Workspace allocated once per
call (not per position) to avoid malloc churn.
- LOOKA measurement harness expanded to 4 slots (prev, skip-attn,
PILOT stale, two-step) with updated reporting at both exit points.
- PILOT_TWO env var wired into main().
Measurements (GLM-5.2 744B, 24GB RAM, cap=2):
LOOKA recall: PILOT stale 73.6% -> two-step 76.7% (+3.1%)
End-to-end tok/s: no change (0.16 tok/s) — cache too small (cap=2)
for prediction quality to matter; disk bandwidth saturated regardless.
On higher-RAM hosts (cap>=32) the +3.1% recall would translate to
measurably fewer disk misses.
Prior art: 'Speculating Experts' (arXiv:2603.19289) independently
developed the same idea as a 'quasi-hidden state' using a static
default vector. Our approach uses the actual computed shared expert,
which is input-dependent and more accurate.
Co-authored-by: woolcoxm <13604288+woolcoxm@users.noreply.github.com>
* Fuse CUDA expert MLP execution
* Group CUDA expert transfers by device
* Instrument grouped CUDA expert execution
* Bound grouped CUDA decode scratch
* Execute expert groups across GPUs in parallel
* Release host backing for multi-GPU experts
* Define quality-preserving memory policies
* Overlap cold expert loading with resident compute
* Adapt expert placement with session LFRU
* Fuse q4 expert gate and up dispatch
* Plan CPU work on physical cores
* Batch grouped expert CUDA kernels
* Separate VRAM and RAM expert placement
* Add ragged multi-sequence decode forward
* feat(runtime): add continuous decode scheduler
* Route concurrent API requests through batch scheduler
* Harden multiplex request lifecycle and framing
* Cancel disconnected multiplex requests
* Bind API port before starting the engine
* fix automatic KV slot allocation
* add native int4 Tensor Core grouped GEMM
* add Tensor Core throughput benchmark
* optimize packed int4 low-row kernels
* add asynchronous CUDA staging streams
* document validated six-GPU dense acceleration
* tune six-GPU expert hot set
* raise validated expert hot-set target
* add CUDA MLA absorption core
* fuse grouped expert gate and up projections
* Warn for explicit lossy routing flags
* Add full-resident expert placement mode
* Adapt VRAM expert slots to live routes
* Accelerate int4 matvec on AVX-512
* Reduce AVX-512 and RoPE decode overhead
* Seed every GPU expert layer after prefill
* Limit live GPU swaps during decode
* CUDA batch MLA attention, kv_b head-sharding, fused o_proj, expert-group dispatch, W4A16 kernels
Lab-qualified on the 6x RTX 5090 machine (914-token request benchmark):
- batch MLA absorption kernel (COLI_CUDA_ATTN=1): whole-batch attention on
device, 154.8s -> 102.4s
- attention -> o_proj fusion on the layer device: -> 97.4s
- kv_b head-sharding across cards (COLI_CUDA_ATTN_SHARD=1), no weight
duplication: -> 94.05s
- per-device expert-group dispatch with pinned-buffer async transfers,
W4A16 tensor-core kernels for the shared expert, OMP hot-thread tuning
Negative results (reverted, kept out): GPU-side weighted scatter-add
(atomics + per-layer D2H lose 43.8%), shared-expert fused small-batch
kernel (-38.8%), W4A4 grouped tensor cores (int4 activations corrupt
output). Details in the lab research log.
* GPU resident pipeline: device-resident prefill attention chain, GPU expert groups in prefill, batched router, W4A16 mixed dispatch
COLI_CUDA_PIPE=1 keeps the prefill data plane on the layer home device;
control flow (routing, cache/pin management) stays on CPU. Any CUDA
failure falls back to the unchanged CPU path.
- Device primitives + unit tests (tests/test_pipe_cuda.cu): rmsnorm
(strided), interleaved RoPE, silu-mul, residual add, fixed-order row
merge (no atomics), device-input GEMM, persistent per-device scratch.
All verified against the engine's CPU math on SM120 (worst 1.2e-5).
- attn_pipe_prefill: q_a -> norm -> q_b -> rope -> kv_a -> norm -> rope ->
batch attention -> o_proj in one device chain (q_a/q_b/kv_a colocated
with kv_b); only the final [S,D] and the new KV rows return to host.
Attention 41.2s -> 30.8s on the 1571-token benchmark.
- Prefill batch-union now uses the GPU expert groups (previously gated to
S<=64, leaving all VRAM-resident experts idle during prefill - measured
21ms of GPU expert time in a 148s prefill). Expert phase 78.9s -> 69.0s.
- Router computed as one batched matmul instead of S sequential rows
(bit-identical math).
- W4A16 tensor-core path for expert groups (COLI_CUDA_TC_W4A16=1) with
row-count mixed dispatch: >=16 rows per expert use tensor cores, smaller
batches keep the naive kernel (tensor cores measured negative below
~16 rows). Expert phase 69.0s -> 64.3s, decode unaffected.
Net on the 1571-token prefill benchmark: 148.8s -> 114.3-126.8s
(component timings stable across runs; wall drifts +-3-5s because
.coli_usage placement learning shifts the expert tiers between runs).
PROFILO now also prints the prefill-phase breakdown.
* Skip OMP hot-thread tuning when CUDA is enabled
The active-spin worker team measured 66.9s->20.9s on the CPU-only Zen5
build, but on the six-GPU full-residency workload the spinning workers
contend with the CUDA dispatch threads: ~4x slower prefill with the
process stuck near 1.8 cores. Gate the tuning on COLI_CUDA so each
configuration keeps the behavior it was measured to prefer.
* Inc.2a: sparse layers fully resident on the layer device, residual hops cards at layer boundaries
COLI_CUDA_PIPE=2 keeps the residual stream on the layer home device for
consecutive sparse layers (cudaMemcpyPeer at boundaries): in/post norms,
attention chain, both residual adds and the shared-expert MLP run on
device. Per layer only the post-norm activations (router + CPU-tier
experts + group gather), the new KV rows and, on DSA indexer layers, the
pre-attention norm leave the card. Per-layer transfers drop from ~130MB
to ~70MB. A device-side snapshot at layer entry makes any mid-layer CUDA
failure fall back to the unchanged CPU path idempotently.
1571-token prefill: 127.1s (PIPE=1 control) -> 117.6/118.9s, components
attention 30.8->26.1, other 31.8->22.5-24.5; output verified coherent
against the control.
* Head-sharded attention inside the pipe: negative on PCIe star topology, gated opt-in
Slicing q per card from the home device and collecting ctx back
serializes ~95MB/layer through the home card's PCIe link: attention
26.1s -> 41.4/44.4s on the 1571-token benchmark (two repeats), wall
117.6 -> 135-138s. The standalone host-path sharding won because six
cards uploaded from host RAM in parallel; a home-device star has no
such parallelism without NVLink. Kept behind COLI_CUDA_PIPE_SHARD=1
for interconnects where peer bandwidth does not share one root port.
* Inc.3: device-resident KV shadow for decode attention
Decode re-uploaded the whole latent+rope window per layer per token
(~300MB/token at 1571 context). Each layer now keeps a device shadow of
the compressed KV on its kv_b card, bulk-synced when behind and appended
incrementally; the host cache stays canonical. Invalidation on kv_bind
(slot switch), kv_alloc (resize) and on any overwrite of mirrored rows,
with the legacy full-upload path as fallback.
Measured (COLI_CUDA_PIPE gate): short-context decode 5.48 -> 5.59/5.87
tok/s, 1571-context decode 4.14 -> 4.22 tok/s. Decode remains CPU-expert
bound; the shadow removes the transfer tax, not the compute.
* tools: unified user-experience benchmark (bench_ux.sh)
Two fixed scenarios (short chat, long-document QA), TTFT + decode tok/s
+ first-line drift check, TEMP=0 DRAFT=0 enforced, medians over REPS
runs. Encodes the measurement discipline from the lab record: same
binary per comparison, judge medians because .coli_usage placement
learning drifts wall times between runs.
* tools: bench_ux.sh executable bit
* gitignore compiled test binaries
* tools: expert_atlas.py — measure per-expert topic affinity (#175)
Diffs .coli_usage across 10 themed probe batches (code/math/chinese/
prose/science/law/poetry/structured/translation/casual, 3 prompts each)
driven through a running API server — one engine load total. Every
touched expert gets a topic-affinity vector, entropy, and a specialist/
generalist label; output experts.json feeds the Brain page hover.
* serve: persist .coli_usage after every turn in mux mode, not only at exit
run_serve_mux saved the learning cache once at shutdown; a crash lost
the whole session's routing history, and live consumers of the file
(expert_atlas.py diffs it between probe batches) saw a frozen snapshot.
Now saved per turn like the interactive path (165KB write, negligible).
* web: Brain hover shows measured expert atlas when published
If /experts.json (from tools/expert_atlas.py, #175) is served next to
the app, the tooltip upgrades from the depth heuristic to measured
data: specialist/generalist label, entropy, and the top-3 topic
affinities. Row index maps to real layer (row+3, last row = MTP 78).
Falls back to the heuristic when no atlas is published.
---------
Co-authored-by: JustVugg <JustVugg@users.noreply.github.com>
Two Windows-only bugs in run_serve_mux left the gateway hanging after READY:
1. No _setmode(_O_BINARY): the CRT collapsed CRLF inside fread() payloads (waits
forever for missing bytes) and expanded LF in the READY/STAT sentinels.
2. WaitForSingleObject on an anonymous pipe is undefined (always-signaled or
WAIT_FAILED) and PeekNamedPipe fails on file/console handles, so the dispatch
gate never opened. New rule: idle -> block in getline (POSIX select(NULL)
semantics); active -> PeekNamedPipe poll.
Reproduced and verified on real Windows via MinGW cross-compile + WSL interop:
old binary writes READY (with CRLF corruption) then hangs forever on a crafted
SUBMIT frame; fixed binary answers DONE + STAT and exits cleanly. Linux path
untouched (0 warnings, oracle-exact).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
#152 carried leftover debug scaffolding from the lm_head measurement posted on that
PR: a `g_lmhead_exact` global, an undocumented `LMHEAD_EXACT` getenv, and 7 lm_head
call sites routed through matmul_qt_ex(..., !g_lmhead_exact). None of it was in the
PR description; it rode in because `git checkout -B` carries uncommitted working-tree
edits forward.
It should not stay:
- it does nothing worth having. Pinning lm_head off IDOT measures -0.03% perplexity
across 5 corpora (in-distribution, isolated) — nothing.
- LMHEAD_EXACT=1 would silently change lm_head numerics, undocumented and untested.
- lm_head is fmt=1 (int8), so the IDOT gate applies to it unconditionally on every
platform anyway; there is no threshold story here to expose.
The flag defaulted to 0 and `!g_lmhead_exact == 1 == matmul_qt`'s own allow_idot, so
this removal is a pure no-op. Verified rather than assumed — summed log-lik over 1023
tokens, in-distribution, CPU (deterministic), dev-with-probe vs dev-minus-probe:
prose -2295.245383 == -2295.245383
markdown -3272.146403 == -3272.146403
Bit-identical. make check green (6 C suites + 60 python tests), no new warnings.
The actual #152 changes are untouched: the three attention input projections and the
DSA indexer's ix_wk stay batched and pinned to the exact int4 kernel via
matmul_qt_ex(..., 0).
SCORE mode scored raw token streams: GLM sees [gMASK]<sop> at the start of
every training sequence, so unprefixed requests run the model out-of-
distribution and silently distort logprobs (#108). The eval harness got the
text-level fix in #194; contributors driving SCORE directly (e.g. the
perplexity work in #153) were still exposed.
Same detection rule as tools/eval_glm.py: config.json model_type contains
"glm" (case-insensitive). The two ids are looked up in the snapshot's
tokenizer.json ([gMASK], <sop>) rather than hardcoded. Requests that already
carry the prefix pass through untouched — the patched eval_glm.py sends
prefixed streams, so no double-prefixing. SCORE_PREFIX=0 restores stock
behavior. One [SCORE] stderr notice when active.
Validated on GLM-5.2 744B (M4 Pro, Metal): default run flips the '2 + 2 ='
smoke question back to ' 4' (-1.006 vs ' 5' -2.950); SCORE_PREFIX=0
reproduces stock output; pre-prefixed requests give bit-identical logprobs
to auto-prefixed ones.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Splits the expert DISK LOADS (LRU misses -> expert_load) two ways in the
final stats line of 'run':
- by context: loads issued while drafting (mtp_draft) vs absorbing
verified tokens into the MTP KV (mtp_absorb) vs verify/main forwards
- by layer kind: the MTP layer (int8 experts, ~2x bytes) vs the 78 main
layers (int4), with exact bytes read (weights + scales)
Gated behind DISK_SPLIT=1 (default OFF, parsed once at startup like the
other g_* flags): when unset no atomic is ever touched and the stats
output is byte-identical to stock. Measurement only - no effect on
routing, verification or output.
Measured on M4 Pro 48GB (GLM-5.2 int4, METAL, --ram 38): the draft path
is ~1.5% of misses and the MTP layer ~4.5% of disk bytes.
Requested in #182.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* SCHEMA=<file.json>: JSON-Schema -> GBNF compiler for grammar-forced drafts (#48/#70 follow-up)
schema_gbnf.h compiles a practical JSON-Schema subset (strict objects, string/
number/integer/boolean/null, enum/const, arrays with items, nesting) into the
byte-level GBNF subset grammar.h parses, so structured-output workloads get
grammar-forced drafts without hand-writing GBNF. Unsupported keywords fail soft:
the engine runs without a grammar and output is unchanged (drafts are verified,
never constraints - a wrong compile can only cost acceptance, not correctness).
grammar_setup: GRAMMAR= (raw GBNF) keeps precedence; SCHEMA= feeds the compiler
into the same gr_parse path. 8 test groups in tests/test_schema_gbnf.c walk
compiled grammars end-to-end through the PDA (forced spans, enum disambiguation,
nested instances, escapes, leading-zero rejection, fail-closed fallbacks).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* schema_gbnf: whitespace-tolerant emission (jws at separators)
Measured on GLM-5.2 current main (#146): the greedy continuation writes sloppy
JSON (spaces after colons, fences, long free text) and a compact-only grammar
desyncs at the first stray space, forfeiting every span after it. jws points are
not forced themselves (two legal bytes) but the multi-byte spans around them
keep drafting and the walker survives non-compact output - strictly
acceptance-positive for a verified draft source. Tests re-derived for the new
span boundaries + a sloppy-instance walk.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: JustVugg <JustVugg@users.noreply.github.com>
* cuda-dll: fix Windows build — MSVC host flags, CUDA_PATH default, POSIX setenv shim in the kernel test (#157)
First hardware validation of the #131 CUDA_DLL path (RTX PRO 6000 Blackwell
sm_120, MSVC 14.44 + CUDA 13.2, MSYS2 UCRT64 host build) found three blockers
that made 'make cuda-dll' unbuildable as shipped:
- NVCCFLAGS passed GCC-style -Xcompiler=-Wall,-Wextra to the MSVC host
compiler (hard error D8021). Use -Xcompiler=-W3 on Windows — dash form,
since MSYS make mangles /W3 into a filesystem path.
- NVCC defaulted to $(CUDA_HOME)/bin/nvcc with CUDA_HOME=/usr/local/cuda;
on Windows default CUDA_HOME from the installer's CUDA_PATH and NVCC to
plain 'nvcc' from PATH (CUDA_PATH contains spaces, which the unquoted
recipe checks cannot survive; an MSVC PATH environment is already required).
- tests/test_backend_cuda.cu used POSIX setenv/unsetenv (undefined under
MSVC); add a two-line _putenv_s shim.
Also corrects the stale '11 API symbols' comment (the header exports 15 and
backend_loader.c resolves all 15).
Validated: make cuda-dll (stock flags) + make glm CUDA_DLL=1 ARCH=native →
[CUDA] device init on sm_120, tiny oracle TF 32/32 + greedy 20/20, kernel
suite 'q8/q4/q2/f32 correctness ok', graceful no-dll fallback, plain build
byte-identical CPU behavior.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* cuda: fix heap corruption in expert_host_release on Windows (CUDA_RELEASE_HOST=1)
expert_host_release() freed the expert slab with plain free(), but the slab
is posix_memalign'd — which compat.h maps to _aligned_malloc on Windows, so
free() corrupts the CRT heap: instant 0xC0000374 crash on the first released
expert. This is the exact pattern the compat.h audit fixed at the original
expert_load site ("l'unico sito che libera memoria aligned e' free(s->slab)");
this call site was added later and reintroduced it. compat_aligned_free is
plain free on POSIX, so non-Windows behavior is unchanged. fslab stays plain
free (malloc/falloc on the CPU path).
Found running the VRAM expert tier on real hardware (80 GB resident on an
RTX PRO 6000, CUDA_RELEASE_HOST=1 to avoid 80 GB of host double-residency —
reproducible crash before, clean generation after).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: lEWFkRAD <186512915+lEWFkRAD@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: JustVugg <JustVugg@users.noreply.github.com>
* prefill: batch the attention input projections (q_a/q_b/kv_a) over the prompt
During prefill the three attention input projections were called one row at a time
(matmul_qt(..., 1)) inside the loop over the S prompt tokens, while o_proj right
below already runs batched and moe() already does a batch-union. Batching them over
all S rows lets each weight row be read once for the whole prompt instead of once
per token.
Measured on Zen5 + RTX 5090, GLM-5.2 int4, 78 layers, DSA on, OMP tuning on,
prefill run to completion (engine's own timings, upstream/dev @ 6d3ed7e):
prompt projection/RoPE total prefill
256 22.7s -> 17.7s (-22%) 120.9s -> 115.4s (-4.5%)
512 45.5s -> 36.3s (-20%) 231.4s -> 222.7s (-3.8%)
Output is BIT-IDENTICAL to dev: summed log-lik of a fixed 1023-token passage is
-4987.458316 before and after (SCORE mode, CPU, deterministic). Decode (S=1) is
untouched — same calls, same kernel.
matmul_qt_ex(..., allow_idot=0) keeps the projections on the EXACT int4 kernel.
This matters: batching alone would push S past the S>=g_i4s gate (2 on x86) and
silently move them onto IDOT's int8-quantized activations, which costs
+0.169 nats/token on markdown/code and +0.084 on prose (perplexity +18.4% / +8.7%,
measured) for only 1.4% more speed now that #95 made the exact kernel fast. The
gate is a decode-era speed threshold; it should not be crossed by a batching change.
Also: matmul_qt documents the IDOT threshold as "configurable via I4S", but the
getenv was missing and the knob did nothing. Wire it up — it is what isolates the
kernel switch from the batching win.
* prefill: batch the DSA indexer's ix_wk projection too
Same per-token bug as the attention projections, one block below: ix_wk was called
one row at a time inside the prefill loop, so its weights were re-read for every
token on every DSA layer.
matmul_qt_ex(..., 0) keeps it on the exact int4 kernel for the same reason as the
attention projections: batching alone would push S past the S>=g_i4s gate and
silently change the activation quantization.
Output stays BIT-IDENTICAL (summed log-lik over 1023 tokens, in-distribution):
prose -2295.245383 (unchanged)
markdown -3272.146403 (unchanged)
Speed, 256-token prompt, 3 alternating reps each (Zen5 + RTX 5090, dev + the
attention-projection commit as the baseline):
projection/RoPE 17.54s (sd 0.20) -> 16.96s (sd 0.15) -3.3%
total prefill 115.34s (sd 0.82) -> 113.69s (sd 0.35) -1.4%
Both separated beyond 2 sd. Small, but it is free and bit-identical.
ix_wq/ix_wp are left alone on purpose: they sit inside an `omp parallel for` and
are skipped entirely below index_topk context, so batching them is a different
change with a different risk profile.
---------
Co-authored-by: JustVugg <JustVugg@users.noreply.github.com>
Routing at layer L strongly constrains routing at L+1/L+2: measured on GLM-5.2,
co-activation lift over independence is median 1.8x / p99 40x in-domain, and the
structure TRANSFERS across workloads (a coupling table trained on prose/code
keeps +33-46% relative prefetch recall on an unseen NDJSON workload, both
depths) - it is a property of the model, not the session.
Three pieces:
- ROUTE_TRACE=<path>: zero-effect routing dump (one line per position/layer,
top-K ids:gates) from moe() FASE A.
- tools/route_pairs.py: traces -> .coli_pairs table (top-16 co-activated
L+1/L+2 experts per (layer, expert)); tools/route_coupling_report.py:
Frechet-bound screen + marginal-vs-coupled prefetch recall, with a
train-on-A/test-on-B transfer mode.
- COUPLE=<.coli_pairs> (+COUPLE_K, COUPLE_D): scores next-layer candidates by
summed pair counts over the position's routed set and enqueues non-resident
ones into the existing pilot ring (same worker, residency re-check, and
safety invariants; hints only - output byte-identical, verified).
End-to-end on M3 Max (fast NVMe, warm page cache), interleaved
baseline/couple/baseline: K=4 D=1 neutral (0.48 vs 0.48/0.50 tok/s brackets);
K=8 D=2 harmful (0.35 tok/s: ~600 hints/token = ~11 GB/token of readahead
thrashing the page cache). WILLNEED cannot move engine hit% by construction.
So the PREDICTOR is validated, the readahead ACTUATOR does not pay on this
hardware class - default OFF, small K default, and the interesting targets are
matched-latency storage (expert read ~ layer compute) and PILOT_REAL-style
loading on small-RAM boxes, both left to owners of such hardware.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Web dashboard: live expert-tier panel (VRAM/RAM/disk), static hosting, coli web command
- Engine: TIERS protocol line (counts + GB for VRAM/RAM/disk experts)
emitted at READY and refreshed after every served turn, computed from
the live pin/LRU/CUDA state.
- Server: dispatcher tracks the latest snapshot, /health exposes it as
"tiers"; the built web UI (web/dist) is served from the same port
(read-only, traversal-safe, SPA fallback) so the dashboard needs no
second process.
- Web: tier bar + legend in the Runtime panel showing where the 19k
experts live right now, with GB per tier; updates with the existing
health polling.
- CLI: `coli web` = serve + auto-open the browser once /health answers
(the 744B load takes minutes; a poller waits for it), --no-browser to
disable, friendly hint when web/dist isn't built.
* coli web: HERE is a path string, not a Path
* web: same-origin default + auto-connect when served by the engine
The page hosted by coli web talked to http://127.0.0.1:8000/v1 by
default while being loaded from http://localhost:8000 - a different
origin, so the very first probe died on CORS ('Failed to fetch').
When the page is engine-served the default (and a stored stale factory
default) now resolve to window.location.origin + /v1, and the console
auto-probes on load; the Vite dev server keeps the classic default.
The server's own port is also whitelisted in DEFAULT_CORS_ORIGINS for
the localhost/127.0.0.1 cross-name case.
* web: fix auto-connect effect returning Promise (React 19 StrictMode crash)
The async connect() call inside useEffect returned a Promise that React
19 StrictMode stored as the effect cleanup. On re-render, React called
destroy_() on that Promise — 'not a function'. Block body with explicit
return undefined prevents any non-function from leaking into the cleanup
slot.
* web: rich streaming metrics — live token counter, tok/s gauge, TTFT, session totals
During generation: a flashing Zap badge counts tokens in real time.
After: Gauge shows tok/s, Timer shows TTFT (time to first token),
Layers shows prompt→completion token counts, Clock shows queue wait.
Session totals (cumulative prompt + completion) appear in the Runtime
panel. All with lucide icons and tabular-nums for stable layout.
Tier bar gains a smoother cubic-bezier transition and slightly taller
height for visual weight.
* web: hardware environment panel — CPU model, GPU count/VRAM, RAM total/free, cores
Engine emits HWINFO protocol line at startup (CPU name from /proc/cpuinfo,
core count, RAM total/available from /proc/meminfo, GPU count + VRAM from
CUDA). Server parses and caches it, /health exposes as 'hwinfo'. Dashboard
renders it with icons at the top of the Runtime section so the user sees
immediately what the engine is running on.
* web: downgrade to React 18 — React 19 effect cleanup bug crashes the dashboard
React 19 treats the return value of every useEffect callback as a
cleanup function. In practice, any async interaction (streamChat's
rapid onDelta re-renders, the health poller, auto-connect) caused
React 19 to store a Promise as inst.destroy and crash with
'destroy_ is not a function' on the next commit cycle. The bug
reproduced in incognito, without StrictMode, and with every effect
converted to explicit block bodies returning undefined — it is a
React 19 regression in the passive-effect unmount path.
React 18 does not exhibit this behavior. Pin to React 18 until
React 19 is fixed or the codebase migrates to a pattern React 19
handles correctly. All 17 vitest pass, ErrorBoundary retained.
* web: Brain page — the expert cortex, live
A 76x256 canvas grid, one cell per expert (19,456 total): colour =
tier (green VRAM / blue RAM / grey disk), brightness = routing heat
(log-bucketed .coli_usage counts), and a white pulse that flashes on
every expert routed in the current turn and decays — you watch the
model think.
- Engine: EMAP protocol line (1 byte/expert: 2bit tier + 6bit heat,
hex) at READY and after each turn; HITS bitmap of this turn's routed
experts (set where eusage increments, cleared on emit).
- Server: parses both, GET /experts serves {rows, cols, map, hits, seq}.
- Web: Brain tab next to Chat; canvas renderer with rAF pulse decay;
hover tooltip shows layer/expert/tier/heat plus an honest depth-role
heuristic (early = surface features ... late = output shaping, MTP
row labelled as the speculative draft head).
* web: Brain page responsive — cells sized from container via ResizeObserver
Cell size derives from the wrapper's actual client box instead of fixed
1400x900, re-rendering on resize; mobile media query tightens padding,
legend and tooltip. The cortex now fills whatever screen it gets.
* win: direct I/O via FILE_FLAG_NO_BUFFERING + compat_fsize + VirtualLock primitives
compat_open_direct() gives Windows the O_DIRECT twin fd st.h already uses
on Linux/macOS: FILE_FLAG_NO_BUFFERING, same 4K-alignment contract as
O_DIRECT (the engine's DIRECT=1 path already aligns offset/len and slabs
are posix_memalign'd).
Measured on GLM-5.2 744B int4, Ryzen 9 9950X3D / 126 GB / PCIe4 NVMe
(5.8 GB/s at the engine's 19MBx8T pattern), Windows 11, MinGW GCC 16.1,
32-token greedy runs at --topp 0.7, 40 GB pin, current dev HEAD:
buffered: 0.38 tok/s (expert-disk dominates)
DIRECT=1: 0.56 tok/s (1.47x) — byte-identical greedy output vs buffered
compat_fsize() (GetFileSizeEx): CRT lseek(SEEK_END) returns -1 on
NO_BUFFERING fds (measured on UCRT); iobench uses it and gains a
NO_BUFFERING branch so disk numbers are comparable across platforms.
compat_mlock/compat_munlock: VirtualLock with working-set growth (bare
VirtualLock caps at the default working-set minimum, a few hundred KB).
Wired into the engine in the next commit.
tests/test_compat_direct.c covers the alignment contract, data integrity,
fsize on both fd kinds; skips cleanly off Windows.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* win: wire VirtualLock into mem_wire, munlock pairing in expert_host_release
MLOCK=1 was a silent no-op on Windows: pinned experts could be paged out
by working-set trimming under memory pressure. mem_wire now uses
compat_mlock (VirtualLock + working-set growth); expert_host_release
unlocks before freeing, mirroring the POSIX branch.
Validated: 39.6 GB pin wired in 17s on a 126 GB machine, zero failures;
TF oracle 32/32 with MLOCK=1.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* cuda: thread-local current-device cache in select_ctx
cudaSetDevice on every call is expensive when the serial expert loop
alternates devices. Measured on RTX 5090 + RTX 4090 (Windows, DLL
backend, pre-#68 dispatch): expert-matmul 14.3s -> 25.4s per 32 tokens
going from 1 to 2 devices, entirely per-call context switching. The
current device is per-thread in the CUDA runtime, so a thread_local
cache skips redundant switches; multi-GPU expert serving becomes
positive-scaling instead of negative.
Kernel suite passes on sm_120 + sm_89; TF oracle 32/32 dual-GPU.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: olorin <io@zyphyr.co>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
run_serve_mux (continuous-batching serve, SERVE_BATCH=1) used select() on
STDIN_FILENO to poll for incoming requests. On native Windows/MinGW,
select() on a non-socket handle routes to winsock and always returns
SOCKET_ERROR (-1), so the batch serve loop compiled and printed READY but
could never accept a request.
Replaced the platform-specific polling with a dual implementation:
- POSIX (Apple/Linux): select() as before
- Windows: WaitForSingleObject on the stdin OS handle (works on both pipe
and console handles), with PeekNamedPipe to confirm bytes are available
The rest of the serve loop is now shared across platforms — no more Windows
stub. The function is no longer guarded behind #if __APPLE__/__linux__.
Co-authored-by: woolcoxm <13604288+woolcoxm@users.noreply.github.com>