Commit Graph

223 Commits

Author SHA1 Message Date
Vincenzo c7531c43d1 Merge pull request #297 from RonitBStudent/fix/portable-target-architecture
fix(build): make portable checks target-aware
2026-07-16 07:35:51 +02:00
Vincenzo 6d190d91e1 Merge pull request #293 from woolcoxm/fix/pipe2-profiling-and-cuda-mtp-optin
cuda: fix pipe2 profile double-count + COLI_CUDA_MTP opt-in (#292)
2026-07-16 07:35:48 +02:00
Vincenzo 5f167e5964 Windows: pipe2 resident pipeline at decode + EXPERT_BUDGET ws_b cache fix (#274)
Windows: 1.07 tok/s decode via GPU resident pipeline + disk-I/O tuning (3.2x over stock)
2026-07-16 07:29:34 +02:00
RonitBStudent 53324fdcf0 fix(build): handle portable detection edge cases 2026-07-15 22:19:25 -05:00
RonitBStudent c82593ad87 fix(build): make portable checks target-aware
Select a portable architecture from the compiler target instead of forcing x86-64-v3 on every platform. On macOS, only enable Homebrew OpenMP when its header and library actually exist, preserving the dependency-free fallback.
2026-07-15 21:13:05 -05:00
woolcoxm 78c675eb8b cuda: fix pipe2 profile double-count + add COLI_CUDA_MTP opt-in (#292)
Two issues found by the diagnostic sweep in #292:

1. Profile double-count in pipe_layer_sparse: the outer t_emm span wrapped
   both the shared-expert GPU dispatch AND the moe() call, but moe()
   self-times its own t_emm internally (6 accumulation sites). Routed-expert
   matmul was counted twice -> accounted > elapsed -> 'other' went negative,
   and expert-matmul showed an inflated ~120s. Fix: split into two narrow
   spans around only the GPU work moe() does NOT cover (shared-expert
   dispatch, routed upload+add), letting moe() self-time as all other
   callers already do. Verified: 'other' now positive (14-17s),
   expert-matmul realistic (5-6s).

2. MTP blanket-disabled under CUDA (g_draft=0 when g_cuda_enabled). This
   was a conservative guard from #163 before the root cause (cold-expert
   fused-pair + IDOT kernel divergence) was fully diagnosed. GPU-resident
   experts have no divergence; the cold subset still does but achieves
   30-50% acceptance anyway (matching non-CUDA builds). Add COLI_CUDA_MTP=1
   opt-in so users can test speculation under CUDA. Default unchanged.
   Verified: COLI_CUDA_MTP=1 -> MTP ACTIVE (draft=3), 44% acceptance,
   3 forwards for 8 tokens (vs 7 without MTP).

Refs #292 #163
2026-07-15 18:33:25 -04:00
JustVugg 550ddcba83 glm(win): ignore cmd.exe's built-in PROMPT template; add COLI_PROMPT (#271)
Windows cmd.exe always exports PROMPT (its prompt template, default "$P$G")
into a child's environment. The engine's `getenv("PROMPT")` picked that up, so
running `glm.exe` from cmd for the oracle self-test instead entered text-
generation mode — loading a tokenizer the tiny model doesn't ship and failing
with "tokenizer.json: No such file". PowerShell has no PROMPT env var, so it
worked there (galmok's 32/32) but not in cmd — same command, different shell.

New coli_user_prompt(): honors COLI_PROMPT everywhere, and on Windows ignores a
PROMPT that carries cmd's $-metacodes ($P,$G,...) — a real prompt has none. cmd
users can still pass a prompt via COLI_PROMPT or a non-$ PROMPT. Verified: $P$G
-> oracle mode, "Explain recursion" -> honored, COLI_PROMPT -> honored.

Third native-Windows bug surfaced by a from-scratch main download (after the
-lpsapi link fix and the bilingual oracle message).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 23:11:35 +02:00
JustVugg 6b4f5e5fd0 glm: bilingual (EN+IT) oracle-mismatch message, lead with English (#271)
Running a real model with no PROMPT lands in oracle self-test mode, which
compares against ref_glm.json — the TINY model's oracle. The guard that
detects the vocab mismatch and points the user at PROMPT=/coli chat was
Italian-only, which read as a crash to English users (#271, galmok). Lead
with English, keep an IT footer, and add the `coli chat` hint.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 23:02:35 +02:00
JustVugg a2391587c7 Makefile(win): link -lpsapi so the MinGW build resolves GetProcessMemoryInfo
compat.h's rss_gb() calls GetProcessMemoryInfo and links psapi via
#pragma comment(lib,"psapi.lib") — an MSVC-ism. MinGW gcc ignores that pragma
(emits -Wunknown-pragmas), and the Windows LDFLAGS never linked psapi, so on a
gcc that doesn't honor the pragma (e.g. 16.1.0 UCRT) the build fails with
`undefined reference to GetProcessMemoryInfo`. Add -lpsapi to the Windows
LDFLAGS; harmless on toolchains where the pragma also resolves it. Found while
building on native Windows 11 with winlibs GCC 16.1.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 22:52:17 +02:00
woolcoxm e56d483a15 run_text: TOKENS=1 env to dump generated token ids for path A/B comparison
Opt-in (default off): prints the generated token ids to stderr after
generation, for exact comparison across decode paths (e.g. the S=1
resident pipeline vs the CPU path). Used to verify the pipe_layer_sparse
gate relaxation is token-exact vs the CPU decode path (see #273).
2026-07-15 16:31:18 -04:00
Vincenzo a37dda9eae Merge pull request #289 from woolcoxm/fix/grouped-quant-detect-arbitrary-gs
quant: generalize fmt=4 group-size detection (g64/g128/g256, not just 128)
2026-07-15 22:01:20 +02:00
Vincenzo d9d0e7b958 Merge pull request #288 from monotophic/fix/mux-1token
serve mux: keep ragged decode batches off the fused Metal kernels
2026-07-15 22:00:32 +02:00
woolcoxm 55576941cc quant: generalize fmt=4 group-size detection (g64/g128/g256, not just 128)
The grouped-int4 (fmt=4) loader hard-coded gs=128 at all three detection
sites (resident weights + the two expert-load paths: mmap and pread-slab).
A checkpoint converted with --group-size 64 (or any non-128 size) wrote a
valid file that the engine then misdetected as plain per-row int4 (fmt=2)
and read the wrong number of scales -> garbage output, silently.

The conversion path (convert_fp8_to_int4.py --group-size) already emits
arbitrary group sizes, and the compute kernel (matmul_i4_grouped) is
fully gs-generic (only constraint: gs multiple of 16, the AVX2 width).
The loader was the sole gap.

Replace the three hardcoded checks with one shared helper, detect_group_size(),
that derives gs from the scale-array byte count by probing candidate sizes
{16,32,48,64,96,128,192,256} finest-first. Data-driven: any listed size
just works; per-row int4 (ns == O*4) correctly returns gs=0 -> stays fmt=2.

Verified against real GLM-5.2 expert dims (gate/up O=2048,I=6144 and down
O=6144,I=2048): g64/g128/g256 all detect correctly, and per-row is never
misdetected as grouped. g128 (the only previously-supported size) is
unchanged -> no regression for existing checkpoints.

This unblocks the g64 lane that ZacharyZcR's #225 ablation identified as
beating shipped per-row int4 (-7.5pp vs -9.3pp) at ~19% fewer bits: the
converter can produce it, and now the engine can load it.

Refs #225
2026-07-15 15:54:37 -04:00
woolcoxm a55cdfd4b8 cuda: gate pipe2 S-threshold on device count (single-GPU S=1, multi-GPU S>=8)
Per ZacharyZcR's 6x5090 A/B on #273: the S=1 resident-pipeline relaxation is
+49% on a single GPU (5070 Ti) but a wash on multi-GPU. With layers sharded
across N devices, each resident forward at S=1 crosses P2P per layer group and
those small hops don't amortize — the same term that killed pipe x head-shard
in #111. Multi-GPU decode walls on disk service, which pipe2 can't touch.

Make the threshold device-count-dependent:
  - single GPU  (g_cuda_ndev<=1): S>=1  (the breakthrough path)
  - multi  GPU  (g_cuda_ndev> 1): S>=8  (the original prefill-only gate)
with COLI_CUDA_PIPE_S_MIN as an env override for anyone who wants to measure.

The two calibration points bracket the design space:
  1x 5070 Ti, modest CPU, "other"-bound decode  -> S=1 (+49%)
  6x 5090,    sharded,   disk-service-bound      -> S=1 a wash, keep S>=8

Refs #273 (comment)
2026-07-15 15:36:10 -04:00
woolcoxm 043014536b docs: CACHE_ROUTE breakthrough — 1.41 tok/s (4.3x over stock), misses 27%->17%
CACHE_ROUTE=1 ROUTE_J=2 ROUTE_M=12 steers the MoE router to prefer
cache-resident experts, reducing the miss rate from 27% to 17% and disk
I/O from 12.4s to 8.5s. Combined with the full optimization stack
(disk tuning + CUDA pipe2 + ws_b cache fix), this achieves 1.41 tok/s
on GLM-5.2 744B int4 / RTX 5070 Ti / 32GB RAM — a 4.3x speedup over stock.

route_agree=94.6% confirms minimal quality cost (94.6% of cache-steered
picks match the true top-K the router would have chosen).

No code change — CACHE_ROUTE is an existing engine feature, opt-in via
env vars. Documented in issue_diskio.md with the full optimization journey.
2026-07-15 13:40:51 -04:00
woolcoxm d0971ff7c1 cache: fix ws_b overcounting under EXPERT_BUDGET (cap 3->4, hit 57%->73%)
The cap_for_ram reserve for the 64-slot expert working set (ws_b = 64 × eb
= 1.21 GB) is overcounted when EXPERT_BUDGET is active. At budget=4 only
ws[0..3] are populated (not all 64), so the actual working set is 4 × eb
= 76 MB — 16x less than reserved. The excess 1.06 GB was starving the LRU
cache, capping it at 3 when budget=4 needs cap>=4.

Fix: clamp ws_b to (budget+4) × eb when EXPERT_BUDGET < 64. This raises
cap from 3 to 4, matching the budget. The cache can now hold all experts
a token needs, eliminating the LRU thrashing that caused excessive disk
re-reads (the SSD hammering).

Measured (pipe2 + full stack, budget=4, RAM_GB=28):
  tok/s: 0.85 -> 1.03  (+21%)
  hit rate: 57% -> 73%  (+28%)
  expert-disk: 18.2s -> 12.4s  (-32%)
  decode: 37.7s -> 31.0s  (-18%)

Correctness: 32/32 oracle positions.

Same fix applied to expert_avail() (the mirror function for pin budgeting).
2026-07-15 13:24:02 -04:00
woolcoxm c5b1d14a74 cuda: enable pipe_layer_sparse resident pipeline at decode (S>=1 gate)
One-line gate change: relax the pipe2 call-site gate from S>=8 to S>=1,
allowing the existing resident-pipeline (pipe_layer_sparse) to run during
single-token decode, not just prefill.

The S>=8 gate was a performance heuristic (prefill-only), not a correctness
constraint — pipe_layer_sparse is fully S-general. At decode it keeps the
residual stream x on the GPU device across all 78 layers, running rmsnorm,
residual adds, and shared-expert matmuls on-device. This eliminates the
~12.5k GPU sync interruptions per decode that caused the expert-matmul
regression (13.3s -> 9.2s), and moves the untracked 'other' CPU work
(rmsnorms, residual adds, routing) onto the GPU (29s -> 17.6s).

Measured (GLM-5.2 744B int4, RTX 5070 Ti, 32GB RAM, budget=4 + full disk stack):
  tok/s: 0.72 -> 1.07  (+49%)
  decode: 44.5s -> 29.9s  (-33%)
  expert-matmul: 13.3s -> 9.2s  (regression fixed)
  'other': 29s -> 17.6s  (-39%)

Correctness: 32/32 oracle positions (3 consecutive runs).
Configuration: COLI_CUDA=1 CUDA_DENSE=1 COLI_CUDA_ATTN=1 COLI_CUDA_PIPE=2
  CUDA_EXPERT_GB=0 EXPERT_BUDGET=4 PIPE=1 RAM_GB=28 PILOT_REAL=1 DIRECT=1
2026-07-15 12:25:16 -04:00
Vincenzo 704b4d7eeb Merge pull request #269 from michael-denyer/arm-i8mm-smmla
ARM i8mm: SMMLA fast path for the int8/int4 IDOT drivers (opt-in via ARCH=native)
2026-07-15 16:20:32 +02:00
Michael Denyer d0c9941a67 ARM i8mm: SMMLA fast path for the int8/int4 IDOT drivers
vmmlaq_s32 computes a 2x2 int32 tile (2 weight rows x 2 activation
rows) per instruction on 8-deep segments. Tile o and s in pairs,
halving weight traffic and doubling per-instruction work at S>=2.
Four independent accumulators over a 64-deep unroll keep the loop
throughput-bound (a single chained accumulator measures no better
than SDOT: latency-bound). S=1 and all tails (odd o, odd s, I not a
multiple of 16/32) keep the existing SDOT/scalar code, and scales
apply in the same order, so results are bit-identical.

Compile-time gated on __ARM_FEATURE_MATMUL_INT8. The default Darwin
build passes no -mcpu and is byte-identical (still SDOT, IDOT_KERNEL
"neon"). Opt in with ARCH=native (new Darwin Makefile knob, appends
-mcpu=<arch>), which reports IDOT_KERNEL "neon-i8mm". The same gate
lights up on any aarch64 with i8mm (Graviton3+, Grace).

test_idot grows a driver-level exactness check through matmul_qt_ex:
fmt 1 and 2, S in {2,3,4,5,8}, O in {1,2,3,64,65}, I in {16,17,100,
1408}, bitwise float equality against a plain-C reference. Green on
both build flavors.

Measured on an M5 Pro (18 threads, matmul_qt_ex microbenchmark at
GLM-5.2 expert shapes, best of 3 process runs, vs the SDOT baseline):
  gateup int4 S=8   499.7 -> 1076.6 GF/s  (+115%)
  gateup int8 S=8   512.9 -> 1166.3 GF/s  (+127%)
  down   int4 S=8   696.9 -> 1186.0 GF/s  (+70%)
  S=1 decode rows unchanged (SDOT path untouched)
2026-07-15 14:59:06 +01:00
Vincenzo 52550ff571 Merge pull request #267 from michael-denyer/fix-profile-disk-wait
profile: account expert-disk stall as wait, not other
2026-07-15 15:52:39 +02:00
Michael Denyer da684b240d profile: account expert-disk stall as wait, not other
The decode/prefill PROFILE line splits expert-disk time into service
(overlapped async dispatch) and wait (blocking stalls), but every
accumulation site wrote t_edisk and nothing ever wrote t_ewait, so the
wait column always printed 0.000s. Worse, the accounted sum used the
dead t_ewait instead of t_edisk, so the entire disk-read stall was
excluded from accounted and silently fell into the other bucket.

On a disk-streaming MoE the effect is large: other reads as ~60% of
decode when it is really the expert-load stall. Route the three
blocking sites (non-PIPE parallel load, the Metal drain barrier, and
the per-expert pipe_wait in the CPU matmul loop) to t_ewait, keep the
async dispatch in t_edisk, and include both in accounted.

Profiling-only; no behavioural change. Before, on a 168-expert model:
  expert-disk 25.077s service / 0.000s wait | ... | other 28.401s
After:
  expert-disk 0.111s service / 26.948s wait | ... | other 3.390s
other now holds only the genuinely-unbucketed work (router, norms).
2026-07-15 14:44:29 +01:00
Vincenzo dcc14217af Merge pull request #259 from withinboredom/uring
[linux] saturate disk (3-10gb/s depending on hardware), become compute bound.
2026-07-15 15:11:11 +02:00
JustVugg 3fdc6d394e Merge remote-tracking branch 'origin/dev' into pr259
# Conflicts:
#	README.md
2026-07-15 15:09:42 +02:00
Vincenzo 39a0777351 Merge pull request #261 from FECO-Admin/fix/randomuuid-fallback
web: fallback when crypto.randomUUID() is unavailable
2026-07-15 14:59:33 +02:00
Vincenzo 0fc18abd24 Merge pull request #265 from skeldoor/feat/interrupt-generation
glm.c/coli: Ctrl-C soft-stops the current turn instead of killing the engine
2026-07-15 14:59:24 +02:00
Vincenzo 87ea13c963 Merge pull request #264 from skeldoor/perf/neon-multiacc-sdot
glm.c: 4-accumulator NEON SDOT for int8/int4 expert dots (2.4x/core on Apple Silicon)
2026-07-15 14:59:12 +02:00
Vincenzo e871fb792d Merge pull request #266 from skeldoor/fix/coli-tiers-leak
coli: consume the engine's TIERS line at startup so it doesn't leak into the first answer
2026-07-15 14:57:08 +02:00
JustVugg da260c38cb server: clamp max_tokens to the server cap instead of rejecting it (fixes #260)
opencode / the ai-sdk OpenAI-compatible client sends a large default max_tokens
(> the server's --max-tokens cap, 1024 by default), and generation_options
returned 400 "must be an integer between 1 and 1024" — even for a trivial
"hello". OpenAI-compatible servers clamp to their own ceiling rather than
reject. Now max_tokens > limit is clamped to limit; only non-int / non-positive
values are a hard error. Test updated to assert the clamp + keep the <1 reject.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 14:56:01 +02:00
Skeldoor b79610852b coli: consume the engine's TIERS line at startup so it doesn't leak into the first answer
After the READY handshake, run_serve emits one TIERS status line (the
web-dashboard expert-pyramid snapshot). cmd_chat never read it, so in
interactive chat it surfaced as literal "TIERS 0 181 19275 0.00 ..." text
prepended to the first response. Read and discard that one line after READY,
matching how the other status frames are consumed. Chat-only; the HTTP serve
path already drains it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:37:36 +01:00
Skeldoor 3745cc680e glm.c/coli: Ctrl-C soft-stops the current turn instead of killing the engine
In serve/chat mode, Ctrl-C during generation killed the whole engine, losing
the loaded model and forcing a full reload. Now a SIGINT handler (armed only
in run_serve / run_serve_mux) sets a flag that spec_decode's token loop treats
exactly like hitting the NGEN cap: the turn ends through the normal path, so
the END sentinel, STAT line, usage_save and KV append all run — and :more can
continue the interrupted answer. The mux loop closes in-flight requests via the
same mux_done path. One-shot ./glm runs and Windows keep default SIGINT (die).

coli: stream_turn survives the first KeyboardInterrupt, forwards SIGINT to the
engine (covers non-TTY), drains to the turn boundary, and reports it. A second
Ctrl-C quits. Help line and per-turn footer updated.

POSIX only (sigaction); no behaviour change on Windows. Verified end-to-end on
Apple M4 + Metal: interrupt mid-decode, engine stays up, next prompt answers,
:q exits 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:37:02 +01:00
Skeldoor af421d6d3a glm.c: 4-accumulator NEON SDOT for int8/int4 expert dots (2.4x/core on Apple Silicon)
dot_i8i8 and dot_i4i8 accumulated the whole SDOT reduction into a single
int32x4_t. SDOT has ~3-4 cycle latency, so the serial dependency on `acc`
capped each core at ~26 GB/s (int8) / ~12 GB/s (int4) of weight throughput
regardless of memory bandwidth. Split into 4 independent accumulators (64
values/iter) so the loads become the bottleneck instead of the reduction
chain; the original single-acc loop is kept as the tail handler.

Measured on an Apple M4 (isolated microbench, expert-shaped 2048x6144):
  int8*int8  26.0 -> 63.2 GB/s/core  (2.4x)
  int4*int8  12.4 -> 29.9 GB/s/core  (2.4x)
Output is bit-identical to the previous kernels (verified over random
inputs). Non-DOTPROD NEON, AVX2/AVX-512/VNNI and VSX paths are untouched;
only the __ARM_FEATURE_DOTPROD branch changed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:28:05 +01:00
rene 869e80ee2d web: fallback when crypto.randomUUID() is unavailable
Some browsers and insecure contexts (plain HTTP) throw on
crypto.randomUUID(), which silently breaks the chat UI.
Fall back to a plain UUID v4 generator when the native call
fails - same behaviour on modern browsers, works everywhere
else.
2026-07-15 14:20:50 +02:00
Vincenzo 5d4c3aa11b Merge pull request #257 from woolcoxm/windows-optimizations
Windows disk I/O: pread + PIPE + compat_fadvise (1.70s/tok, mmap reverted)
2026-07-15 13:42:23 +02:00
Vincenzo 4d0438314a Merge pull request #246 from woolcoxm/feat/diskio-kv-batching
Disk I/O: KV cache write batching + persistent file handle
2026-07-15 13:35:03 +02:00
Vincenzo a6895445a9 Merge pull request #254 from woolcoxm/feat/expert-budget
EXPERT_BUDGET=N: miss-aware cap on distinct experts/layer (+75% decode tok/s, 6x prefill on low-RAM hosts)
2026-07-15 13:34:19 +02:00
Vincenzo d9266e39e3 Merge pull request #248 from ZacharyZcR/serve/mux-kv-diag
serve: emit the [API] KV prefix-reuse diagnostic in mux mode (#153)
2026-07-15 13:22:57 +02:00
Vincenzo d99d1a7a0e Merge pull request #255 from ZacharyZcR/tools/quant-e8
tools: E8 lattice quantization in the ablation harness — the missing QuIP# ingredient (#81)
2026-07-15 13:20:44 +02:00
Vincenzo e51ce378ca Merge pull request #245 from ZacharyZcR/tools/atlas-unify
tools: one expert atlas, not two — retire expert_atlas.py, analyze.py gains --web output for the dashboard
2026-07-15 13:20:37 +02:00
Robert Landers 0b9ce242e4 fix wording
Signed-off-by: Robert Landers <landers.robert@gmail.com>
2026-07-15 12:42:13 +02:00
Robert Landers cf062010d8 use async and fix workers
Signed-off-by: Robert Landers <landers.robert@gmail.com>
2026-07-15 12:42:13 +02:00
Robert Landers 5d1eb142ab port uring implementation
Signed-off-by: Robert Landers <landers.robert@gmail.com>
2026-07-15 12:42:10 +02:00
woolcoxm 25219de45b windows: PIPE default ON + compat_fadvise WILLNEED cache-warmer (pread path)
Cut Windows decode disk I/O from 2.06s/tok to 1.70s/tok (budget=4), meeting the
<=2s/tok target. Two changes on the pread expert-load path:

1. compat.h: replace the posix_fadvise no-op with a real WILLNEED cache-warmer
   (overlapped ReadFile into a scratch buffer -> populates the standby page cache
   so the later synchronous pread faults from RAM). Re-arms the existing
   expert_prefetch/PILOT/next-block prefetch chain on Windows. DONTNEED stays a
   no-op (matches macOS; Windows standby-list trimming self-regulates).
   Measured: hit rate 16.4% -> 27.6%.

2. glm.c: flip PIPE (async expert-load thread pool) from default OFF to default ON
   on Windows. Dispatches expert pread onto worker threads so loads overlap the
   matmul, instead of blocking serial load-then-compute. PIPE=0 opts out.
   Measured: expert-disk 65.9s -> 54.3s (-18%).

Also adds compat_fadvise assertions to tests/test_compat_direct.c (data integrity
after cache-warmer, safe no-op on bad fd / non-WILLNEED).

mmap (CreateFileMapping/MapViewOfFile) was implemented and tested at length but
reverted: it regressed on Windows (RSS bloat from touched mapped pages collapsed
the expert cache via ullAvailPhys — a fundamental Windows-vs-Linux difference).
Full findings + the dead-end analysis recorded in issue_diskio.md.
2026-07-15 06:14:15 -04:00
monotophic 0811730845 serve mux: keep ragged decode batches off the fused Metal kernels
With METAL=1 + COLI_METAL=1, run_serve_mux (SERVE_BATCH=1) truncated
every completion to exactly 1 token: step_decode_batch passes per-row
kvs[]/positions[] with pos_base=0, but the two Metal decode fast paths
(attention_rows and the FULL-LAYER CB in layer_forward_rows) ignored
them and dispatched coli_metal_attn_decode/coli_metal_layer_decode with
the model-bound Lc/Rc and the hardcoded pos_base. The kernels' contract
is one sequence, row s at pos_base+s: ragged rows got roped at position
0 and attended over a T=1 window of the wrong cache, so greedy decode
hit a stop token on the first batched step (DONE ... STAT 1).

Gate both fast paths on !kvs. Ragged mux rows now take the CPU absorb
path, which already reads kvs[s]/positions[s]/ks->kv_start per row;
plain serve, chat/run, prefill and MTP verification (kvs==NULL) keep
the fused GPU kernels unchanged.

Verified on GLM-5.2 int4 (M5 Max): mux with COLI_METAL=1 went from
1-token DONE to full 16/16-token greedy completions, byte-identical to
the plain-serve comparator on both test prompts; CPU-only mux was
already correct (bisection); test-c and metal-test pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 05:54:50 -04:00
ZacharyZcR 0189a9f0ae tools: E8 lattice quantization for the ablation harness — int{2,3,4,8}[-gN][-e8|-e8u][-rot] (#81) 2026-07-15 17:24:16 +08:00
woolcoxm 9c24b1eceb Merge branch 'feat/expert-budget' into windows-dev 2026-07-15 03:23:25 -04:00
woolcoxm dd0d60692d experiment: miss-aware EXPERT_BUDGET — keep all cache hits, only drop misses
The original budget dropped experts blindly — even cached ones that cost
zero disk I/O. The miss-aware version pre-scans pin/ecache residency
before applying the budget:
  - ALL cache hits are kept (free to compute, no disk I/O)
  - Only misses compete for the remaining budget slots
  - Miss budget = EXPERT_BUDGET - nhits (min 0)

Results (budget=4, same prompt/config as before):
                    Original    Miss-aware
  tok/s             0.33        0.36 (+9%)
  hit rate          16.4%       38.2% (2.3x)
  prefill           8.9s        5.6s (1.6x faster)
  decode            97.4s       88.9s (9% faster)

The hit rate doubling is the key quality signal: the model now gets the
full contribution from all resident experts plus the top-4 new loads,
instead of losing some hits to the budget.
2026-07-15 02:34:58 -04:00
woolcoxm 413b370fbf experiment: EXPERT_BUDGET=N — cap distinct experts/layer, up to 1.8x faster decode
Add EXPERT_BUDGET env var that caps the number of distinct experts loaded
per layer across the batch-union. When the union exceeds the budget, keeps
only the highest-aggregate-gate-weight experts and drops the rest from
idxs[] so they're never loaded from disk.

Complementary to TOPP (per-position) — this trims the cross-position union
that multiplies under MTP/prefill. Based on MoE-Spec (arXiv 2602.16052):
'top 32 of 64 experts capture 93% of routing weight.'

Measurements (GLM-5.2 744B, 24GB RAM, cap=2, MTP=0, 32 tokens):
  Baseline (budget=0):  0.18 tok/s, 9.3% hit,  176s decode, 39s prefill
  EXPERT_BUDGET=12:     0.19 tok/s, 14.0% hit, 171s decode, 12s prefill
  EXPERT_BUDGET=6:      0.26 tok/s, 21.0% hit, 123s decode,  7s prefill
  EXPERT_BUDGET=4:      0.33 tok/s, 16.4% hit,  97s decode,  9s prefill

Budget=4 nearly doubles decode speed (+83%) and 4x's prefill speed by
halving disk reads per layer. Default OFF (EXPERT_BUDGET=0).
2026-07-15 02:34:57 -04:00
woolcoxm 71c262ce1a diskio: KV write batching + persistent handle; generators: unfuse experts
Two independent fixes validated end-to-end on fresh fixtures:

1. KV cache disk I/O (issue_diskio.md opportunities #1 + #4):
   - kv_disk_append: fopen/fclose every turn -> persistent FILE* kept open
     for the engine lifetime, lazy open on first append, closed in
     serve_ctx_free. Eliminates per-turn handle creation overhead.
   - kv_disk_append: ~157 small fwrites per position -> one contiguous
     record memcpy'd into a staging buffer then a single fwrite per
     position. The staging buffer grows on demand via realloc.
   - kv_disk_truncate: closes the persistent handle before truncating
     so the file actually shrinks on disc, then reopens lazily.
   - KVState gains disk_fp, disk_buf, disk_buf_cap fields.
   - Verified: serve-mode round-trip, write 11 tokens then reload and
     resume with no re-prefill, then append 8 more and reload to 19.

2. Expert weight unfusing in test-model generators:
   - The real GLM-5.2-FP8 checkpoint stores routed experts UNFUSED as
     per-expert 2-D tensors, each with its own _scale_inv. HF fuses
     gate+up into a single 3-D gate_up_proj for compute efficiency.
   - The converter and C engine both expect the unfused layout. The
     fused 3-D tensors were silently skipped by the converter, and the
     engine crashed with missing-tensor errors.
   - New unfuse_experts in glm_fp8_emit.py splits gate_up_proj and
     down_proj into per-expert 2-D tensors. Called after reference
     generation but before saving, in both generators, both FP8 and bf16.
   - Also fixed: make_glm_oracle.py FP8 round-trip guard used p.dim()<2
     which let 3-D fused experts through and crashed fp8_block_quantize.
     Changed to p.dim()!=2 to match the converter ndim!=2 guard.

Validated full chain on fresh fixtures:
  generator --fp8 -> 570 e4m3 tensors + 629 scale_inv, was 90 when fused
  converter --group-size 0  -> per-row int4 fmt=2, engine loads clean
  converter --group-size 128 -> grouped int4 fmt=4, 8-16x more scales,
    engine loads clean, fmt=4 auto-detected in both mmap and slab paths
  dequant error: grouped 1.14-1.22x lower than per-row vs FP8 source
2026-07-15 02:32:12 -04:00
woolcoxm e71d4fbe29 windows-dev: CUDA build script, expert budget + grouped quant research
- build_cuda.bat: one-shot nvcc DLL compilation for sm_120 (RTX 5070 Ti)
- issue_budget.md: EXPERT_BUDGET research (cap distinct experts/layer, 2x tok/s)
- issue_grouped_quant.md: root cause of int4 incoherence (per-row vs group-128 scales)
- glm_fp8_emit.py: FP8 e4m3 test weight generator for converter validation
2026-07-15 02:32:12 -04:00
woolcoxm 2c6946c478 test-models: add --fp8 emission for the FP8->int4 converter test path
Both test-model generators (make_glm_oracle.py, make_glm_bench_model.py) can now
emit weights as FP8 e4m3 + 128x128 block scale_inv, in the same layout as the real
GLM-5.2-FP8 checkpoint. This lets convert_fp8_to_int4.py exercise its FP8->int4
dequant path on a local fixture without the 379 GB download.

- New shared helper glm_fp8_emit.py: FP8 block quantize/dequantize (FBGEMM/TE
  scale=amax/448 convention) + state_dict emitter. Only exactly-2-D tensors are
  quantized; 1-D/3-D and norms/router/e_score_correction_bias are kept as f32,
  mirroring the converter's classify() + ndim!=2 guard.
- make_glm_bench_model.py: opt-in --fp8 writes model.safetensors in FP8 layout
  (config.json written explicitly since the FP8 path bypasses save_pretrained);
  manifest gains a 'format' field. Default bf16 behavior unchanged.
- make_glm_oracle.py: opt-in --fp8 round-trips quantizable weights through FP8
  before computing ref_glm.json, so the reference reflects exactly the FP8 model
  the converter ingests. Default bf16 oracle contract unchanged.

Verified end-to-end: FP8 model -> converter --indir -> int4 U8 + .qs F32 output,
bit-identical dequant between helper and converter (maxdiff 0.0).
2026-07-15 02:32:12 -04:00