Commit Graph

265 Commits

Author SHA1 Message Date
woolcoxm d0971ff7c1 cache: fix ws_b overcounting under EXPERT_BUDGET (cap 3->4, hit 57%->73%)
The cap_for_ram reserve for the 64-slot expert working set (ws_b = 64 × eb
= 1.21 GB) is overcounted when EXPERT_BUDGET is active. At budget=4 only
ws[0..3] are populated (not all 64), so the actual working set is 4 × eb
= 76 MB — 16x less than reserved. The excess 1.06 GB was starving the LRU
cache, capping it at 3 when budget=4 needs cap>=4.

Fix: clamp ws_b to (budget+4) × eb when EXPERT_BUDGET < 64. This raises
cap from 3 to 4, matching the budget. The cache can now hold all experts
a token needs, eliminating the LRU thrashing that caused excessive disk
re-reads (the SSD hammering).

Measured (pipe2 + full stack, budget=4, RAM_GB=28):
  tok/s: 0.85 -> 1.03  (+21%)
  hit rate: 57% -> 73%  (+28%)
  expert-disk: 18.2s -> 12.4s  (-32%)
  decode: 37.7s -> 31.0s  (-18%)

Correctness: 32/32 oracle positions.

Same fix applied to expert_avail() (the mirror function for pin budgeting).
2026-07-15 13:24:02 -04:00
KingIcyCreamProjects 939b689a10 glm: hwinfo_emit Windows support — CPU brand, cores, RAM were hardcoded-zero
The HWINFO snapshot read /proc/cpuinfo, /proc/meminfo and
sysconf(_SC_NPROCESSORS_ONLN) — all absent on native Windows — so the
dashboard runtime panel showed "0 GB RAM / 0 cores" while the CUDA branch
populated VRAM fine. Now: CPUID brand string (0x80000002..4, no new
deps), GetSystemInfo for logical cores, and the existing compat_meminfo
(GlobalMemoryStatusEx) for RAM. Verified live: panel reads "AMD Ryzen 9
9950X3D 16-Core Processor / 66 GB RAM 24 GB free / 32 cores".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 12:14:27 -05:00
KingIcyCreamProjects 1a243cfc3e coli: measured Windows launcher defaults — OMP tuning parity + DIRECT/PIPE/PILOT_REAL
On Windows the engine self-exec OMP tuning never runs (Linux/FreeBSD-only)
and posix_fadvise readahead is a compat.h no-op, so a stock Windows run
leaves large measured wins on the table. The launcher now setdefaults, on
win32 only, each independently overridable by setting the variable:

- OMP_WAIT_POLICY=active, GOMP_SPINCOUNT=200000, OMP_DYNAMIC=FALSE,
  OMP_NUM_THREADS=<physical cores> (parity with the glm.c self-exec block;
  COLI_NO_OMP_TUNE disables exactly this block, presence-based like the
  engine). OMP_PROC_BIND/OMP_PLACES deliberately omitted and also removed
  from environment_for_plan on win32: MinGW libgomp has no affinity support
  ("Affinity not supported on this configuration").
- DIRECT=1: unbuffered expert reads. Measured on a 9950X3D + Samsung 9100
  PRO Gen5 + Win11: iobench 10.68 GB/s O_DIRECT vs 9.03 buffered (warm);
  end-to-end REPLAY 0.48 -> 1.02 tok/s. Matches #162 (1.47x same class).
- PIPE=1: load/matmul overlap, byte-identical output; +8% on top of DIRECT
  (PIPE_WORKERS untouched at 8 - 4/8/16 swept flat on Gen5).
- PILOT_REAL=1: real cross-layer prefetch, the only working prefetch on
  Windows; +11% and expert hit rate +19 points.

Full ladder methodology and numbers: 96-token greedy REPLAY, one lever per
step, medians of 3-4 runs (see the fork tuning doc referenced in the PR).
tests/test_env_defaults.py covers the defaults, explicit-override-wins,
the kill-switch scope, and the non-win32 no-op.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 12:06:48 -05:00
KingIcyCreamProjects 4059e10761 docs(readme): mechanical fixes — expert count, IO_THREADS, line counts, desktop/
- Expert count 21,504 -> 19,456 (75x256 + MTP head), matching the rest of the repo.
- Replace phantom IO_THREADS with PIPE_WORKERS (default 8); state the pool only
  engages under PIPE=1.
- Drop rotting precise line-count claims (glm.c ~2,400; web ~390) — keep the prose.
- Add the desktop/ directory to the repo-layout section.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 12:06:14 -05:00
KingIcyCreamProjects 51dd4805af benchmark_cuda_fixture: accept the current PROFILE service/wait format
profile_print now emits 'expert-disk N.NNNs service / N.NNNs wait' but
the harness regex still expected the old single expert-disk number, so
every run died with 'benchmark output missing'. Accept both formats and
report disk = service + wait.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 12:06:14 -05:00
KingIcyCreamProjects f72ee6532d resource_plan: count physical cores on Windows (GetLogicalProcessorInformationEx)
os.cpu_count() returns logical processors, so on SMT machines the plan
sets OMP_NUM_THREADS to 2 threads/core, which thrashes the AVX-512 units
during expert matmul (9950X3D: 32 logical vs 16 physical). Count
RelationProcessorCore records instead, with the existing lscpu/cpu_count
fallbacks intact.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 12:05:06 -05:00
woolcoxm c5b1d14a74 cuda: enable pipe_layer_sparse resident pipeline at decode (S>=1 gate)
One-line gate change: relax the pipe2 call-site gate from S>=8 to S>=1,
allowing the existing resident-pipeline (pipe_layer_sparse) to run during
single-token decode, not just prefill.

The S>=8 gate was a performance heuristic (prefill-only), not a correctness
constraint — pipe_layer_sparse is fully S-general. At decode it keeps the
residual stream x on the GPU device across all 78 layers, running rmsnorm,
residual adds, and shared-expert matmuls on-device. This eliminates the
~12.5k GPU sync interruptions per decode that caused the expert-matmul
regression (13.3s -> 9.2s), and moves the untracked 'other' CPU work
(rmsnorms, residual adds, routing) onto the GPU (29s -> 17.6s).

Measured (GLM-5.2 744B int4, RTX 5070 Ti, 32GB RAM, budget=4 + full disk stack):
  tok/s: 0.72 -> 1.07  (+49%)
  decode: 44.5s -> 29.9s  (-33%)
  expert-matmul: 13.3s -> 9.2s  (regression fixed)
  'other': 29s -> 17.6s  (-39%)

Correctness: 32/32 oracle positions (3 consecutive runs).
Configuration: COLI_CUDA=1 CUDA_DENSE=1 COLI_CUDA_ATTN=1 COLI_CUDA_PIPE=2
  CUDA_EXPERT_GB=0 EXPERT_BUDGET=4 PIPE=1 RAM_GB=28 PILOT_REAL=1 DIRECT=1
2026-07-15 12:25:16 -04:00
Vincenzo 704b4d7eeb Merge pull request #269 from michael-denyer/arm-i8mm-smmla
ARM i8mm: SMMLA fast path for the int8/int4 IDOT drivers (opt-in via ARCH=native)
2026-07-15 16:20:32 +02:00
Michael Denyer d0c9941a67 ARM i8mm: SMMLA fast path for the int8/int4 IDOT drivers
vmmlaq_s32 computes a 2x2 int32 tile (2 weight rows x 2 activation
rows) per instruction on 8-deep segments. Tile o and s in pairs,
halving weight traffic and doubling per-instruction work at S>=2.
Four independent accumulators over a 64-deep unroll keep the loop
throughput-bound (a single chained accumulator measures no better
than SDOT: latency-bound). S=1 and all tails (odd o, odd s, I not a
multiple of 16/32) keep the existing SDOT/scalar code, and scales
apply in the same order, so results are bit-identical.

Compile-time gated on __ARM_FEATURE_MATMUL_INT8. The default Darwin
build passes no -mcpu and is byte-identical (still SDOT, IDOT_KERNEL
"neon"). Opt in with ARCH=native (new Darwin Makefile knob, appends
-mcpu=<arch>), which reports IDOT_KERNEL "neon-i8mm". The same gate
lights up on any aarch64 with i8mm (Graviton3+, Grace).

test_idot grows a driver-level exactness check through matmul_qt_ex:
fmt 1 and 2, S in {2,3,4,5,8}, O in {1,2,3,64,65}, I in {16,17,100,
1408}, bitwise float equality against a plain-C reference. Green on
both build flavors.

Measured on an M5 Pro (18 threads, matmul_qt_ex microbenchmark at
GLM-5.2 expert shapes, best of 3 process runs, vs the SDOT baseline):
  gateup int4 S=8   499.7 -> 1076.6 GF/s  (+115%)
  gateup int8 S=8   512.9 -> 1166.3 GF/s  (+127%)
  down   int4 S=8   696.9 -> 1186.0 GF/s  (+70%)
  S=1 decode rows unchanged (SDOT path untouched)
2026-07-15 14:59:06 +01:00
Vincenzo 52550ff571 Merge pull request #267 from michael-denyer/fix-profile-disk-wait
profile: account expert-disk stall as wait, not other
2026-07-15 15:52:39 +02:00
Michael Denyer da684b240d profile: account expert-disk stall as wait, not other
The decode/prefill PROFILE line splits expert-disk time into service
(overlapped async dispatch) and wait (blocking stalls), but every
accumulation site wrote t_edisk and nothing ever wrote t_ewait, so the
wait column always printed 0.000s. Worse, the accounted sum used the
dead t_ewait instead of t_edisk, so the entire disk-read stall was
excluded from accounted and silently fell into the other bucket.

On a disk-streaming MoE the effect is large: other reads as ~60% of
decode when it is really the expert-load stall. Route the three
blocking sites (non-PIPE parallel load, the Metal drain barrier, and
the per-expert pipe_wait in the CPU matmul loop) to t_ewait, keep the
async dispatch in t_edisk, and include both in accounted.

Profiling-only; no behavioural change. Before, on a 168-expert model:
  expert-disk 25.077s service / 0.000s wait | ... | other 28.401s
After:
  expert-disk 0.111s service / 26.948s wait | ... | other 3.390s
other now holds only the genuinely-unbucketed work (router, norms).
2026-07-15 14:44:29 +01:00
Vincenzo dcc14217af Merge pull request #259 from withinboredom/uring
[linux] saturate disk (3-10gb/s depending on hardware), become compute bound.
2026-07-15 15:11:11 +02:00
JustVugg 3fdc6d394e Merge remote-tracking branch 'origin/dev' into pr259
# Conflicts:
#	README.md
2026-07-15 15:09:42 +02:00
Vincenzo 39a0777351 Merge pull request #261 from FECO-Admin/fix/randomuuid-fallback
web: fallback when crypto.randomUUID() is unavailable
2026-07-15 14:59:33 +02:00
Vincenzo 0fc18abd24 Merge pull request #265 from skeldoor/feat/interrupt-generation
glm.c/coli: Ctrl-C soft-stops the current turn instead of killing the engine
2026-07-15 14:59:24 +02:00
Vincenzo 87ea13c963 Merge pull request #264 from skeldoor/perf/neon-multiacc-sdot
glm.c: 4-accumulator NEON SDOT for int8/int4 expert dots (2.4x/core on Apple Silicon)
2026-07-15 14:59:12 +02:00
Vincenzo e871fb792d Merge pull request #266 from skeldoor/fix/coli-tiers-leak
coli: consume the engine's TIERS line at startup so it doesn't leak into the first answer
2026-07-15 14:57:08 +02:00
JustVugg da260c38cb server: clamp max_tokens to the server cap instead of rejecting it (fixes #260)
opencode / the ai-sdk OpenAI-compatible client sends a large default max_tokens
(> the server's --max-tokens cap, 1024 by default), and generation_options
returned 400 "must be an integer between 1 and 1024" — even for a trivial
"hello". OpenAI-compatible servers clamp to their own ceiling rather than
reject. Now max_tokens > limit is clamped to limit; only non-int / non-positive
values are a hard error. Test updated to assert the clamp + keep the <1 reject.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 14:56:01 +02:00
Skeldoor b79610852b coli: consume the engine's TIERS line at startup so it doesn't leak into the first answer
After the READY handshake, run_serve emits one TIERS status line (the
web-dashboard expert-pyramid snapshot). cmd_chat never read it, so in
interactive chat it surfaced as literal "TIERS 0 181 19275 0.00 ..." text
prepended to the first response. Read and discard that one line after READY,
matching how the other status frames are consumed. Chat-only; the HTTP serve
path already drains it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:37:36 +01:00
Skeldoor 3745cc680e glm.c/coli: Ctrl-C soft-stops the current turn instead of killing the engine
In serve/chat mode, Ctrl-C during generation killed the whole engine, losing
the loaded model and forcing a full reload. Now a SIGINT handler (armed only
in run_serve / run_serve_mux) sets a flag that spec_decode's token loop treats
exactly like hitting the NGEN cap: the turn ends through the normal path, so
the END sentinel, STAT line, usage_save and KV append all run — and :more can
continue the interrupted answer. The mux loop closes in-flight requests via the
same mux_done path. One-shot ./glm runs and Windows keep default SIGINT (die).

coli: stream_turn survives the first KeyboardInterrupt, forwards SIGINT to the
engine (covers non-TTY), drains to the turn boundary, and reports it. A second
Ctrl-C quits. Help line and per-turn footer updated.

POSIX only (sigaction); no behaviour change on Windows. Verified end-to-end on
Apple M4 + Metal: interrupt mid-decode, engine stays up, next prompt answers,
:q exits 0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:37:02 +01:00
Skeldoor af421d6d3a glm.c: 4-accumulator NEON SDOT for int8/int4 expert dots (2.4x/core on Apple Silicon)
dot_i8i8 and dot_i4i8 accumulated the whole SDOT reduction into a single
int32x4_t. SDOT has ~3-4 cycle latency, so the serial dependency on `acc`
capped each core at ~26 GB/s (int8) / ~12 GB/s (int4) of weight throughput
regardless of memory bandwidth. Split into 4 independent accumulators (64
values/iter) so the loads become the bottleneck instead of the reduction
chain; the original single-acc loop is kept as the tail handler.

Measured on an Apple M4 (isolated microbench, expert-shaped 2048x6144):
  int8*int8  26.0 -> 63.2 GB/s/core  (2.4x)
  int4*int8  12.4 -> 29.9 GB/s/core  (2.4x)
Output is bit-identical to the previous kernels (verified over random
inputs). Non-DOTPROD NEON, AVX2/AVX-512/VNNI and VSX paths are untouched;
only the __ARM_FEATURE_DOTPROD branch changed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 13:28:05 +01:00
rene 869e80ee2d web: fallback when crypto.randomUUID() is unavailable
Some browsers and insecure contexts (plain HTTP) throw on
crypto.randomUUID(), which silently breaks the chat UI.
Fall back to a plain UUID v4 generator when the native call
fails - same behaviour on modern browsers, works everywhere
else.
2026-07-15 14:20:50 +02:00
Vincenzo 5d4c3aa11b Merge pull request #257 from woolcoxm/windows-optimizations
Windows disk I/O: pread + PIPE + compat_fadvise (1.70s/tok, mmap reverted)
2026-07-15 13:42:23 +02:00
Vincenzo 4d0438314a Merge pull request #246 from woolcoxm/feat/diskio-kv-batching
Disk I/O: KV cache write batching + persistent file handle
2026-07-15 13:35:03 +02:00
Vincenzo a6895445a9 Merge pull request #254 from woolcoxm/feat/expert-budget
EXPERT_BUDGET=N: miss-aware cap on distinct experts/layer (+75% decode tok/s, 6x prefill on low-RAM hosts)
2026-07-15 13:34:19 +02:00
Vincenzo d9266e39e3 Merge pull request #248 from ZacharyZcR/serve/mux-kv-diag
serve: emit the [API] KV prefix-reuse diagnostic in mux mode (#153)
2026-07-15 13:22:57 +02:00
Vincenzo d99d1a7a0e Merge pull request #255 from ZacharyZcR/tools/quant-e8
tools: E8 lattice quantization in the ablation harness — the missing QuIP# ingredient (#81)
2026-07-15 13:20:44 +02:00
Vincenzo e51ce378ca Merge pull request #245 from ZacharyZcR/tools/atlas-unify
tools: one expert atlas, not two — retire expert_atlas.py, analyze.py gains --web output for the dashboard
2026-07-15 13:20:37 +02:00
Robert Landers 0b9ce242e4 fix wording
Signed-off-by: Robert Landers <landers.robert@gmail.com>
2026-07-15 12:42:13 +02:00
Robert Landers cf062010d8 use async and fix workers
Signed-off-by: Robert Landers <landers.robert@gmail.com>
2026-07-15 12:42:13 +02:00
Robert Landers 5d1eb142ab port uring implementation
Signed-off-by: Robert Landers <landers.robert@gmail.com>
2026-07-15 12:42:10 +02:00
woolcoxm 25219de45b windows: PIPE default ON + compat_fadvise WILLNEED cache-warmer (pread path)
Cut Windows decode disk I/O from 2.06s/tok to 1.70s/tok (budget=4), meeting the
<=2s/tok target. Two changes on the pread expert-load path:

1. compat.h: replace the posix_fadvise no-op with a real WILLNEED cache-warmer
   (overlapped ReadFile into a scratch buffer -> populates the standby page cache
   so the later synchronous pread faults from RAM). Re-arms the existing
   expert_prefetch/PILOT/next-block prefetch chain on Windows. DONTNEED stays a
   no-op (matches macOS; Windows standby-list trimming self-regulates).
   Measured: hit rate 16.4% -> 27.6%.

2. glm.c: flip PIPE (async expert-load thread pool) from default OFF to default ON
   on Windows. Dispatches expert pread onto worker threads so loads overlap the
   matmul, instead of blocking serial load-then-compute. PIPE=0 opts out.
   Measured: expert-disk 65.9s -> 54.3s (-18%).

Also adds compat_fadvise assertions to tests/test_compat_direct.c (data integrity
after cache-warmer, safe no-op on bad fd / non-WILLNEED).

mmap (CreateFileMapping/MapViewOfFile) was implemented and tested at length but
reverted: it regressed on Windows (RSS bloat from touched mapped pages collapsed
the expert cache via ullAvailPhys — a fundamental Windows-vs-Linux difference).
Full findings + the dead-end analysis recorded in issue_diskio.md.
2026-07-15 06:14:15 -04:00
monotophic 0811730845 serve mux: keep ragged decode batches off the fused Metal kernels
With METAL=1 + COLI_METAL=1, run_serve_mux (SERVE_BATCH=1) truncated
every completion to exactly 1 token: step_decode_batch passes per-row
kvs[]/positions[] with pos_base=0, but the two Metal decode fast paths
(attention_rows and the FULL-LAYER CB in layer_forward_rows) ignored
them and dispatched coli_metal_attn_decode/coli_metal_layer_decode with
the model-bound Lc/Rc and the hardcoded pos_base. The kernels' contract
is one sequence, row s at pos_base+s: ragged rows got roped at position
0 and attended over a T=1 window of the wrong cache, so greedy decode
hit a stop token on the first batched step (DONE ... STAT 1).

Gate both fast paths on !kvs. Ragged mux rows now take the CPU absorb
path, which already reads kvs[s]/positions[s]/ks->kv_start per row;
plain serve, chat/run, prefill and MTP verification (kvs==NULL) keep
the fused GPU kernels unchanged.

Verified on GLM-5.2 int4 (M5 Max): mux with COLI_METAL=1 went from
1-token DONE to full 16/16-token greedy completions, byte-identical to
the plain-serve comparator on both test prompts; CPU-only mux was
already correct (bisection); test-c and metal-test pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 05:54:50 -04:00
ZacharyZcR 0189a9f0ae tools: E8 lattice quantization for the ablation harness — int{2,3,4,8}[-gN][-e8|-e8u][-rot] (#81) 2026-07-15 17:24:16 +08:00
woolcoxm 9c24b1eceb Merge branch 'feat/expert-budget' into windows-dev 2026-07-15 03:23:25 -04:00
woolcoxm dd0d60692d experiment: miss-aware EXPERT_BUDGET — keep all cache hits, only drop misses
The original budget dropped experts blindly — even cached ones that cost
zero disk I/O. The miss-aware version pre-scans pin/ecache residency
before applying the budget:
  - ALL cache hits are kept (free to compute, no disk I/O)
  - Only misses compete for the remaining budget slots
  - Miss budget = EXPERT_BUDGET - nhits (min 0)

Results (budget=4, same prompt/config as before):
                    Original    Miss-aware
  tok/s             0.33        0.36 (+9%)
  hit rate          16.4%       38.2% (2.3x)
  prefill           8.9s        5.6s (1.6x faster)
  decode            97.4s       88.9s (9% faster)

The hit rate doubling is the key quality signal: the model now gets the
full contribution from all resident experts plus the top-4 new loads,
instead of losing some hits to the budget.
2026-07-15 02:34:58 -04:00
woolcoxm 413b370fbf experiment: EXPERT_BUDGET=N — cap distinct experts/layer, up to 1.8x faster decode
Add EXPERT_BUDGET env var that caps the number of distinct experts loaded
per layer across the batch-union. When the union exceeds the budget, keeps
only the highest-aggregate-gate-weight experts and drops the rest from
idxs[] so they're never loaded from disk.

Complementary to TOPP (per-position) — this trims the cross-position union
that multiplies under MTP/prefill. Based on MoE-Spec (arXiv 2602.16052):
'top 32 of 64 experts capture 93% of routing weight.'

Measurements (GLM-5.2 744B, 24GB RAM, cap=2, MTP=0, 32 tokens):
  Baseline (budget=0):  0.18 tok/s, 9.3% hit,  176s decode, 39s prefill
  EXPERT_BUDGET=12:     0.19 tok/s, 14.0% hit, 171s decode, 12s prefill
  EXPERT_BUDGET=6:      0.26 tok/s, 21.0% hit, 123s decode,  7s prefill
  EXPERT_BUDGET=4:      0.33 tok/s, 16.4% hit,  97s decode,  9s prefill

Budget=4 nearly doubles decode speed (+83%) and 4x's prefill speed by
halving disk reads per layer. Default OFF (EXPERT_BUDGET=0).
2026-07-15 02:34:57 -04:00
woolcoxm 71c262ce1a diskio: KV write batching + persistent handle; generators: unfuse experts
Two independent fixes validated end-to-end on fresh fixtures:

1. KV cache disk I/O (issue_diskio.md opportunities #1 + #4):
   - kv_disk_append: fopen/fclose every turn -> persistent FILE* kept open
     for the engine lifetime, lazy open on first append, closed in
     serve_ctx_free. Eliminates per-turn handle creation overhead.
   - kv_disk_append: ~157 small fwrites per position -> one contiguous
     record memcpy'd into a staging buffer then a single fwrite per
     position. The staging buffer grows on demand via realloc.
   - kv_disk_truncate: closes the persistent handle before truncating
     so the file actually shrinks on disc, then reopens lazily.
   - KVState gains disk_fp, disk_buf, disk_buf_cap fields.
   - Verified: serve-mode round-trip, write 11 tokens then reload and
     resume with no re-prefill, then append 8 more and reload to 19.

2. Expert weight unfusing in test-model generators:
   - The real GLM-5.2-FP8 checkpoint stores routed experts UNFUSED as
     per-expert 2-D tensors, each with its own _scale_inv. HF fuses
     gate+up into a single 3-D gate_up_proj for compute efficiency.
   - The converter and C engine both expect the unfused layout. The
     fused 3-D tensors were silently skipped by the converter, and the
     engine crashed with missing-tensor errors.
   - New unfuse_experts in glm_fp8_emit.py splits gate_up_proj and
     down_proj into per-expert 2-D tensors. Called after reference
     generation but before saving, in both generators, both FP8 and bf16.
   - Also fixed: make_glm_oracle.py FP8 round-trip guard used p.dim()<2
     which let 3-D fused experts through and crashed fp8_block_quantize.
     Changed to p.dim()!=2 to match the converter ndim!=2 guard.

Validated full chain on fresh fixtures:
  generator --fp8 -> 570 e4m3 tensors + 629 scale_inv, was 90 when fused
  converter --group-size 0  -> per-row int4 fmt=2, engine loads clean
  converter --group-size 128 -> grouped int4 fmt=4, 8-16x more scales,
    engine loads clean, fmt=4 auto-detected in both mmap and slab paths
  dequant error: grouped 1.14-1.22x lower than per-row vs FP8 source
2026-07-15 02:32:12 -04:00
woolcoxm e71d4fbe29 windows-dev: CUDA build script, expert budget + grouped quant research
- build_cuda.bat: one-shot nvcc DLL compilation for sm_120 (RTX 5070 Ti)
- issue_budget.md: EXPERT_BUDGET research (cap distinct experts/layer, 2x tok/s)
- issue_grouped_quant.md: root cause of int4 incoherence (per-row vs group-128 scales)
- glm_fp8_emit.py: FP8 e4m3 test weight generator for converter validation
2026-07-15 02:32:12 -04:00
woolcoxm 2c6946c478 test-models: add --fp8 emission for the FP8->int4 converter test path
Both test-model generators (make_glm_oracle.py, make_glm_bench_model.py) can now
emit weights as FP8 e4m3 + 128x128 block scale_inv, in the same layout as the real
GLM-5.2-FP8 checkpoint. This lets convert_fp8_to_int4.py exercise its FP8->int4
dequant path on a local fixture without the 379 GB download.

- New shared helper glm_fp8_emit.py: FP8 block quantize/dequantize (FBGEMM/TE
  scale=amax/448 convention) + state_dict emitter. Only exactly-2-D tensors are
  quantized; 1-D/3-D and norms/router/e_score_correction_bias are kept as f32,
  mirroring the converter's classify() + ndim!=2 guard.
- make_glm_bench_model.py: opt-in --fp8 writes model.safetensors in FP8 layout
  (config.json written explicitly since the FP8 path bypasses save_pretrained);
  manifest gains a 'format' field. Default bf16 behavior unchanged.
- make_glm_oracle.py: opt-in --fp8 round-trips quantizable weights through FP8
  before computing ref_glm.json, so the reference reflects exactly the FP8 model
  the converter ingests. Default bf16 oracle contract unchanged.

Verified end-to-end: FP8 model -> converter --indir -> int4 U8 + .qs F32 output,
bit-identical dequant between helper and converter (maxdiff 0.0).
2026-07-15 02:32:12 -04:00
woolcoxm 21eb86a0dd diskio: research on disk I/O minimization for MoE expert streaming
Audit of all disk I/O paths in the engine (expert pread, KV persistence,
config/tokenizer loads) and research into techniques used by llama.cpp,
vLLM, AirLLM, PRESERVE, HOBBIT, SolidAttention. Findings:

- Expert path is already well-batched (one coalesced ~19MB O_DIRECT pread)
- 76% of decode time is expert-disk I/O on RAM-constrained hosts
- posix_fadvise(WILLNEED) is a no-op on Windows (compat.h:107)
- I/O-to-compute ratio is 3.6x — the binding constraint
- Levers: hit-rate (cache cap), cross-layer prefetch (PILOT_REAL),
  storage (VHDX vs direct NVMe), batched decode

Ranked opportunities and source links documented in issue_diskio.md.
2026-07-15 02:32:12 -04:00
woolcoxm 69f65d5173 Windows-dev: grouped quantization + mixed precision + expert budget + download tool
Consolidates all experiment branches into one Windows-dev branch:

1. Group-scaled int4 (fmt=4, gs=128) — glm.c
   - QT struct: added gs field
   - matmul_i4_grouped: AVX2 kernel, verified to 3e-08 vs f32
   - Format detection: auto-detects from .qs scale array size
   - expert_load: both mmap and slab+pread paths handle fmt=4
   - qt_bytes: fmt=4 case added

2. Per-tensor-type mixed precision — convert_fp8_to_int4.py
   - Split classify() into sh/o/kvb/attn/dmlp sub-types
   - New args: --shared-bits, --o-bits, --kvb-bits, --attn-bits, --dmlp-bits
   - Plan: shared expert + o_proj + kv_b_proj at int8, rest grouped int4
   - Only +5.3 GB RAM vs +0 for pure int4

3. EXPERT_BUDGET (miss-aware) — glm.c
   - Caps distinct experts per layer across batch-union
   - Always keeps cache hits, only drops misses
   - Up to 1.8x faster decode on low-RAM hosts

4. Two-step shared-expert prediction (PILOT_TWO) — glm.c
   - la_predict kind==2 + pilot_prefetch integration
   - +3.1% recall over baseline PILOT

5. FP8 download tool — download_fp8.py
   - ModelScope + HuggingFace dual-source
   - Parallel shard download with stall recovery

6. Tiny model generation — make_glm_oracle.py
   - Generated and tested locally for pipeline validation
2026-07-15 02:32:12 -04:00
woolcoxm e141db047d converter: per-tensor-type mixed-precision control
Split the resident weight classification into 5 sub-types so each can
get different precision:
  sh   = shared expert (highest sensitivity, fires every token)
  o    = o_proj (reconstructs output, biggest attn tensor)
  kvb  = kv_b_proj (reconstructs KV cache on every decode)
  attn = q_a/q_b/kv_a (other attention projections)
  dmlp = dense MLP (first 3 layers)

New args: --shared-bits, --o-bits, --kvb-bits, --attn-bits, --dmlp-bits
Each defaults to ebits (backward compat). When set, the converter applies
that precision to just that tensor type.

Research-backed plan: put the 3 compounding tensors (shared expert, o_proj,
kv_b_proj) at int8 and everything else at grouped int4. Extra RAM cost:
only +5.3 GB (those tensors are small vs the 372 GB expert pool on disk).
2026-07-15 02:32:12 -04:00
ZacharyZcR d04d99e039 serve: print the KV prefix-reuse diagnostic in mux mode too — [API] KV slot line existed only in the legacy \x02PROMPT path (#153) 2026-07-15 14:31:36 +08:00
woolcoxm e2ad6c72e6 diskio: KV write batching + persistent handle; generators: unfuse experts
Two independent fixes validated end-to-end on fresh fixtures:

1. KV cache disk I/O (issue_diskio.md opportunities #1 + #4):
   - kv_disk_append: fopen/fclose every turn -> persistent FILE* kept open
     for the engine lifetime, lazy open on first append, closed in
     serve_ctx_free. Eliminates per-turn handle creation overhead.
   - kv_disk_append: ~157 small fwrites per position -> one contiguous
     record memcpy'd into a staging buffer then a single fwrite per
     position. The staging buffer grows on demand via realloc.
   - kv_disk_truncate: closes the persistent handle before truncating
     so the file actually shrinks on disc, then reopens lazily.
   - KVState gains disk_fp, disk_buf, disk_buf_cap fields.
   - Verified: serve-mode round-trip, write 11 tokens then reload and
     resume with no re-prefill, then append 8 more and reload to 19.

2. Expert weight unfusing in test-model generators:
   - The real GLM-5.2-FP8 checkpoint stores routed experts UNFUSED as
     per-expert 2-D tensors, each with its own _scale_inv. HF fuses
     gate+up into a single 3-D gate_up_proj for compute efficiency.
   - The converter and C engine both expect the unfused layout. The
     fused 3-D tensors were silently skipped by the converter, and the
     engine crashed with missing-tensor errors.
   - New unfuse_experts in glm_fp8_emit.py splits gate_up_proj and
     down_proj into per-expert 2-D tensors. Called after reference
     generation but before saving, in both generators, both FP8 and bf16.
   - Also fixed: make_glm_oracle.py FP8 round-trip guard used p.dim()<2
     which let 3-D fused experts through and crashed fp8_block_quantize.
     Changed to p.dim()!=2 to match the converter ndim!=2 guard.

Validated full chain on fresh fixtures:
  generator --fp8 -> 570 e4m3 tensors + 629 scale_inv, was 90 when fused
  converter --group-size 0  -> per-row int4 fmt=2, engine loads clean
  converter --group-size 128 -> grouped int4 fmt=4, 8-16x more scales,
    engine loads clean, fmt=4 auto-detected in both mmap and slab paths
  dequant error: grouped 1.14-1.22x lower than per-row vs FP8 source
2026-07-15 02:28:33 -04:00
Vincenzo 3fd47b7bbd Merge pull request #242 from woolcoxm/feat/grouped-quant-fmt4
Group-scaled int4 (fmt=4): one scale per 128 elements — fixes incoherent output
2026-07-15 08:25:29 +02:00
ZacharyZcR e7cb7df501 tools: one expert atlas, not two — retire expert_atlas.py, analyze.py gains --web for the dashboard 2026-07-15 14:16:36 +08:00
JustVugg ff04b320d0 security(win): load coli_cuda.dll by absolute path, never from the CWD (DLL hijack)
coli_cuda_load did LoadLibraryA("coli_cuda.dll") with a bare name. Windows'
default search order includes the current working directory (and, without
SafeDllSearchMode, other writable locations), so an attacker who plants a
coli_cuda.dll where the user launches glm.exe — or inside a downloaded model
directory the user cd's into — gets their DllMain executed at load time:
DLL hijacking -> arbitrary code execution.

Now the loader resolves the path next to glm.exe via GetModuleFileNameA and
loads that exact file with LOAD_WITH_ALTERED_SEARCH_PATH, so both the DLL and
its dependency search are anchored to the trusted install directory. Fallback
(if GetModuleFileNameA ever fails) uses LOAD_LIBRARY_SEARCH_APPLICATION_DIR |
LOAD_LIBRARY_SEARCH_SYSTEM32 — which also excludes the CWD. Cross-compiles clean
under mingw-w64; the CPU path is unaffected (this file is _WIN32-only).

Note: grammar.h gr__rule memcpy was reviewed in the same pass and is safe (len
is clamped to 63 into a name[64] buffer).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 08:16:20 +02:00
JustVugg 4268e00fa1 security: harden safetensors/JSON parsers against malicious model files
A downloaded (supply-chain) model file was fully trusted by the loader. Three
memory-safety holes, all reachable by pointing the engine at a crafted shard —
demonstrated crashing on pre-fix, now rejected fail-closed:

st.h (safetensors):
- header length `hlen` (u64 from the file) was unbounded before malloc(hlen+1):
  a crafted value overflows (malloc(0) then hdr[hlen]=0 OOB) or forces a giant
  allocation. Now bounded to the file size and a 512 MB cap; malloc NULL-checked.
- json_get() returns NULL for missing/mistyped fields, but dtype/data_offsets/
  shape were dereferenced blind (off->kids[0]) — a header omitting data_offsets
  SIGSEGV'd (verified). Now type/arity-checked before use.
- data_offsets [a0,b0] were trusted: b0<a0 gave a negative nbytes -> malloc((size_t))
  giant and an oversized memcpy into the caller's buffer in st_read_f32 (heap
  overflow); off could point outside the file. Now validated 0<=a0<=b0 and
  data_start+b0<=filesize.

json.h: j_parse_val recursed with no depth limit -> stack overflow on nested
input like [[[[...]]]]. Added J_MAX_DEPTH=1024 (headers are ~3 deep); wide-but-
flat objects like the GLM header are unaffected (depth is decremented per return).

eval_glm.py: tempfile.mktemp() -> mkstemp() — closes the TOCTOU/symlink race on
a shared tmp dir (CWE-377).

Network path (openai_server.py + serve SUBMIT parser) audited separately and is
already sound: hmac.compare_digest auth, MAX_BODY cap, resolve()+relative_to
traversal guard, list-form subprocess, bounded/validated SUBMIT header. All 62
tests pass; valid GLM/OLMoE shards load unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 08:13:20 +02:00
woolcoxm 498ab0c20e experiment: group-scaled int4 (fmt=4) — one scale per 128 elements, not per row
Root cause of gibberish output: the int4 quantization uses one F32 scale per
output row (2048 scales for a 2048x6144 matrix). The FP8 source has 128x128
block scales — 48x finer. This destroys reasoning while keeping surface fluency.

Changes:
- Converter: quant_int4_grouped() with --group-size 128 arg. Same nibble
  packing, but one scale per group of 128 elements along the input dim.
- Engine QT struct: added 'gs' field (group size, 0=per-row backward compat)
- Engine qt_from_disk: auto-detects fmt=4 when scale array is O*ceil(I/128)
  elements instead of O. Old per-row models (fmt=2) work unchanged.
- Engine matmul_i4_grouped(): AVX2 kernel that applies per-group scales.
  Accumulator resets at each group boundary: dot(x[grp],w[grp]) * scale[grp].
- Engine matmul_qt_ex: dispatches to grouped kernel for fmt=4 (always exact,
  no IDOT approximation since the point is quality)
- Engine expert_load: both mmap and slab+pread paths detect fmt=4 from
  scale array size and set gs=128
- qt_bytes: fmt=4 reports correct memory including group scales

Backward compatible: existing per-row int4 models work unchanged.
The fused gate+up pair path (matmul_i4_pair) falls back to separate
matmul_qt calls for fmt=4 — minor perf cost, correctness preserved.
2026-07-15 02:05:18 -04:00