Commit Graph

301 Commits

Author SHA1 Message Date
ZacharyZcR 786f98d471 cuda: load ragged attention entry point on Windows 2026-07-18 03:50:48 +08:00
KingIcyCreamProjects e2d39abd2d sampling: NaN-skip the mx scan too (review follow-up on #369)
Per review: the collapse starts one line before the sum — seeding
mx=lo[0] means a NaN at index 0 makes mx NaN, every (lo[i]-mx) NaN, and
the softmax is doomed at the max-finding, not the normalize. Seed mx
from -INFINITY and skip NaNs (x==x), mirroring the argmax_v change; if
nothing finite survives, fall back to mx=0 and let the post-sum guard
decide. The isfinite(s) guard is now the second line of defense rather
than the only one.

Clean logits take a byte-identical path (the extra x==x compare is
noise next to V expf calls). test_logit_nan gains the NaN-at-index-0
and all-NaN dist_build cases; test_topp's 123-case sweep still passes
on this tree, confirming no interaction with the #354 heap select.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 14:39:14 -05:00
KingIcyCreamProjects 7f70a8db5d sampling: guard against non-finite logits (was silent token-0 spew)
On the default serve path (TEMP>0, 0<NUCLEUS<1) a single NaN or +Inf in
the logits — a bad streamed expert tile, or an fp overflow in the matmul
at a low-RAM eviction boundary — poisoned softmax: g_pbuf became all-NaN,
dist_sample never satisfied cum>=u, and the fallback returned token 0. The
engine then emitted an unbroken run of token 0 with NO error. The greedy
path was equally blind: argmax_v started bv=lo[0] and `lo[i]>NaN` is always
false, so a NaN at index 0 pinned the argmax to 0.

- argmax_v: skip NaN (x==x) and seed from -inf, so it returns the max
  finite/+Inf entry instead of being NaN-pinned to 0. Covers greedy decode
  and the speculative-verify argmax path.
- dist_build: after the softmax sum, if s is non-finite or <=0, collapse
  g_pbuf to a one-hot over the finite argmax and warn once, instead of
  dividing every entry into NaN. Covers the nucleus and verify paths.

Both are O(1)/free on the happy path (one branch after the existing loop;
one extra comparison inside the existing argmax loop). Degrade + diagnose,
never silently corrupt.

test_logit_nan (wired into TEST_BINS): asserts argmax_v skips NaN/picks
+Inf, dist_build yields a finite normalized one-hot on the max finite
logit, dist_sample emits that token (not 0), and clean logits still give a
valid distribution. Fails on stock dev, passes with this change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 14:37:30 -05:00
volognamor ef97f3cf2b Windows: fix non-ASCII chat prompt corruption (ANSI codepage vs UTF-8)
coli passes the chat prompt to glm.exe through the PROMPT/COLI_PROMPT
environment variable. On Windows, plain getenv() is populated by the CRT
from the ANSI-codepage view of the environment block, not UTF-8 — so any
non-ASCII prompt text (Cyrillic, CJK, ...) is silently mangled before the
byte-level tokenizer ever sees it, even though the parent process (coli's
Python subprocess call) sets the value correctly via the wide env block.

Add compat_getenv_utf8() in compat.h: reads the variable through
GetEnvironmentVariableW and converts straight to UTF-8, bypassing the ANSI
codepage entirely. No-op passthrough to getenv() on non-Windows platforms.
coli_user_prompt() in glm.c now uses it for both COLI_PROMPT and PROMPT.

Verified: tiny-oracle self-test still 32/32 after rebuild, and a real run
against the full GLM-5.2-int4 model with a Cyrillic prompt now round-trips
correctly end to end (input echoed intact, coherent Cyrillic output),
where it previously produced replacement-character garbage on both input
and output.
2026-07-17 21:49:28 +03:00
JustVugg 8b36736e5d Merge branch 'p357' into trial357
# Conflicts:
#	c/Makefile
2026-07-17 20:47:21 +02:00
JustVugg 70a58799d6 Merge branch 'p354' into trialsamp
# Conflicts:
#	c/Makefile
2026-07-17 20:45:53 +02:00
Vincenzo 1a956b0119 Merge pull request #331 from nbeerbower/st-pread-full
st: chunked pread with EINTR retry and honest short-read errors
2026-07-17 20:42:42 +02:00
Vincenzo 3d0a28b7bf Merge pull request #350 from woolcoxm/fix/build-warnings-dev
fix(warnings): silence two -Wall/-Wextra warnings on the MinGW build
2026-07-17 20:42:25 +02:00
Vincenzo 59d74ae435 Merge pull request #340 from tonuonu/fix/fs-detect-9p-statfs
glm.c: detect 9p via statfs f_type, not the /mnt/ path prefix
2026-07-17 20:42:22 +02:00
Vincenzo 1af9435760 Merge pull request #328 from mohamedmastouri2000-boop/fix/olmoe-ebits
convert_olmoe.py: wire --ebits through to quantization (fixes #323)
2026-07-17 20:42:18 +02:00
ZacharyZcR 1223f9ca2e cuda: batch ragged attention across independent streams 2026-07-18 02:01:12 +08:00
Vincenzo 1552db2323 Merge pull request #347 from ZacharyZcR/tools/e8-rate-scaled
tools: rate-scale the E8 lattice ball by bit-width (#81) — int3-e8 is now real int3
2026-07-17 19:17:33 +02:00
JustVugg 211d4488c3 win32: stop erasing the real ReadFile error — pread failures name their WinErr (#307)
compat_pread mapped every non-EOF ReadFile failure to EIO, so the field report
was always the same three words: "Input/output error". #307 then burned three
rounds of guessing across three people — storage? contention? alignment? —
because the actual GetLastError code never appeared anywhere.

compat_pread now stashes the code in a per-thread slot and pread_full appends
it to the failure line:

    pread qs: Input/output error (off 780592, 0/8192 bytes, WinErr=1450)

That one number is the difference between "insufficient system resources"
(1450 -> memory pressure), "sharing violation" (32 -> AV interference),
"device not ready" (21 -> the T: drive itself) and a dozen other distinct
diagnoses. Diagnostic only: no behavior change, no retry policy — that
discussion lives in #361 and should be settled AFTER the first report tells
us which error we are actually retrying.

Linux path untouched byte-for-byte; the Windows compile is certified by the
check.yml MSYS2 job.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 18:49:28 +02:00
woolcoxm 5d16368817 sampling: partial top-keep select in attention_rows DSA — O(nk) quickselect, not O(nk log nk) qsort (#356)
The DSA lightning indexer selects the top-index_topk (2048) context keys to
attend to by finding the threshold = keep-th largest attention score. It
previously full-qsorted all nk scores per layer per token (O(nk log nk)) just
to read one pivot value, then scanned the original array in position order to
build the kept set.

Replace the qsort with partial_select_desc (Hoare quickselect, median-of-three,
descending): O(nk) average to partition the keep largest into a[0..k), then the
threshold is min of that block. The two position-order scans (>thr then ==thr)
are UNCHANGED, so the kept-position set is bit-identical -- a stronger contract
than #335's sampling heap (which was multiset-only because the heap was unstable
and changed accumulation order). The quickselect pivot IS by definition the
keep-th largest, so the new threshold equals old tmp[keep-1] exactly.

Measured (bench_dsa_select, keep=2048, median of 2000 reps):
  nk=2049:   119us -> 5.7us   (21x)
  nk=8192:   626us -> 43us    (15x)
  nk=32768:  2.8ms -> 0.28ms  (10x)
  nk=65536:  6.6ms -> 0.47ms  (14x)
The gap widens with context (linear vs n-log-n). DSA only fires past index_topk,
so this is precisely the long-conversation regime where decode latency matters.

Adds test_dsa_select (in TEST_BINS): 129 cases asserting element-wise identical
kept-set vs an independent qsort reference across shapes (random, peaked,
sorted, reverse-sorted, tie-plateau, all-equal) and edges (keep==1, keep==nk).
Also directly checks the partition invariant.

Adds bench_dsa_select (on-demand, NOT a gate): reproduces the table above.
2026-07-17 09:54:40 -04:00
woolcoxm 37c96ee07a bench: add bench_topp -- head-to-head old-qsort vs new-heap timing on V=151936 (#335)
test_topp proves correctness; bench_topp measures the latency claim. It re-implements
the OLD dist_build (full-vocab qsort) inline on a private buffer and times it against
the REAL new dist_build over identical inputs in one process: V=151936, temp=0.7,
3 shapes (realistic / uniform / plateau) x 4 nuc values (0.5/0.9/0.95/0.99), 2000
timed reps each, median reported. Deliberately NOT in TEST_BINS -- it's a microbench,
not a gate. Build on demand: make tests/bench_topp && ./tests/bench_topp
2026-07-17 09:22:26 -04:00
woolcoxm ba00889fb9 sampling: partial top-p select in dist_build — O(V) heapify + k pops, not O(V log V) qsort (#335)
dist_build() sorted the entire 151936-entry vocab by probability (qsort) on every
sampled token whenever 0 < g_nuc < 1 — the serve default — and again per draft
position under rejection sampling. Measured cost: 5.6-8.0 ms/call; the actual work
is finding the few-hundred-token head whose cumulative mass reaches g_nuc.

Replace the full qsort + linear scan with a max-heap partial select:
  - Floyd heapify g_pidx over V by descending g_pbuf prob  (O(V), cache-friendly)
  - pop winners to the array's high end until cum >= g_nuc  (k * O(log V))
  - the remaining heap prefix IS the tail -> zero it, renormalize the head

Winners land in g_pidx[out..V-1] in descending order, so s2 accumulates in the
same order as before -> head is unchanged on tie-free shapes (ties were already
unspecified under the unstable qsort). All four dist_build/dist_sample contract
properties hold: g_pbuf stays id-indexed, g_pidx stays internal, the tail is
fully zeroed, the head renormalizes to 1.

No API change, no caller change, no new globals.

c/tests/test_topp.c (new): drives the REAL dist_build via the include-glm.c
pattern against an independent double-precision reimplementation of the OLD
algorithm. 123 cases: 6 sizes (1..1519) x 5 shapes (uniform/peaked/geometric/
plateau/sharptail) x 4 nuc values, plus the g_nuc>=1 guard-off paths and V=1.
Tie-free shapes compare head values within 1e-6 rel (float vs double renorm
noise); tie shapes compare value-multisets. No scratch files -> builds clean on
Windows MinGW without the unmerged mkdtemp shim (#352).
2026-07-17 09:01:23 -04:00
woolcoxm 5e2be61a2c fix(test): make test_stops build on Windows (mkdtemp compat shim)
test_stops.c uses POSIX mkdtemp() to make a scratch dir in the CWD, but
MinGW-w64 does not declare mkdtemp, so the test failed to compile on the
Windows job - and only there:

  tests/test_stops.c:73:9: error: implicit declaration of function 'mkdtemp';
     did you mean 'mktemp'? [-Wimplicit-function-declaration]
  make: *** [Makefile:318: tests/test_stops.exe] Error 1

That halted `make test` at the C-test stage on Windows (test_uring is
correctly Linux-only, so test_stops was the only blocker).

Added a compat_mkdtemp shim to compat.h, following the file's existing
convention (every platform difference lives there; the .c stays clean):
_mktemp fills the trailing X's in place (same contract as mkdtemp), then
_mkdir creates the directory. Also added <direct.h> for _mkdir. On Linux
compat.h is a complete no-op, so POSIX mkdtemp is untouched there.

Verified: test_stops builds clean on MinGW and all 5 sub-cases pass:
  tokenizer special-flag parsing, config/generation_config eos union,
  no-generation_config fallback, both-configs-mutilated tokenizer sweep,
  and the T=NULL validation path.
2026-07-17 08:27:51 -04:00
woolcoxm 411f237f94 fix(warnings): silence two -Wall/-Wextra warnings on the MinGW build
A clean 'make glm' on MinGW emitted exactly two warnings, both real:

1. compat.h:240 - ignoring pragma comment [-Wunknown-pragmas]
   #pragma comment(lib, "psapi.lib") is an MSVC directive; MinGW/GCC
   warns about it. Guarded with ifdef _MSC_VER - MinGW links psapi via
   -lpsapi (already in the Makefile), MSVC keeps the pragma.

2. glm.c:1210 - g_numa_nodes defined but not used [-Wunused-variable]
   g_numa_nodes is only read/written inside ifdef __linux__ blocks, so on
   every non-Linux build it is a static that is never used. Moved the
   definition under the same __linux__ guard; nothing references it off-Linux.

Verified: rm -f *.o glm.exe && make glm -> 0 warnings, 0 errors.
2026-07-17 08:26:32 -04:00
JustVugg d5327e2252 omp: cap LLVM libomp idle spin — a parked engine must not burn 3000% CPU (#341)
The hot-thread tuning sets OMP_WAIT_POLICY=active and caps libgomp's spin
with GOMP_SPINCOUNT=200000. LLVM libomp reads NEITHER: under active policy it
sets KMP_BLOCKTIME=infinite, so once the answer ends and the engine parks on
stdin, the whole OMP team spins forever — ~100% per thread, the 3000% figure
reported on FreeBSD 15.1 in #341 (clang/libomp is the default toolchain
neighborhood there; the same applies to macOS builds).

KMP_BLOCKTIME=200 keeps the team hot for 200 ms after each parallel region —
plenty to bridge back-to-back per-expert matmuls — and then sleeps it.
setenv with overwrite=0, so a user-provided KMP_BLOCKTIME still wins, and
libgomp builds ignore the variable entirely: on GCC toolchains this commit
is a no-op by construction.

Not reproduced locally (this box is gcc/libgomp and idles clean); shipped as
the mechanism-matching candidate fix with a request on the issue for the
reporter to verify on dev and to name their OpenMP runtime (ldd | grep omp).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 13:53:20 +02:00
JustVugg 7eb239328d coli: serve pidfile + coli stop — one command to shut engine and server down
No more manual pkill: cmd_serve writes a pidfile, cmd_stop finds the server
(pidfile, then /proc cmdline) and its engine (comm glm/exe/olmoe with SERVE=1
in environ — the engine re-execs for OMP tuning so its comm is 'exe', which is
why every 'pkill -x glm' in history killed nothing). SIGTERM, wait, SIGKILL.
--dry-run lists targets without acting. First real use took down a live
serve+engine pair cleanly and released 16 GB.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 41a872c331a2a0a8655699e0171c68dd2bcda186)
2026-07-17 13:27:02 +02:00
JustVugg 8bf4cb9a98 coli: chat attaches to a running serve — the engine survives the chat (local)
The cold-chat cost, measured on this box: every `coli chat` spawns a private
engine (34-136 s of resident load) and starts with an empty expert cache (hit
4% cold vs 55% warm). Quitting throws both away.

`coli chat` now probes localhost:8000 first (~1 ms when nothing listens) and,
if a `coli serve` answers, runs the REPL over plain OpenAI SSE against it:
stdlib urllib only, engine byte-protocol untouched. --attach [URL] forces it,
--no-attach restores a private engine. reasoning_content keepalive pings are
filtered; :reset starts a new conversation client-side (the server's KV slots
reuse prefixes per conversation on their own).

Verified against a mock SSE server (pings ignored, markdown rendered, :reset,
clean exit) and against the real 744B model: two consecutive sessions, second
attach instant with zero reload — the engine stayed resident at 15.7 GB across
both. Honest limit: warmth carry-over BETWEEN different conversations is small
here because cap=3 slots/layer is a short memory; the structural wins are the
load never being repaid and same-conversation continuation.

LOCAL ONLY for now — not pushed, per the current working rule.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit aa406ccab8a4501b924bc3a9f4725dd1d18a685d)
2026-07-17 13:26:59 +02:00
ZacharyZcR 3f239dbb92 tools: rate-scale the E8 lattice ball by bit-width — int3-e8 is now a real int3 codebook, not a fixed 2-bit ball 2026-07-17 18:39:53 +08:00
Tonu Samuel 81a56777df glm.c: detect 9p by statfs f_type, not the "/mnt/" path prefix
The slow-filesystem warning fired for any model path under /mnt/, which
false-positives on native-Linux mounts (ZFS/ext4/xfs/NFS) that commonly
live there. Check the actual filesystem type via statfs() against the 9p
superblock magic (0x01021997) instead, so it only warns for a genuine
WSL 9p mount. Linux-only (statfs); no behavior change on other platforms.
2026-07-17 08:00:26 +03:00
Nicholas Beerbower ed8dab4da4 test_st_pread: Windows-compatible — relative tmpdir, fork subtest gated POSIX
Same fix pattern as test_stops: mkdtemp with a CWD-relative template
(MinGW resolves Windows paths; /tmp is not one), and the fork/pipe/
truncate-based truncation subtest compiles only where those exist —
Windows still runs the cross-platform chunk-loop content check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 16:32:43 -04:00
Nicholas Beerbower 6507817d9f st: chunked pread with EINTR retry and honest short-read errors
Two latent bugs in every st.h reader, both hit in the field:

 - a single pread caps at ~2^31 bytes on Linux, so any tensor past
   2.1 GB (bf16 embed/unembed tensors of large models qualify) came
   back silently truncated with perror printing '... : Success'
   (errno untouched by a short read) — the same misleading-error
   symptom glm.c fixed for its own reads in #236;
 - no EINTR retry.

st_pread_full loops in ST_PREAD_CHUNK pieces (1 GB default, override
for tests), retries EINTR, and reports offset + progress on failure.
All five read sites converted; behavior on well-formed files is
byte-identical (GLM oracle re-verified on this branch: 32/32).

tests/test_st_pread builds with -DST_PREAD_CHUNK=7 so a 96-byte tensor
exercises the multi-chunk loop, and forks a child against a shard
truncated after st_init (init's static bounds check means the pread
path only fires when a file shrinks under a live handle) asserting
exit(1) with a 'short read' message and no 'Success'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 15:20:43 -04:00
Mohamed Mastouri 80f886cd22 convert_olmoe.py: wire --ebits through to quantization (#323) 2026-07-16 21:59:16 +03:00
Vincenzo cead3d8bf3 Merge pull request #324 from dawnfield-institute/feat/olmoe-ppl-mode
olmoe: PPL=1 teacher-forced perplexity mode (loss meter for throughput experiments)
2026-07-16 20:47:37 +02:00
JustVugg e7188df16a tests: test_stops must not assume /tmp exists (fixes the windows job)
My test_stops.c called mkdtemp("/tmp/coli_stops_XXXXXX"). These binaries are
built by MinGW into NATIVE Windows .exe files, which resolve Windows paths —
"/tmp" is not one, so mkdtemp failed ENOENT and `make check` went red on the
windows job the moment #143 gave us cross-platform CI.

Now a relative template in the CWD, which is what test_compat_direct.c already
does (`#define TMPF "test_direct.tmp"`). test_uring.c uses /tmp but is
Linux-only by construction (Makefile guards it behind $(LINUX)); I copied the
wrong neighbour.

Worth being precise about what this was, because the red job looked scarier
than it was: Windows is fine. `make check` built the engine, ran the whole C
suite, and passed everything else — test_i4_grouped, kv_alloc, compat_direct —
before tripping on my temp path. The CI caught a bug in the test, not in the
port. That is #143 earning its keep on day one.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 20:08:02 +02:00
Vincenzo ee5d273cd4 Merge pull request #232 from nbeerbower/profiling-upstream
Profiling: PROF=1 opt-in performance profile + live per-turn Profiling page in the web dashboard
2026-07-16 19:58:31 +02:00
JustVugg 98f7c88ca9 numa: parenthesise the page-align arithmetic (#313)
`(uintptr_t)p+n+4095 & ~(uintptr_t)4095` is CORRECT — `+` binds tighter than
`&` in C, so the rounding happens before the mask, which is what was meant.
But gcc -Wall warns (`suggest parentheses around '+' in operand of '&'`), and
this repo builds clean at -Wall -Wextra. Warning-free is a property worth more
than the two characters it costs to keep: it is what makes a NEW warning
visible instead of scrolling past in a wall of noise.

No behaviour change; the emitted code is identical.

Co-Authored-By: ZacharyZcR <#313>

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 19:50:20 +02:00
Vincenzo 3cec90cd45 Merge pull request #313 from ZacharyZcR/feat/numa-expert-interleave
glm: COLI_NUMA=1 — interleave resident weights across NUMA nodes (+13% decode on 2-socket, raw mbind, zero deps) (#82)
2026-07-16 19:49:17 +02:00
Vincenzo 2dca36a8a0 Merge pull request #314 from mohamedmastouri2000-boop/fix/windows-cuda-dll-build
Add build-config stamp: flag changes (e.g. CUDA_DLL=1) force a relink instead of a silent stale binary
2026-07-16 19:49:14 +02:00
Vincenzo 644e0d16c6 Merge pull request #302 from alekseysorokin68/main
Fix: set stdout to O_BINARY on Windows to fix READY sentinel
2026-07-16 19:49:11 +02:00
Vincenzo b04c2ca50e Merge pull request #322 from woolcoxm/fix/oracle-transformers-pin
oracle: hard-fail on transformers < 5.11.0 (interleaved-RoPE floor, #281)
2026-07-16 19:46:20 +02:00
Vincenzo 95b059b15b Merge pull request #318 from tt1203/fix/oracle-mkdir-glm-tiny
fix: make_glm_oracle.py creates glm_tiny/ before saving (fresh-checkout failure)
2026-07-16 19:46:13 +02:00
Peter Groom c5f2027fa1 olmoe: PPL=1 teacher-forced perplexity mode (loss meter for throughput experiments)
Feeds reference tokens through the normal step() decode path and reports
NLL/ppl + hit rate + tok/s + RSS. Inert unless PPL=1; default path
untouched (12/12 vs ref.json verified with patch applied). Cross-checked
vs HF transformers bf16 on identical token ids: 12.11 vs 12.25 ppl (#108).
2026-07-16 13:10:50 -03:00
mohamedmastouri2000-boop c70d94368e Merge branch 'main' into fix/windows-cuda-dll-build 2026-07-16 19:05:43 +03:00
woolcoxm fa821a15a2 oracle: hard-fail on transformers < 5.11.0 (interleaved-RoPE floor, #281)
GLM-5.2 MLA uses interleaved (DeepSeek-style) RoPE, which the C engine
implements. transformers < 5.11.0 applied split-half (Llama-style) RoPE in
GlmMoeDsa* instead; an oracle built on those versions silently drifts and the
engine scores 25/32 instead of the documented 32/32 (#281). Weights come out
identical across versions -- only the forward pass differs -- so a too-old
transformers produces an invalid ref_glm.json with no warning.

Add a version gate at the top of make_glm_oracle.py: hard sys.exit with an
actionable message citing the issue and the upgrade command. Reads the version
from importlib.metadata (authoritative installed-dist version) rather than the
mutable transformers.__version__ attribute -- the latter gets reset by the lazy
model-class import (from transformers import GlmMoeDsaForCausalLM), so reading
it after that import is unreliable. The gate runs before the heavy import and
falls back to the attribute only if the dist metadata lookup fails (editable/
src installs).

Validated end-to-end on transformers 5.13.1: script runs, ref_glm.json and
model.safetensors are byte-identical to the shipped versions, engine scores
32/32. With the floor raised to (5,14) the gate blocks with the expected
message.
2026-07-16 12:03:15 -04:00
JustVugg 6de32c55f6 Don't trust a converted checkpoint's config for the stop set
The engine armed its stop tokens from config.json's eos_token_id and nothing
else. That trusts metadata written by third-party conversion tooling, which is
a thing we already know goes wrong: the README documents a mirror shipping
int4 MTP heads that silently give 0% draft acceptance. GLM-5.2 declares THREE
eos ids (<|endoftext|>, <|user|>, <|observation|>); a converter that rewrites
config.json with a reduced list leaves the engine stopping on fewer tokens
than the model emits, and the missed ones get detokenized and printed into the
chat as literal text while generation runs past the end of the turn.

Two independent defenses:

  - eos_token_id is now unioned with generation_config.json, which is
    HuggingFace's authority for generation (config.json often carries a
    partial legacy copy). An extra stop is harmless; a missing one is not.

  - every added-token the TOKENIZER marks "special":true is armed as a stop,
    whatever the configs say. Those are control tokens (<|user|>, <|assistant|>,
    <sop>, [gMASK], the image/video/audio markers) and are never legitimate
    content in a reply -- GLM itself lists three of them as official eos.
    <think>/<tool_call>/<arg_key> are "special":false and are deliberately NOT
    swept up: they are real output. tok.h was parsing added_tokens but throwing
    the "special" flag away, so the distinction wasn't available to anyone.

On the real per-row checkpoint this takes the armed set from 3 to 18:
  [stop] 18 stop tokens: 154820 154827 154829 154821 ... (15 from the
  tokenizer's special set)

Honesty about scope: this is hygiene for a class of bug, NOT a fix for the
trailing-junk report on #298 that prompted it. I hypothesised @woolcoxm's g64
checkpoint had lost eos ids in conversion; he checked, and it hadn't -- his
config arms all three correctly. The emit path is also innocent: is_stop() is
checked BEFORE emit() at every one of the four call sites (4215, 4256, 4908,
4987), so a correctly-armed stop cannot be printed. His trailing junk is still
unexplained and is more likely quantization noise. What this commit buys is
that a checkpoint we don't control cannot leak control tokens into a reply,
which was true before and is not now.

tests/test_stops.c covers both defenses: the union, a missing
generation_config.json, BOTH configs mutilated (the tokenizer still stops all
five control tokens while leaving <think> alone), and T=NULL (the validation
path keeps config-only behaviour).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 17:59:34 +02:00
JustVugg ac7103fe9c tests: extend the fmt=4 oracle to the fused gate+up kernel (#298)
@woolcoxm's matmul_i4_grouped_pair reads x once instead of twice for the
gate+up pair. Verified here against his branch (86e91b1 merged onto dev):
correct to ~2e-8 relative vs the double reference, and BIT-EXACT against two
separate matmul_i4_grouped calls on aligned shapes -- which is the shape the
real g64 checkpoints have (I = 2048 / 6144, gs = 64). His kernel is good.

Guarded behind COLI_HAVE_GROUPED_PAIR since the function only exists on that
branch; add -DCOLI_HAVE_GROUPED_PAIR to the test's Makefile rule when #298
lands and the pair cases activate.

The checks are deliberately asymmetric, and the reason is worth recording.
Bit-exactness is asserted ONLY when I % gs == 0: there every group is covered
by the AVX2 body, whose accumulation order matches the unfused kernel, so any
difference is a real bug. With a partial last group the tail falls to scalar
code and the compiler may contract/reassociate the fused body differently,
producing ~1e-7 differences -- rounding, not logic. My first version demanded
bit-exactness everywhere and duly "found" a bug in his kernel that did not
exist; the tell was that only `up` differed and never `gate`, which is FP luck
rather than a code path. Correctness is checked everywhere against the double
reference; identity only where identity is actually implied.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 17:27:06 +02:00
JustVugg 2a5961a01b tests: exactness oracle for the grouped-int4 kernel (fmt=4)
matmul_i4_grouped is the reference the CUDA fmt=4 port (#298) is expected to
reproduce, and it had no test of its own. @woolcoxm is currently debugging a
CUDA backend against an oracle nobody had verified, which is two moving
targets at once -- and he can't cross-check on CPU, since a 5-prompt run
takes 8 hours on the 744B model.

This checks matmul_i4_grouped against a plain-C reference that dequantizes
nibble -> (v-8)*scale[i/gs] and accumulates in double, over 11 shapes: I a
clean multiple of gs, a partial last group (the glen clamp), odd I (the
scalar nibble tail), gs > I, gs=16/64/128, S>1, and the nibble extremes
0x00/0xFF -- which decode to -8/+7 because the format is offset-encoded, not
two's complement. Reading that backwards turns 15 into -1 and looks like
data-dependent noise rather than a bug.

All 11 shapes match to ~1e-8 relative, so the CPU kernel is exact and can be
trusted as the reference.

One note on the tolerance, because the first draft of this test got it wrong
and "found" a bug that wasn't there: the error is compared against the sum of
|terms|, not against |result|. A dot product of signed terms can land near
zero through cancellation, and then a 1e-6 absolute error -- ordinary f32
accumulator precision -- reads as a 1e-3 relative one. A wrong scale index or
a wrong group boundary shifts the result by a fraction of the terms, so it is
still caught at 1e-6.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 16:37:39 +02:00
Egon Ruiter c769e04d13 fix(prefetch): address fourth round of copilot review comments
- Fix queue flush race: clear is_queued under mutex only; never move
  pilot_w backwards (would break r<=w ring-buffer invariant). Worker
  skips stale entries via new is_queued guard at start of pilot_realload
- Add early-exit in pilot_realload when is_queued==0 (entry flushed
  between enqueue and worker pickup), preventing unnecessary loads
- Fix misleading EMA struct comment: momentum_logits is used only by
  the PILOT prefetcher, not blended into actual MoE routing decisions
- Fix Slot pinned comment: 'never evicted' was too strong; clarify that
  pinned slots may be displaced under extreme all-pinned cache pressure
- Use st_read_f32() for scale tensor (.qs) instead of st_read_raw() to
  handle potential future BF16/F16 dtype changes robustly
2026-07-16 15:58:38 +02:00
tt1203 a3e942516c fix(oracle): create glm_tiny/ before saving so a fresh checkout works
make_glm_oracle.py wrote glm_tiny/model.safetensors and glm_tiny/config.json
without creating the directory first. safetensors.save_file writes a temp file
inside the target dir before the atomic rename, so on a clean checkout (no
pre-existing glm_tiny/) it aborts with an opaque error:

    SafetensorError: Error while serializing: I/O error:
    The system cannot find the path specified. (os error 3)
    at path ".../glm_tiny/.tmpXXXXXX"

The directory only ever existed because it was left over from a previous run,
so the first-ever `python tools/make_glm_oracle.py` fails for every new user
following the README's verify step.

Create glm_tiny/ with Path.mkdir(parents=True, exist_ok=True) before the save
branch — covers the fp8 path, the bf16 path, and config.json. Path is already
imported; no new dependency, no change to the CPU build.
2026-07-16 14:52:16 +01:00
Egon Ruiter b6bae91b66 fix(prefetch): address third round of copilot review comments
- Fix LRU fallback: when all evictable slots are in-flight, find oldest
  non-in-flight slot (pinned ok) before falling back to slot 0
- Fix pin_hot_experts: guard enqueue behind g_pilot>0, call
  ensure_pilot_worker_started(), and set is_queued flag to prevent
  duplicate in-flight loads from pilot_prefetch()
- Fix token counting: increment token_count/freq_token_count by S (batch
  size) instead of 1 so prefill tokens are counted accurately and warmup
  threshold triggers at the right time
2026-07-16 15:46:26 +02:00
Egon Ruiter 2d8d2951ee fix(prefetch): address second round of copilot review comments
- Fix ENV VARS header: document PILOT=0-3, SMOOTH, CONF_LIMIT; remove stale REBAL entry
- Fix per-layer EMA: apply routing momentum to all layers (not just layer 0) with correct offset
- Fix in-flight slot race in expert_get: LRU eviction now skips slots with eid==-1 (being loaded)
- Fix in-flight slot race in pilot_realload: same fix, prevents concurrent writes into active slot
- Fix idx[] buffer overflow: clamp max_cand to 128 before E in pilot_prefetch
2026-07-16 15:37:20 +02:00
Claude 36394bc317 Profiling: fix the phase accounting so 'Other' stops swallowing the disk stall
The dashboard's wall-time stack assumed ewait = felt I/O stall and
edisk = overlapped read service, but the engine measured neither:
t_ewait was declared and reported yet never incremented (the blue
'I/O wait' segment always read 0.00s), while t_edisk accumulated on
the COMPUTE thread — blocking OMP loads, PIPE dispatch, pipe_wait
spins — i.e. exactly the felt stall. The UI excluded it from the
stack as 'overlapped', so the whole disk stall landed in 'Other'
(e.g. 46.9s Other vs 45.3s disk on a 54.1s turn).

Now the semantics match the contract:
- t_ewait accumulates the compute-thread stall (blocking loads,
  PIPE dispatch, pipe_wait spins) and feeds the in-stack I/O wait.
- Disk service is real: expert_load is timed on whichever thread
  runs the read (PIPE workers, OMP loaders, pilot) into an atomic
  ns counter (g_edisk_ns) — thread-seconds, overlapped with compute.
- prof_report's I/O share and verdict use the felt wait only, and
  the expert I/O line prints 'read service / felt wait' so the
  wait:service gap shows how well PIPE is hiding the reads.

PROF wire format, /profile JSON and the web UI are unchanged —
the dashboard's existing assumptions are simply true now.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dd4hHD2Dsvgr7VAwEVFqWK
2026-07-16 09:00:31 -04:00
Claude 6afffbcbf2 Profiling page: per-turn phase timings, live in the web dashboard
The engine already tracks where each turn's wall time goes (expert-disk
service, I/O wait, expert matmul, attention, lm_head) — it just only spoke
at exit or under PROF=1. Stream it instead:

- glm.c: mux serve emits a per-turn "PROF" protocol line next to TIERS/HITS
  (window deltas per request, same convention as the STAT hit%); the phase
  window base is now always captured (a few loads per request).
- openai_server.py: parses PROF into a 120-turn rolling window and serves it
  at /profile (read-only, same trust level as /health).
- web: new Profiling tab — stat tiles (tok/s, wall, tokens/forward, disk
  service), wall-time composition bars for the last turn and the window,
  per-turn throughput and stacked phase columns with hover readouts, and a
  table of recent turns. Disk service is shown apart from the stack: it
  overlaps with compute, so only the I/O wait the compute thread felt counts
  inside wall time. Phase colours are a CVD-validated set with gaps + legend
  + table so identity never rides on colour alone.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WhTmF8yvEBgSkUKSVfZF7P
2026-07-16 08:56:12 -04:00
Claude 63a6824881 PROF=1: per-request reports in mux serve mode too (the web-dashboard path)
run_serve_mux (SERVE_BATCH, used by openai_server.py / coli web) completes
requests in mux_done, so the run_serve per-turn hook never fired there.
Snapshot the window where hits0 is taken, record batched-forward latency
around step_decode_batch, report on stderr at DONE. With KV_SLOTS>1 the
window shares batched forwards across slots — same convention as the
existing STAT hit%.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TCBNCciBaHea41QmidLUMn
2026-07-16 08:56:12 -04:00
Claude 09bf17f001 PROF=1: opt-in performance profile (latency percentiles, expert I/O, tuning verdict)
Answers 'where does it slow down on THIS machine with THIS config' so users
can tune RAM_GB/PIPE/DIRECT/PIN for their hardware without folklore:

- startup header: CPU/cores/RAM/backend + effective knobs (cache cap, pin,
  DRAFT/PIPE/DIRECT/MMAP/IDOT/DSA/PILOT/CACHE_ROUTE) — every saved log is
  self-describing when comparing runs across configs or machines
- per-forward decode latency ring (32k) -> p50/p90/p99/max, plus a tail
  diagnosis when p99 >> p50 (cold-cache expert loads)
- expert I/O accounting at the pread/mmap-touch sites: GB fetched, MB/token,
  GB/s, hit rate, loads/token, pinned/LRU tier fill
- phase shares of wall time and a plain-language verdict naming the knob
  most likely to move tok/s (I/O-bound vs compute-bound vs attention-bound)
- reports after REPLAY / PROMPT / oracle runs (stdout) and per turn in serve
  mode (stderr; stdout stays the framed protocol)

Additive only: with PROF unset every mode's output is byte-identical.
hwinfo_emit's /proc probe is factored into hw_probe() and shared.
Validated end-to-end on a tiny-random unquantized fixture (REPLAY, PROF on/
off, RAM_GB squeeze flips hit 96.9%->26.6% and the verdict follows);
make check clean, 0 warnings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TCBNCciBaHea41QmidLUMn
2026-07-16 08:56:12 -04:00
Egon Ruiter 1ac2e7b487 fix(prefetch): address copilot code quality reviews on safety and concurrency 2026-07-16 14:36:53 +02:00