Commit Graph

150 Commits

Author SHA1 Message Date
JustVugg e5b1e19022 Merge branch 'p366' into trialsafe
# Conflicts:
#	.github/workflows/ci.yml
2026-07-19 01:26:51 +02:00
Vincenzo Fornaro 3fed654473 Merge pull request #372 from ZacharyZcR/feat/paged-ragged-kv
cuda: keep paged ragged KV resident across decode steps
2026-07-18 12:48:07 +02:00
JustVugg 4fcbea9056 Merge branch 'p370' into t370
# Conflicts:
#	c/Makefile
#	c/glm.c
2026-07-18 12:41:07 +02:00
ZacharyZcR 6d3a6168e4 cuda: keep ragged KV resident across decode steps 2026-07-18 18:34:46 +08:00
Vincenzo Fornaro 605b0e28f9 Merge pull request #365 from ZacharyZcR/feat/ragged-batch-attention-upstream
cuda: batch ragged attention across independent streams
2026-07-18 02:04:43 +02:00
JustVugg b09273f73d serve: MTP status message tells the truth in multiplexed mode (#358)
`coli serve` runs the engine with SERVE_BATCH=1 (openai_server.py), which
selects run_serve_mux, which sets g_draft=0 — speculation isn't ragged-safe
across the multi-slot batch. But the "[MTP] active: native speculative decoding
(draft=N)" line is printed in main() BEFORE the serve path is chosen, so
`coli serve DRAFT=8` announced draft=8 and then silently disabled it.
@LordMZTE reported the misleading message.

The load line and the [MTP] stderr line now detect the mux case (SERVE +
SERVE_BATCH) and report it truthfully: "MTP DISABLED (multiplexed serve)" with
draft=0, plus a one-line explanation that single-client interactive use
(`coli chat`, which spawns run_serve without SERVE_BATCH) keeps MTP. No more
draft=8 claim on a path that runs draft=0.

This is the honesty half of #358. The feature half — MTP inside the HTTP
server for a single client — is a real enhancement but needs engine work: the
mux decode kernel would have to run the speculative path when exactly one KV
slot is active (S=1 is not actually ragged), and it needs a server round-trip
test. Tracked separately; the message no longer lies in the meantime.

Reported-by: LordMZTE <#358>

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 01:57:17 +02:00
JustVugg 679c0742b2 telemetry: split expert hits into pin-tier vs LRU ecache (#336)
The expert lookup counted a hit identically whether the pinned hot-store or the
LRU ecache served it — both Phase C branches bumped the single `m->hits`. With a
warm pin profile the pin tier absorbs most hot experts, so the ecache could be
serving anywhere from ~0% to most hits and the logs couldn't tell which. That
made every cache-policy question unanswerable, including #223's "at what cap
does the eviction policy start to win?" (a flat A/B can mean "pin absorbed
everything" or "genuine floor" and the lumped counter can't distinguish them).

Two counters `hit_pin`/`hit_ecache` bumped at the two lookup branches (the
existing `m->hits++` stays, so all existing math is unchanged), snapshotted in
ProfBase and reset alongside `hits`. Surfaced in the human-readable summaries:
  decode:  expert hit rate 31.3% (pin 22.1% + lru 9.2%)
  [PROF]:  hit 31.3% (X pin + Y lru / Z load)
  tiny:    hit rate 88.1% (0 pin + 74 lru / 10 miss)
The serve-mux STAT protocol line is untouched (openai_server.py parses it
positionally).

Invariant hit_pin + hit_ecache == hits holds by construction — exactly two
sites bump hits and each also bumps one split counter, nothing else mutates
hits. Verified numerically on the tiny oracle: 0 pin + 74 lru = 74 hits, 88.1%.
Zero cost: two increments on paths that already increment.

Reported-by: KingIcyCreamProjects <#336>

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 01:43:50 +02:00
JustVugg 4fb5b4e975 sampling: survive non-finite logits instead of emitting token 0 forever (#369)
A single NaN or +Inf logit silently broke the default sampling path. +Inf
became `mx`, then `expf((Inf-mx))`/`expf((NaN-mx))` is NaN, the softmax sum went
NaN, every probability went NaN — and dist_sample's fallback loop
`if(g_pbuf[i]>0)` is false for NaN at every index, so it returned 0. Every
subsequent token: 0. No error, no warning. @KingIcyCreamProjects found it.

dist_build now takes `mx` over finite logits only, gives a non-finite logit
probability 0, and when the distribution is unusable (no finite logit, or a
non-finite/zero sum) collapses to a delta on the finite argmax and warns ONCE
on stderr — degraded, but a valid token and a visible cause, never a silent
stream of zeros. The finite argmax uses the index found during the mx pass
(robust even when lo[0] itself is NaN, where argmax_v would wrongly return 0).

tests/test_sample_nan.c: healthy logits still sample correctly; NaN/+Inf
injected at lo[0], the middle, and the last position all pick the finite
argmax; an all-non-finite vocab leaves no NaN in the buffer and doesn't crash.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 01:10:33 +02:00
Vincenzo 6ce6cc4380 Merge pull request #374 from ZacharyZcR/perf/p0-execution-profile
profile CPU/GPU tier execution and residual P2P
2026-07-18 00:57:35 +02:00
ZacharyZcR 570d738ed5 profile effective CPU expert bandwidth 2026-07-18 05:34:41 +08:00
bokiko 946fcd4f9f glm: fmt=4 support in qt_addrow/qt_matvec_rows — CPU absorb path decoded grouped int4 as int2
An all-grouped container (kv_b_proj at fmt=4) generates one correct token
and then EOS: qt_addrow and qt_matvec_rows handle fmt 0/1/2 and fall
through to the int2 decoder, so grouped-int4 kv_b was unpacked as 2-bit
pairs under a per-row scale that does not exist in the [O,ng] layout.
Prefill (S>4, reconstruction) is unaffected, which made the failure look
like an EOS bug rather than an attention bug.

Same class as #298 (CUDA absorb kernels missing fmt=4), CPU side. Existing
containers escape it because the recommended mixed-precision recipe keeps
kv_b at int8 (#237).

Adds per-group branches mirroring matmul_i4_grouped semantics. fmt 0/1/2/3
paths are untouched.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-18 00:19:39 +03:00
ZacharyZcR c3a90eca36 profile CPU GPU tier execution costs 2026-07-18 04:54:59 +08:00
KingIcyCreamProjects e2d39abd2d sampling: NaN-skip the mx scan too (review follow-up on #369)
Per review: the collapse starts one line before the sum — seeding
mx=lo[0] means a NaN at index 0 makes mx NaN, every (lo[i]-mx) NaN, and
the softmax is doomed at the max-finding, not the normalize. Seed mx
from -INFINITY and skip NaNs (x==x), mirroring the argmax_v change; if
nothing finite survives, fall back to mx=0 and let the post-sum guard
decide. The isfinite(s) guard is now the second line of defense rather
than the only one.

Clean logits take a byte-identical path (the extra x==x compare is
noise next to V expf calls). test_logit_nan gains the NaN-at-index-0
and all-NaN dist_build cases; test_topp's 123-case sweep still passes
on this tree, confirming no interaction with the #354 heap select.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 14:39:14 -05:00
KingIcyCreamProjects 7f70a8db5d sampling: guard against non-finite logits (was silent token-0 spew)
On the default serve path (TEMP>0, 0<NUCLEUS<1) a single NaN or +Inf in
the logits — a bad streamed expert tile, or an fp overflow in the matmul
at a low-RAM eviction boundary — poisoned softmax: g_pbuf became all-NaN,
dist_sample never satisfied cum>=u, and the fallback returned token 0. The
engine then emitted an unbroken run of token 0 with NO error. The greedy
path was equally blind: argmax_v started bv=lo[0] and `lo[i]>NaN` is always
false, so a NaN at index 0 pinned the argmax to 0.

- argmax_v: skip NaN (x==x) and seed from -inf, so it returns the max
  finite/+Inf entry instead of being NaN-pinned to 0. Covers greedy decode
  and the speculative-verify argmax path.
- dist_build: after the softmax sum, if s is non-finite or <=0, collapse
  g_pbuf to a one-hot over the finite argmax and warn once, instead of
  dividing every entry into NaN. Covers the nucleus and verify paths.

Both are O(1)/free on the happy path (one branch after the existing loop;
one extra comparison inside the existing argmax loop). Degrade + diagnose,
never silently corrupt.

test_logit_nan (wired into TEST_BINS): asserts argmax_v skips NaN/picks
+Inf, dist_build yields a finite normalized one-hot on the max finite
logit, dist_sample emits that token (not 0), and clean logits still give a
valid distribution. Fails on stock dev, passes with this change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 14:37:30 -05:00
volognamor ef97f3cf2b Windows: fix non-ASCII chat prompt corruption (ANSI codepage vs UTF-8)
coli passes the chat prompt to glm.exe through the PROMPT/COLI_PROMPT
environment variable. On Windows, plain getenv() is populated by the CRT
from the ANSI-codepage view of the environment block, not UTF-8 — so any
non-ASCII prompt text (Cyrillic, CJK, ...) is silently mangled before the
byte-level tokenizer ever sees it, even though the parent process (coli's
Python subprocess call) sets the value correctly via the wide env block.

Add compat_getenv_utf8() in compat.h: reads the variable through
GetEnvironmentVariableW and converts straight to UTF-8, bypassing the ANSI
codepage entirely. No-op passthrough to getenv() on non-Windows platforms.
coli_user_prompt() in glm.c now uses it for both COLI_PROMPT and PROMPT.

Verified: tiny-oracle self-test still 32/32 after rebuild, and a real run
against the full GLM-5.2-int4 model with a Cyrillic prompt now round-trips
correctly end to end (input echoed intact, coherent Cyrillic output),
where it previously produced replacement-character garbage on both input
and output.
2026-07-17 21:49:28 +03:00
JustVugg 8b36736e5d Merge branch 'p357' into trial357
# Conflicts:
#	c/Makefile
2026-07-17 20:47:21 +02:00
JustVugg 70a58799d6 Merge branch 'p354' into trialsamp
# Conflicts:
#	c/Makefile
2026-07-17 20:45:53 +02:00
Vincenzo 3d0a28b7bf Merge pull request #350 from woolcoxm/fix/build-warnings-dev
fix(warnings): silence two -Wall/-Wextra warnings on the MinGW build
2026-07-17 20:42:25 +02:00
Vincenzo 59d74ae435 Merge pull request #340 from tonuonu/fix/fs-detect-9p-statfs
glm.c: detect 9p via statfs f_type, not the /mnt/ path prefix
2026-07-17 20:42:22 +02:00
ZacharyZcR 1223f9ca2e cuda: batch ragged attention across independent streams 2026-07-18 02:01:12 +08:00
JustVugg 211d4488c3 win32: stop erasing the real ReadFile error — pread failures name their WinErr (#307)
compat_pread mapped every non-EOF ReadFile failure to EIO, so the field report
was always the same three words: "Input/output error". #307 then burned three
rounds of guessing across three people — storage? contention? alignment? —
because the actual GetLastError code never appeared anywhere.

compat_pread now stashes the code in a per-thread slot and pread_full appends
it to the failure line:

    pread qs: Input/output error (off 780592, 0/8192 bytes, WinErr=1450)

That one number is the difference between "insufficient system resources"
(1450 -> memory pressure), "sharing violation" (32 -> AV interference),
"device not ready" (21 -> the T: drive itself) and a dozen other distinct
diagnoses. Diagnostic only: no behavior change, no retry policy — that
discussion lives in #361 and should be settled AFTER the first report tells
us which error we are actually retrying.

Linux path untouched byte-for-byte; the Windows compile is certified by the
check.yml MSYS2 job.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 18:49:28 +02:00
woolcoxm 5d16368817 sampling: partial top-keep select in attention_rows DSA — O(nk) quickselect, not O(nk log nk) qsort (#356)
The DSA lightning indexer selects the top-index_topk (2048) context keys to
attend to by finding the threshold = keep-th largest attention score. It
previously full-qsorted all nk scores per layer per token (O(nk log nk)) just
to read one pivot value, then scanned the original array in position order to
build the kept set.

Replace the qsort with partial_select_desc (Hoare quickselect, median-of-three,
descending): O(nk) average to partition the keep largest into a[0..k), then the
threshold is min of that block. The two position-order scans (>thr then ==thr)
are UNCHANGED, so the kept-position set is bit-identical -- a stronger contract
than #335's sampling heap (which was multiset-only because the heap was unstable
and changed accumulation order). The quickselect pivot IS by definition the
keep-th largest, so the new threshold equals old tmp[keep-1] exactly.

Measured (bench_dsa_select, keep=2048, median of 2000 reps):
  nk=2049:   119us -> 5.7us   (21x)
  nk=8192:   626us -> 43us    (15x)
  nk=32768:  2.8ms -> 0.28ms  (10x)
  nk=65536:  6.6ms -> 0.47ms  (14x)
The gap widens with context (linear vs n-log-n). DSA only fires past index_topk,
so this is precisely the long-conversation regime where decode latency matters.

Adds test_dsa_select (in TEST_BINS): 129 cases asserting element-wise identical
kept-set vs an independent qsort reference across shapes (random, peaked,
sorted, reverse-sorted, tie-plateau, all-equal) and edges (keep==1, keep==nk).
Also directly checks the partition invariant.

Adds bench_dsa_select (on-demand, NOT a gate): reproduces the table above.
2026-07-17 09:54:40 -04:00
woolcoxm ba00889fb9 sampling: partial top-p select in dist_build — O(V) heapify + k pops, not O(V log V) qsort (#335)
dist_build() sorted the entire 151936-entry vocab by probability (qsort) on every
sampled token whenever 0 < g_nuc < 1 — the serve default — and again per draft
position under rejection sampling. Measured cost: 5.6-8.0 ms/call; the actual work
is finding the few-hundred-token head whose cumulative mass reaches g_nuc.

Replace the full qsort + linear scan with a max-heap partial select:
  - Floyd heapify g_pidx over V by descending g_pbuf prob  (O(V), cache-friendly)
  - pop winners to the array's high end until cum >= g_nuc  (k * O(log V))
  - the remaining heap prefix IS the tail -> zero it, renormalize the head

Winners land in g_pidx[out..V-1] in descending order, so s2 accumulates in the
same order as before -> head is unchanged on tie-free shapes (ties were already
unspecified under the unstable qsort). All four dist_build/dist_sample contract
properties hold: g_pbuf stays id-indexed, g_pidx stays internal, the tail is
fully zeroed, the head renormalizes to 1.

No API change, no caller change, no new globals.

c/tests/test_topp.c (new): drives the REAL dist_build via the include-glm.c
pattern against an independent double-precision reimplementation of the OLD
algorithm. 123 cases: 6 sizes (1..1519) x 5 shapes (uniform/peaked/geometric/
plateau/sharptail) x 4 nuc values, plus the g_nuc>=1 guard-off paths and V=1.
Tie-free shapes compare head values within 1e-6 rel (float vs double renorm
noise); tie shapes compare value-multisets. No scratch files -> builds clean on
Windows MinGW without the unmerged mkdtemp shim (#352).
2026-07-17 09:01:23 -04:00
woolcoxm 411f237f94 fix(warnings): silence two -Wall/-Wextra warnings on the MinGW build
A clean 'make glm' on MinGW emitted exactly two warnings, both real:

1. compat.h:240 - ignoring pragma comment [-Wunknown-pragmas]
   #pragma comment(lib, "psapi.lib") is an MSVC directive; MinGW/GCC
   warns about it. Guarded with ifdef _MSC_VER - MinGW links psapi via
   -lpsapi (already in the Makefile), MSVC keeps the pragma.

2. glm.c:1210 - g_numa_nodes defined but not used [-Wunused-variable]
   g_numa_nodes is only read/written inside ifdef __linux__ blocks, so on
   every non-Linux build it is a static that is never used. Moved the
   definition under the same __linux__ guard; nothing references it off-Linux.

Verified: rm -f *.o glm.exe && make glm -> 0 warnings, 0 errors.
2026-07-17 08:26:32 -04:00
JustVugg d5327e2252 omp: cap LLVM libomp idle spin — a parked engine must not burn 3000% CPU (#341)
The hot-thread tuning sets OMP_WAIT_POLICY=active and caps libgomp's spin
with GOMP_SPINCOUNT=200000. LLVM libomp reads NEITHER: under active policy it
sets KMP_BLOCKTIME=infinite, so once the answer ends and the engine parks on
stdin, the whole OMP team spins forever — ~100% per thread, the 3000% figure
reported on FreeBSD 15.1 in #341 (clang/libomp is the default toolchain
neighborhood there; the same applies to macOS builds).

KMP_BLOCKTIME=200 keeps the team hot for 200 ms after each parallel region —
plenty to bridge back-to-back per-expert matmuls — and then sleeps it.
setenv with overwrite=0, so a user-provided KMP_BLOCKTIME still wins, and
libgomp builds ignore the variable entirely: on GCC toolchains this commit
is a no-op by construction.

Not reproduced locally (this box is gcc/libgomp and idles clean); shipped as
the mechanism-matching candidate fix with a request on the issue for the
reporter to verify on dev and to name their OpenMP runtime (ldd | grep omp).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 13:53:20 +02:00
Tonu Samuel 81a56777df glm.c: detect 9p by statfs f_type, not the "/mnt/" path prefix
The slow-filesystem warning fired for any model path under /mnt/, which
false-positives on native-Linux mounts (ZFS/ext4/xfs/NFS) that commonly
live there. Check the actual filesystem type via statfs() against the 9p
superblock magic (0x01021997) instead, so it only warns for a genuine
WSL 9p mount. Linux-only (statfs); no behavior change on other platforms.
2026-07-17 08:00:26 +03:00
Vincenzo ee5d273cd4 Merge pull request #232 from nbeerbower/profiling-upstream
Profiling: PROF=1 opt-in performance profile + live per-turn Profiling page in the web dashboard
2026-07-16 19:58:31 +02:00
JustVugg 98f7c88ca9 numa: parenthesise the page-align arithmetic (#313)
`(uintptr_t)p+n+4095 & ~(uintptr_t)4095` is CORRECT — `+` binds tighter than
`&` in C, so the rounding happens before the mask, which is what was meant.
But gcc -Wall warns (`suggest parentheses around '+' in operand of '&'`), and
this repo builds clean at -Wall -Wextra. Warning-free is a property worth more
than the two characters it costs to keep: it is what makes a NEW warning
visible instead of scrolling past in a wall of noise.

No behaviour change; the emitted code is identical.

Co-Authored-By: ZacharyZcR <#313>

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 19:50:20 +02:00
Vincenzo 3cec90cd45 Merge pull request #313 from ZacharyZcR/feat/numa-expert-interleave
glm: COLI_NUMA=1 — interleave resident weights across NUMA nodes (+13% decode on 2-socket, raw mbind, zero deps) (#82)
2026-07-16 19:49:17 +02:00
Vincenzo 644e0d16c6 Merge pull request #302 from alekseysorokin68/main
Fix: set stdout to O_BINARY on Windows to fix READY sentinel
2026-07-16 19:49:11 +02:00
JustVugg 6de32c55f6 Don't trust a converted checkpoint's config for the stop set
The engine armed its stop tokens from config.json's eos_token_id and nothing
else. That trusts metadata written by third-party conversion tooling, which is
a thing we already know goes wrong: the README documents a mirror shipping
int4 MTP heads that silently give 0% draft acceptance. GLM-5.2 declares THREE
eos ids (<|endoftext|>, <|user|>, <|observation|>); a converter that rewrites
config.json with a reduced list leaves the engine stopping on fewer tokens
than the model emits, and the missed ones get detokenized and printed into the
chat as literal text while generation runs past the end of the turn.

Two independent defenses:

  - eos_token_id is now unioned with generation_config.json, which is
    HuggingFace's authority for generation (config.json often carries a
    partial legacy copy). An extra stop is harmless; a missing one is not.

  - every added-token the TOKENIZER marks "special":true is armed as a stop,
    whatever the configs say. Those are control tokens (<|user|>, <|assistant|>,
    <sop>, [gMASK], the image/video/audio markers) and are never legitimate
    content in a reply -- GLM itself lists three of them as official eos.
    <think>/<tool_call>/<arg_key> are "special":false and are deliberately NOT
    swept up: they are real output. tok.h was parsing added_tokens but throwing
    the "special" flag away, so the distinction wasn't available to anyone.

On the real per-row checkpoint this takes the armed set from 3 to 18:
  [stop] 18 stop tokens: 154820 154827 154829 154821 ... (15 from the
  tokenizer's special set)

Honesty about scope: this is hygiene for a class of bug, NOT a fix for the
trailing-junk report on #298 that prompted it. I hypothesised @woolcoxm's g64
checkpoint had lost eos ids in conversion; he checked, and it hadn't -- his
config arms all three correctly. The emit path is also innocent: is_stop() is
checked BEFORE emit() at every one of the four call sites (4215, 4256, 4908,
4987), so a correctly-armed stop cannot be printed. His trailing junk is still
unexplained and is more likely quantization noise. What this commit buys is
that a checkpoint we don't control cannot leak control tokens into a reply,
which was true before and is not now.

tests/test_stops.c covers both defenses: the union, a missing
generation_config.json, BOTH configs mutilated (the tokenizer still stops all
five control tokens while leaving <think> alone), and T=NULL (the validation
path keeps config-only behaviour).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 17:59:34 +02:00
Claude 36394bc317 Profiling: fix the phase accounting so 'Other' stops swallowing the disk stall
The dashboard's wall-time stack assumed ewait = felt I/O stall and
edisk = overlapped read service, but the engine measured neither:
t_ewait was declared and reported yet never incremented (the blue
'I/O wait' segment always read 0.00s), while t_edisk accumulated on
the COMPUTE thread — blocking OMP loads, PIPE dispatch, pipe_wait
spins — i.e. exactly the felt stall. The UI excluded it from the
stack as 'overlapped', so the whole disk stall landed in 'Other'
(e.g. 46.9s Other vs 45.3s disk on a 54.1s turn).

Now the semantics match the contract:
- t_ewait accumulates the compute-thread stall (blocking loads,
  PIPE dispatch, pipe_wait spins) and feeds the in-stack I/O wait.
- Disk service is real: expert_load is timed on whichever thread
  runs the read (PIPE workers, OMP loaders, pilot) into an atomic
  ns counter (g_edisk_ns) — thread-seconds, overlapped with compute.
- prof_report's I/O share and verdict use the felt wait only, and
  the expert I/O line prints 'read service / felt wait' so the
  wait:service gap shows how well PIPE is hiding the reads.

PROF wire format, /profile JSON and the web UI are unchanged —
the dashboard's existing assumptions are simply true now.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dd4hHD2Dsvgr7VAwEVFqWK
2026-07-16 09:00:31 -04:00
Claude 6afffbcbf2 Profiling page: per-turn phase timings, live in the web dashboard
The engine already tracks where each turn's wall time goes (expert-disk
service, I/O wait, expert matmul, attention, lm_head) — it just only spoke
at exit or under PROF=1. Stream it instead:

- glm.c: mux serve emits a per-turn "PROF" protocol line next to TIERS/HITS
  (window deltas per request, same convention as the STAT hit%); the phase
  window base is now always captured (a few loads per request).
- openai_server.py: parses PROF into a 120-turn rolling window and serves it
  at /profile (read-only, same trust level as /health).
- web: new Profiling tab — stat tiles (tok/s, wall, tokens/forward, disk
  service), wall-time composition bars for the last turn and the window,
  per-turn throughput and stacked phase columns with hover readouts, and a
  table of recent turns. Disk service is shown apart from the stack: it
  overlaps with compute, so only the I/O wait the compute thread felt counts
  inside wall time. Phase colours are a CVD-validated set with gaps + legend
  + table so identity never rides on colour alone.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WhTmF8yvEBgSkUKSVfZF7P
2026-07-16 08:56:12 -04:00
Claude 63a6824881 PROF=1: per-request reports in mux serve mode too (the web-dashboard path)
run_serve_mux (SERVE_BATCH, used by openai_server.py / coli web) completes
requests in mux_done, so the run_serve per-turn hook never fired there.
Snapshot the window where hits0 is taken, record batched-forward latency
around step_decode_batch, report on stderr at DONE. With KV_SLOTS>1 the
window shares batched forwards across slots — same convention as the
existing STAT hit%.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TCBNCciBaHea41QmidLUMn
2026-07-16 08:56:12 -04:00
Claude 09bf17f001 PROF=1: opt-in performance profile (latency percentiles, expert I/O, tuning verdict)
Answers 'where does it slow down on THIS machine with THIS config' so users
can tune RAM_GB/PIPE/DIRECT/PIN for their hardware without folklore:

- startup header: CPU/cores/RAM/backend + effective knobs (cache cap, pin,
  DRAFT/PIPE/DIRECT/MMAP/IDOT/DSA/PILOT/CACHE_ROUTE) — every saved log is
  self-describing when comparing runs across configs or machines
- per-forward decode latency ring (32k) -> p50/p90/p99/max, plus a tail
  diagnosis when p99 >> p50 (cold-cache expert loads)
- expert I/O accounting at the pread/mmap-touch sites: GB fetched, MB/token,
  GB/s, hit rate, loads/token, pinned/LRU tier fill
- phase shares of wall time and a plain-language verdict naming the knob
  most likely to move tok/s (I/O-bound vs compute-bound vs attention-bound)
- reports after REPLAY / PROMPT / oracle runs (stdout) and per turn in serve
  mode (stderr; stdout stays the framed protocol)

Additive only: with PROF unset every mode's output is byte-identical.
hwinfo_emit's /proc probe is factored into hw_probe() and shared.
Validated end-to-end on a tiny-random unquantized fixture (REPLAY, PROF on/
off, RAM_GB squeeze flips hit 96.9%->26.6% and the verdict follows);
make check clean, 0 warnings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TCBNCciBaHea41QmidLUMn
2026-07-16 08:56:12 -04:00
ZacharyZcR 2b182ae7f6 glm: COLI_NUMA=1 interleaves resident weights across NUMA nodes — +13% decode on 2S, raw mbind, no libnuma (#82) 2026-07-16 20:32:03 +08:00
aleks 6728029555 fix: set stdout to O_BINARY on Windows to prevent READY sentinel corruption
Both sentinels in c/coli (READY and END) end with \n. Under CRT text
mode on Windows, printf() translates \n to \r\n, so the engine emits
\x01\x01READY\x01\x01\r\n. The Python coli wrapper checks
endswith(SENTINEL) which expects bare \n — the match never fires and
chat hangs forever.

_setmode(fileno(stdout), O_BINARY) at engine startup switches stdout
to binary mode so \n passes through unchanged. This matches the
belt-and-braces reasoning already in compat.h:88 (O_BINARY for file
I/O). The fix is defense-in-depth: on the documented w64devkit/MinGW
build path, CRT text mode may or may not apply depending on how the
terminal is attached, so this ensures correctness regardless.

Note: I was unable to compile-test this on the full codebase because
glm.c uses POSIX-only symbols (mmap, madvise, select, fd_set, struct
stat/fstat) that are unavailable on Windows without platform guards.
The codebase compiles on Linux (Ubuntu 24.04, GCC 13) and macOS
(Apple clang). Windows compilation requires wrapping those POSIX calls
in #if defined(__APPLE__) || defined(__linux__) || defined(__FreeBSD__)
guards — I opened this as a separate concern for the maintainer.
2026-07-16 14:10:27 +03:00
Vincenzo 0d297ee85b Merge pull request #301 from NeuralNotwerk/pr/pin-auto
PIN=auto: seed the pin from the live usage history
2026-07-16 12:38:09 +02:00
JustVugg 35f90b9e76 Quarantine EXPERT_BUDGET: the operating window is measured empty (#303)
@bokiko benchmarked the feature on three hosts and found every setting is
either no faster or no longer coherent. Reproduced here on a 25 GB box, and
the numbers are worse than "a quality/speed tradeoff":

  - budget=8 -> hellaswag 30% vs 90% with it off (25% is the chance floor)
  - budget=4 -> decode is literal noise ("The **1...: s2151:")
  - MTP acceptance 0%. Which experts survive the cap depends on cache
    residency at the moment of the forward, so draft and verify do not
    compute the same function -- the invariant #294 just established for
    #163, violated through cache state instead of kernel dispatch. SPEC_PIN
    cannot fix that and should not try: it is feature semantics, not
    dispatch.
  - 0.13 tok/s vs 0.30 baseline, while loading 14.66 experts per layer
    against a topk=8 baseline. The cap does MORE disk I/O than not using it.
    "~335 GB I/O saved" counts dropped experts, not bytes not read.

EXPERT_BUDGET>0 is now ignored unless EXPERT_BUDGET_EXPERIMENTAL=1, with the
measurements printed. The code stays compiled and developable: MoE-Spec
(arXiv 2602.16052) is not a wrong idea, this implementation just has no
point where it is both faster and correct. Re-enabling it by default needs a
quality number next to every speed number.

Also gates the cap to decode (S<=4), @woolcoxm's fix from #292/#298: during
prefill the batch union is 30-100+ experts, and capping to 4-8 drops most of
them, corrupting the hidden state and therefore the KV cache. Necessary but
not sufficient -- the run above already includes it.

Removes issue_budget.md, issue_diskio.md and issue_grouped_quant.md: design
notes in the repo root, unlinked from any README, whose only user-facing
content was "EXPERT_BUDGET=6-8 -- good speedup, minimal quality loss" and
"+83% decode at budget=4". Measurement says otherwise, and they were the
only thing on main telling anyone to switch this on. They stay in history.

Reported-by: bokiko <#303>
Co-Authored-By: woolcoxm <#292>

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 12:26:17 +02:00
JustVugg d57955e95a Say why the engine died instead of dying mute (#305)
Two silent failures that compound into "[engine terminated]" with no cause.

cap_for_ram() floors the expert cache at cap=1. When resident+slack have
already blown the budget, avail is negative and capmax would be 0 -- "I do
not fit in your budget". Flooring it to 1 and carrying on turns "I do not
fit" into "I overshoot", which is precisely the mid-generation OOM-kill the
function exists to prevent: it printed "projected peak 25.1 GB" against a
22 GB budget and started anyway. Now it says so, names PIN_GB when that is
what inflated the resident set, and refuses to start when the peak also
exceeds the memory actually available on the machine (COLI_RAM_OVERCOMMIT=1
overrides).

The kernel kills with SIGKILL: no error, no log, stdout just closes. coli
read that EOF and printed "[engine terminated]" without ever reaping the
child, so an OOM-kill was indistinguishable from a clean exit -- the report
in #305 (Debian 12, dies mid-generation, no message). engine_diag() now
reports the signal or exit code, names the OOM-killer when it was SIGKILL,
and shows the tail of the engine's stderr.

Reported-by: Ne00n <#305>

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 12:25:53 +02:00
NeuralNotwerk 81c0398877 PIN=auto — seed the pin from the live usage history
PIN=auto resolves to <model>/.coli_usage, the history that serve mode
appends after every turn, so each restart's hot-store placement follows the
accumulated REAL workload instead of a frozen one-shot profile; stats.txt is
the fallback for a virgin model dir, and with neither present the run simply
starts unpinned. Same magic-value convention as PIN_GB=all (#80); explicit
paths and the AUTOPIN flow are untouched (PIN set skips AUTOPIN as before).
This lifts deployment entrypoints' "prefer .coli_usage over stats.txt" shell
plumbing into the engine, plus the ENVIRONMENT.md row.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 07:08:28 +00:00
JustVugg d4b4f33f22 Merge main into dev: Windows fixes (#275-279, #290) alongside engine work (#274, #293, #294, #297) 2026-07-16 07:59:53 +02:00
Vincenzo 12d3bd5140 Merge pull request #290 from NeuralNotwerk/fix/mlock-under-mmap
glm: make MLOCK=1 actually wire pinned experts under COLI_MMAP
2026-07-16 07:51:48 +02:00
Vincenzo e8479da97d Merge pull request #278 from KingIcyCreamProjects/pr/win32-hwinfo
glm: hwinfo_emit Windows support — CPU brand, cores, RAM were hardcoded-zero
2026-07-16 07:51:44 +02:00
Vincenzo a6a99925a2 Merge pull request #294 from bokiko/fix/spec-kernel-pin
glm: pin draft+verify to one kernel family during speculation; soften the MTP guard (#163)
2026-07-16 07:37:32 +02:00
Vincenzo 6d190d91e1 Merge pull request #293 from woolcoxm/fix/pipe2-profiling-and-cuda-mtp-optin
cuda: fix pipe2 profile double-count + COLI_CUDA_MTP opt-in (#292)
2026-07-16 07:35:48 +02:00
Vincenzo 5f167e5964 Windows: pipe2 resident pipeline at decode + EXPERT_BUDGET ws_b cache fix (#274)
Windows: 1.07 tok/s decode via GPU resident pipeline + disk-I/O tuning (3.2x over stock)
2026-07-16 07:29:34 +02:00
bokiko 37da111bb4 glm: pin draft+verify to one kernel family during speculation; soften MTP guard (#163)
MTP acceptance collapses when the draft (S=1) and verify (S>=2) forwards do
not compute the same function. Three switches were S-dependent: the int4
IDOT gate (S>=g_i4s, asymmetric exactly on ISAs where g_i4s>1), the
S==1-only fused gate+up pair, and the Metal GEMM row threshold.

SPEC_PIN=1 (default) pins every forward issued while model drafts are live
to the platform S=1 kernel family, so draft and verify agree by
construction. Prefill and non-speculative decode keep their gates.
SPEC_PIN=0 restores the old behavior for A/B.

The 24-proposal <10% guard now pauses drafting for 256 tokens and re-arms
on a fresh window instead of latching off for the whole session.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-16 01:36:15 +03:00
woolcoxm 78c675eb8b cuda: fix pipe2 profile double-count + add COLI_CUDA_MTP opt-in (#292)
Two issues found by the diagnostic sweep in #292:

1. Profile double-count in pipe_layer_sparse: the outer t_emm span wrapped
   both the shared-expert GPU dispatch AND the moe() call, but moe()
   self-times its own t_emm internally (6 accumulation sites). Routed-expert
   matmul was counted twice -> accounted > elapsed -> 'other' went negative,
   and expert-matmul showed an inflated ~120s. Fix: split into two narrow
   spans around only the GPU work moe() does NOT cover (shared-expert
   dispatch, routed upload+add), letting moe() self-time as all other
   callers already do. Verified: 'other' now positive (14-17s),
   expert-matmul realistic (5-6s).

2. MTP blanket-disabled under CUDA (g_draft=0 when g_cuda_enabled). This
   was a conservative guard from #163 before the root cause (cold-expert
   fused-pair + IDOT kernel divergence) was fully diagnosed. GPU-resident
   experts have no divergence; the cold subset still does but achieves
   30-50% acceptance anyway (matching non-CUDA builds). Add COLI_CUDA_MTP=1
   opt-in so users can test speculation under CUDA. Default unchanged.
   Verified: COLI_CUDA_MTP=1 -> MTP ACTIVE (draft=3), 44% acceptance,
   3 forwards for 8 tokens (vs 7 without MTP).

Refs #292 #163
2026-07-15 18:33:25 -04:00
NeuralNotwerk 97fa52698f glm: make MLOCK=1 actually wire pinned experts under COLI_MMAP
pin_wire() only locked the legacy private-slab path (s->slab / s->fslab),
which stay NULL under COLI_MMAP -- so MLOCK=1 on an mmap build silently
wired 0 bytes and the 'pinned' set was just warm page cache, evictable
under memory pressure (measured: hundreds of MB/s of re-reads from disk
mid-generation on a 503 GB host once the cache tightened).

Three pieces:
- qt_wire_mmap(): mlock each pinned expert's weight + scale ranges inside
  the file mappings. Skips cuda_eligible QTs: VRAM-tier experts compute
  from device memory, and expert_host_release() early-returns for mmap
  experts (no slab) WITHOUT nulling q8/q4, so a host-pointer check alone
  wires the whole VRAM tier too (~137 GB of never-touched locked pages on
  a 6x32GB-class rig -- enough to starve the kernel into thrashing).
- qt_unwire_mmap(): REPIN gpu_swap promotions drop the promoted expert's
  host lock instead of leaking it on every swap.
- expert_load() deliberately does NOT wire (it also runs for the transient
  VRAM-staging pass); wiring happens once in pin_wire() on the final set.

Measured on GLM-5.2 int4, 4x RTX 5090 + 1x RTX 4090, 503 GB RAM,
PIN_GB=all: wired goes 0 -> 226 GB (exactly the RAM tier), decode-time
disk reads drop to zero once warm.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 22:12:31 +00:00