Two independent defects in the CUDA syntax job, both caught on its first run.
1. `cuda: '12.6.3'` — Jimver/cuda-toolkit@v0.2.19's version table stops at
12.6.2, so the install step died with "Version not available: 12.6.3"
before nvcc was ever invoked. One digit.
2. `nvcc ... 2>&1 | head -40` — a pipeline exits with the status of its LAST
command, so head's 0 masked every nvcc error. The job printed "CUDA syntax
check passed" unconditionally: it could not fail. Defect 1 is why we found
out, since it broke the step *before* the pipe.
The second one is the one that matters. backend_cuda.cu is compiled by nothing
else in this repo — no local build, no test — so this job is the only thing
standing between a CUDA change and a user's GPU. A check that cannot fail is
worse than no check: it buys false confidence in exactly the file that most
needs the real thing. #298 spent a night debugging CUDA against no oracle at
all; this job is supposed to be that oracle.
It is also, precisely, the disease of the week in YAML form: a signal that
measures its own intention rather than the thing it claims to measure. See the
`~335 GB I/O saved` counter that counted dropped experts instead of bytes not
read (#303), and `route_agree: 95.3%` cited as "quality preserved" when it
only measures which experts coincide.
The other three jobs (engine, web, python) passed on the first run.
Co-Authored-By: ZacharyZcR <#144>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The engine armed its stop tokens from config.json's eos_token_id and nothing
else. That trusts metadata written by third-party conversion tooling, which is
a thing we already know goes wrong: the README documents a mirror shipping
int4 MTP heads that silently give 0% draft acceptance. GLM-5.2 declares THREE
eos ids (<|endoftext|>, <|user|>, <|observation|>); a converter that rewrites
config.json with a reduced list leaves the engine stopping on fewer tokens
than the model emits, and the missed ones get detokenized and printed into the
chat as literal text while generation runs past the end of the turn.
Two independent defenses:
- eos_token_id is now unioned with generation_config.json, which is
HuggingFace's authority for generation (config.json often carries a
partial legacy copy). An extra stop is harmless; a missing one is not.
- every added-token the TOKENIZER marks "special":true is armed as a stop,
whatever the configs say. Those are control tokens (<|user|>, <|assistant|>,
<sop>, [gMASK], the image/video/audio markers) and are never legitimate
content in a reply -- GLM itself lists three of them as official eos.
<think>/<tool_call>/<arg_key> are "special":false and are deliberately NOT
swept up: they are real output. tok.h was parsing added_tokens but throwing
the "special" flag away, so the distinction wasn't available to anyone.
On the real per-row checkpoint this takes the armed set from 3 to 18:
[stop] 18 stop tokens: 154820 154827 154829 154821 ... (15 from the
tokenizer's special set)
Honesty about scope: this is hygiene for a class of bug, NOT a fix for the
trailing-junk report on #298 that prompted it. I hypothesised @woolcoxm's g64
checkpoint had lost eos ids in conversion; he checked, and it hadn't -- his
config arms all three correctly. The emit path is also innocent: is_stop() is
checked BEFORE emit() at every one of the four call sites (4215, 4256, 4908,
4987), so a correctly-armed stop cannot be printed. His trailing junk is still
unexplained and is more likely quantization noise. What this commit buys is
that a checkpoint we don't control cannot leak control tokens into a reply,
which was true before and is not now.
tests/test_stops.c covers both defenses: the union, a missing
generation_config.json, BOTH configs mutilated (the tokenizer still stops all
five control tokens while leaving <think> alone), and T=NULL (the validation
path keeps config-only behaviour).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@woolcoxm's matmul_i4_grouped_pair reads x once instead of twice for the
gate+up pair. Verified here against his branch (86e91b1 merged onto dev):
correct to ~2e-8 relative vs the double reference, and BIT-EXACT against two
separate matmul_i4_grouped calls on aligned shapes -- which is the shape the
real g64 checkpoints have (I = 2048 / 6144, gs = 64). His kernel is good.
Guarded behind COLI_HAVE_GROUPED_PAIR since the function only exists on that
branch; add -DCOLI_HAVE_GROUPED_PAIR to the test's Makefile rule when #298
lands and the pair cases activate.
The checks are deliberately asymmetric, and the reason is worth recording.
Bit-exactness is asserted ONLY when I % gs == 0: there every group is covered
by the AVX2 body, whose accumulation order matches the unfused kernel, so any
difference is a real bug. With a partial last group the tail falls to scalar
code and the compiler may contract/reassociate the fused body differently,
producing ~1e-7 differences -- rounding, not logic. My first version demanded
bit-exactness everywhere and duly "found" a bug in his kernel that did not
exist; the tell was that only `up` differed and never `gate`, which is FP luck
rather than a code path. Correctness is checked everywhere against the double
reference; identity only where identity is actually implied.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
matmul_i4_grouped is the reference the CUDA fmt=4 port (#298) is expected to
reproduce, and it had no test of its own. @woolcoxm is currently debugging a
CUDA backend against an oracle nobody had verified, which is two moving
targets at once -- and he can't cross-check on CPU, since a 5-prompt run
takes 8 hours on the 744B model.
This checks matmul_i4_grouped against a plain-C reference that dequantizes
nibble -> (v-8)*scale[i/gs] and accumulates in double, over 11 shapes: I a
clean multiple of gs, a partial last group (the glen clamp), odd I (the
scalar nibble tail), gs > I, gs=16/64/128, S>1, and the nibble extremes
0x00/0xFF -- which decode to -8/+7 because the format is offset-encoded, not
two's complement. Reading that backwards turns 15 into -1 and looks like
data-dependent noise rather than a bug.
All 11 shapes match to ~1e-8 relative, so the CPU kernel is exact and can be
trusted as the reference.
One note on the tolerance, because the first draft of this test got it wrong
and "found" a bug that wasn't there: the error is compared against the sum of
|terms|, not against |result|. A dot product of signed terms can land near
zero through cancellation, and then a 1e-6 absolute error -- ordinary f32
accumulator precision -- reads as a 1e-3 relative one. A wrong scale index or
a wrong group boundary shifts the result by a fraction of the terms, so it is
still caught at 1e-6.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@bokiko benchmarked the feature on three hosts and found every setting is
either no faster or no longer coherent. Reproduced here on a 25 GB box, and
the numbers are worse than "a quality/speed tradeoff":
- budget=8 -> hellaswag 30% vs 90% with it off (25% is the chance floor)
- budget=4 -> decode is literal noise ("The **1...: s2151:")
- MTP acceptance 0%. Which experts survive the cap depends on cache
residency at the moment of the forward, so draft and verify do not
compute the same function -- the invariant #294 just established for
#163, violated through cache state instead of kernel dispatch. SPEC_PIN
cannot fix that and should not try: it is feature semantics, not
dispatch.
- 0.13 tok/s vs 0.30 baseline, while loading 14.66 experts per layer
against a topk=8 baseline. The cap does MORE disk I/O than not using it.
"~335 GB I/O saved" counts dropped experts, not bytes not read.
EXPERT_BUDGET>0 is now ignored unless EXPERT_BUDGET_EXPERIMENTAL=1, with the
measurements printed. The code stays compiled and developable: MoE-Spec
(arXiv 2602.16052) is not a wrong idea, this implementation just has no
point where it is both faster and correct. Re-enabling it by default needs a
quality number next to every speed number.
Also gates the cap to decode (S<=4), @woolcoxm's fix from #292/#298: during
prefill the batch union is 30-100+ experts, and capping to 4-8 drops most of
them, corrupting the hidden state and therefore the KV cache. Necessary but
not sufficient -- the run above already includes it.
Removes issue_budget.md, issue_diskio.md and issue_grouped_quant.md: design
notes in the repo root, unlinked from any README, whose only user-facing
content was "EXPERT_BUDGET=6-8 -- good speedup, minimal quality loss" and
"+83% decode at budget=4". Measurement says otherwise, and they were the
only thing on main telling anyone to switch this on. They stay in history.
Reported-by: bokiko <#303>
Co-Authored-By: woolcoxm <#292>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two silent failures that compound into "[engine terminated]" with no cause.
cap_for_ram() floors the expert cache at cap=1. When resident+slack have
already blown the budget, avail is negative and capmax would be 0 -- "I do
not fit in your budget". Flooring it to 1 and carrying on turns "I do not
fit" into "I overshoot", which is precisely the mid-generation OOM-kill the
function exists to prevent: it printed "projected peak 25.1 GB" against a
22 GB budget and started anyway. Now it says so, names PIN_GB when that is
what inflated the resident set, and refuses to start when the peak also
exceeds the memory actually available on the machine (COLI_RAM_OVERCOMMIT=1
overrides).
The kernel kills with SIGKILL: no error, no log, stdout just closes. coli
read that EOF and printed "[engine terminated]" without ever reaping the
child, so an OOM-kill was indistinguishable from a clean exit -- the report
in #305 (Debian 12, dies mid-generation, no message). engine_diag() now
reports the signal or exit code, names the OOM-killer when it was SIGKILL,
and shows the tail of the engine's stderr.
Reported-by: Ne00n <#305>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
One HF 504 killed the whole bench. Now: load_dataset retries with
exponential backoff (hf_hub resumes partial downloads from cache); a task
that still fails is skipped instead of killing the rest; JSONLs are written
atomically (coli only checks existence, so a truncated file from an
interrupted run would block re-download forever); coli bench drops
still-missing tasks with a warning and refuses to run eval with none.
Co-Authored-By: Claude <noreply@anthropic.com>
Apple clang 16 (clang-1600.0.26.6) defaults objective-c++ to a pre-C++11
dialect, so the raw string literal holding the Metal shader in
backend_metal.mm fails to parse:
backend_metal.mm:12:29: error: use of undeclared identifier 'R'
backend_metal.mm:13:10: fatal error: 'metal_stdlib' file not found
Pinning gnu++17 on METALXX fixes 'make glm METAL=1' and 'make metal-test'.
Verified on macOS 15 / M4 Max: both targets build and all metal backend
tests pass.
PIN=auto resolves to <model>/.coli_usage, the history that serve mode
appends after every turn, so each restart's hot-store placement follows the
accumulated REAL workload instead of a frozen one-shot profile; stats.txt is
the fallback for a virgin model dir, and with neither present the run simply
starts unpinned. Same magic-value convention as PIN_GB=all (#80); explicit
paths and the AUTOPIN flow are untouched (PIN set skips AUTOPIN as before).
This lifts deployment entrypoints' "prefer .coli_usage over stats.txt" shell
plumbing into the engine, plus the ENVIRONMENT.md row.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Select a portable architecture from the compiler target instead of forcing x86-64-v3 on every platform. On macOS, only enable Homebrew OpenMP when its header and library actually exist, preserving the dependency-free fallback.
MTP acceptance collapses when the draft (S=1) and verify (S>=2) forwards do
not compute the same function. Three switches were S-dependent: the int4
IDOT gate (S>=g_i4s, asymmetric exactly on ISAs where g_i4s>1), the
S==1-only fused gate+up pair, and the Metal GEMM row threshold.
SPEC_PIN=1 (default) pins every forward issued while model drafts are live
to the platform S=1 kernel family, so draft and verify agree by
construction. Prefill and non-speculative decode keep their gates.
SPEC_PIN=0 restores the old behavior for A/B.
The 24-proposal <10% guard now pauses drafting for 256 tokens and re-arms
on a fresh window instead of latching off for the whole session.
Co-Authored-By: Claude <noreply@anthropic.com>
Two issues found by the diagnostic sweep in #292:
1. Profile double-count in pipe_layer_sparse: the outer t_emm span wrapped
both the shared-expert GPU dispatch AND the moe() call, but moe()
self-times its own t_emm internally (6 accumulation sites). Routed-expert
matmul was counted twice -> accounted > elapsed -> 'other' went negative,
and expert-matmul showed an inflated ~120s. Fix: split into two narrow
spans around only the GPU work moe() does NOT cover (shared-expert
dispatch, routed upload+add), letting moe() self-time as all other
callers already do. Verified: 'other' now positive (14-17s),
expert-matmul realistic (5-6s).
2. MTP blanket-disabled under CUDA (g_draft=0 when g_cuda_enabled). This
was a conservative guard from #163 before the root cause (cold-expert
fused-pair + IDOT kernel divergence) was fully diagnosed. GPU-resident
experts have no divergence; the cold subset still does but achieves
30-50% acceptance anyway (matching non-CUDA builds). Add COLI_CUDA_MTP=1
opt-in so users can test speculation under CUDA. Default unchanged.
Verified: COLI_CUDA_MTP=1 -> MTP ACTIVE (draft=3), 44% acceptance,
3 forwards for 8 tokens (vs 7 without MTP).
Refs #292#163
pin_wire() only locked the legacy private-slab path (s->slab / s->fslab),
which stay NULL under COLI_MMAP -- so MLOCK=1 on an mmap build silently
wired 0 bytes and the 'pinned' set was just warm page cache, evictable
under memory pressure (measured: hundreds of MB/s of re-reads from disk
mid-generation on a 503 GB host once the cache tightened).
Three pieces:
- qt_wire_mmap(): mlock each pinned expert's weight + scale ranges inside
the file mappings. Skips cuda_eligible QTs: VRAM-tier experts compute
from device memory, and expert_host_release() early-returns for mmap
experts (no slab) WITHOUT nulling q8/q4, so a host-pointer check alone
wires the whole VRAM tier too (~137 GB of never-touched locked pages on
a 6x32GB-class rig -- enough to starve the kernel into thrashing).
- qt_unwire_mmap(): REPIN gpu_swap promotions drop the promoted expert's
host lock instead of leaking it on every swap.
- expert_load() deliberately does NOT wire (it also runs for the transient
VRAM-staging pass); wiring happens once in pin_wire() on the final set.
Measured on GLM-5.2 int4, 4x RTX 5090 + 1x RTX 4090, 503 GB RAM,
PIN_GB=all: wired goes 0 -> 226 GB (exactly the RAM tier), decode-time
disk reads drop to zero once warm.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Windows cmd.exe always exports PROMPT (its prompt template, default "$P$G")
into a child's environment. The engine's `getenv("PROMPT")` picked that up, so
running `glm.exe` from cmd for the oracle self-test instead entered text-
generation mode — loading a tokenizer the tiny model doesn't ship and failing
with "tokenizer.json: No such file". PowerShell has no PROMPT env var, so it
worked there (galmok's 32/32) but not in cmd — same command, different shell.
New coli_user_prompt(): honors COLI_PROMPT everywhere, and on Windows ignores a
PROMPT that carries cmd's $-metacodes ($P,$G,...) — a real prompt has none. cmd
users can still pass a prompt via COLI_PROMPT or a non-$ PROMPT. Verified: $P$G
-> oracle mode, "Explain recursion" -> honored, COLI_PROMPT -> honored.
Third native-Windows bug surfaced by a from-scratch main download (after the
-lpsapi link fix and the bilingual oracle message).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Running a real model with no PROMPT lands in oracle self-test mode, which
compares against ref_glm.json — the TINY model's oracle. The guard that
detects the vocab mismatch and points the user at PROMPT=/coli chat was
Italian-only, which read as a crash to English users (#271, galmok). Lead
with English, keep an IT footer, and add the `coli chat` hint.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
compat.h's rss_gb() calls GetProcessMemoryInfo and links psapi via
#pragma comment(lib,"psapi.lib") — an MSVC-ism. MinGW gcc ignores that pragma
(emits -Wunknown-pragmas), and the Windows LDFLAGS never linked psapi, so on a
gcc that doesn't honor the pragma (e.g. 16.1.0 UCRT) the build fails with
`undefined reference to GetProcessMemoryInfo`. Add -lpsapi to the Windows
LDFLAGS; harmless on toolchains where the pragma also resolves it. Found while
building on native Windows 11 with winlibs GCC 16.1.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Opt-in (default off): prints the generated token ids to stderr after
generation, for exact comparison across decode paths (e.g. the S=1
resident pipeline vs the CPU path). Used to verify the pipe_layer_sparse
gate relaxation is token-exact vs the CPU decode path (see #273).
The grouped-int4 (fmt=4) loader hard-coded gs=128 at all three detection
sites (resident weights + the two expert-load paths: mmap and pread-slab).
A checkpoint converted with --group-size 64 (or any non-128 size) wrote a
valid file that the engine then misdetected as plain per-row int4 (fmt=2)
and read the wrong number of scales -> garbage output, silently.
The conversion path (convert_fp8_to_int4.py --group-size) already emits
arbitrary group sizes, and the compute kernel (matmul_i4_grouped) is
fully gs-generic (only constraint: gs multiple of 16, the AVX2 width).
The loader was the sole gap.
Replace the three hardcoded checks with one shared helper, detect_group_size(),
that derives gs from the scale-array byte count by probing candidate sizes
{16,32,48,64,96,128,192,256} finest-first. Data-driven: any listed size
just works; per-row int4 (ns == O*4) correctly returns gs=0 -> stays fmt=2.
Verified against real GLM-5.2 expert dims (gate/up O=2048,I=6144 and down
O=6144,I=2048): g64/g128/g256 all detect correctly, and per-row is never
misdetected as grouped. g128 (the only previously-supported size) is
unchanged -> no regression for existing checkpoints.
This unblocks the g64 lane that ZacharyZcR's #225 ablation identified as
beating shipped per-row int4 (-7.5pp vs -9.3pp) at ~19% fewer bits: the
converter can produce it, and now the engine can load it.
Refs #225
Per ZacharyZcR's 6x5090 A/B on #273: the S=1 resident-pipeline relaxation is
+49% on a single GPU (5070 Ti) but a wash on multi-GPU. With layers sharded
across N devices, each resident forward at S=1 crosses P2P per layer group and
those small hops don't amortize — the same term that killed pipe x head-shard
in #111. Multi-GPU decode walls on disk service, which pipe2 can't touch.
Make the threshold device-count-dependent:
- single GPU (g_cuda_ndev<=1): S>=1 (the breakthrough path)
- multi GPU (g_cuda_ndev> 1): S>=8 (the original prefill-only gate)
with COLI_CUDA_PIPE_S_MIN as an env override for anyone who wants to measure.
The two calibration points bracket the design space:
1x 5070 Ti, modest CPU, "other"-bound decode -> S=1 (+49%)
6x 5090, sharded, disk-service-bound -> S=1 a wash, keep S>=8
Refs #273 (comment)
CACHE_ROUTE=1 ROUTE_J=2 ROUTE_M=12 steers the MoE router to prefer
cache-resident experts, reducing the miss rate from 27% to 17% and disk
I/O from 12.4s to 8.5s. Combined with the full optimization stack
(disk tuning + CUDA pipe2 + ws_b cache fix), this achieves 1.41 tok/s
on GLM-5.2 744B int4 / RTX 5070 Ti / 32GB RAM — a 4.3x speedup over stock.
route_agree=94.6% confirms minimal quality cost (94.6% of cache-steered
picks match the true top-K the router would have chosen).
No code change — CACHE_ROUTE is an existing engine feature, opt-in via
env vars. Documented in issue_diskio.md with the full optimization journey.
The cap_for_ram reserve for the 64-slot expert working set (ws_b = 64 × eb
= 1.21 GB) is overcounted when EXPERT_BUDGET is active. At budget=4 only
ws[0..3] are populated (not all 64), so the actual working set is 4 × eb
= 76 MB — 16x less than reserved. The excess 1.06 GB was starving the LRU
cache, capping it at 3 when budget=4 needs cap>=4.
Fix: clamp ws_b to (budget+4) × eb when EXPERT_BUDGET < 64. This raises
cap from 3 to 4, matching the budget. The cache can now hold all experts
a token needs, eliminating the LRU thrashing that caused excessive disk
re-reads (the SSD hammering).
Measured (pipe2 + full stack, budget=4, RAM_GB=28):
tok/s: 0.85 -> 1.03 (+21%)
hit rate: 57% -> 73% (+28%)
expert-disk: 18.2s -> 12.4s (-32%)
decode: 37.7s -> 31.0s (-18%)
Correctness: 32/32 oracle positions.
Same fix applied to expert_avail() (the mirror function for pin budgeting).
The HWINFO snapshot read /proc/cpuinfo, /proc/meminfo and
sysconf(_SC_NPROCESSORS_ONLN) — all absent on native Windows — so the
dashboard runtime panel showed "0 GB RAM / 0 cores" while the CUDA branch
populated VRAM fine. Now: CPUID brand string (0x80000002..4, no new
deps), GetSystemInfo for logical cores, and the existing compat_meminfo
(GlobalMemoryStatusEx) for RAM. Verified live: panel reads "AMD Ryzen 9
9950X3D 16-Core Processor / 66 GB RAM 24 GB free / 32 cores".
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
On Windows the engine self-exec OMP tuning never runs (Linux/FreeBSD-only)
and posix_fadvise readahead is a compat.h no-op, so a stock Windows run
leaves large measured wins on the table. The launcher now setdefaults, on
win32 only, each independently overridable by setting the variable:
- OMP_WAIT_POLICY=active, GOMP_SPINCOUNT=200000, OMP_DYNAMIC=FALSE,
OMP_NUM_THREADS=<physical cores> (parity with the glm.c self-exec block;
COLI_NO_OMP_TUNE disables exactly this block, presence-based like the
engine). OMP_PROC_BIND/OMP_PLACES deliberately omitted and also removed
from environment_for_plan on win32: MinGW libgomp has no affinity support
("Affinity not supported on this configuration").
- DIRECT=1: unbuffered expert reads. Measured on a 9950X3D + Samsung 9100
PRO Gen5 + Win11: iobench 10.68 GB/s O_DIRECT vs 9.03 buffered (warm);
end-to-end REPLAY 0.48 -> 1.02 tok/s. Matches #162 (1.47x same class).
- PIPE=1: load/matmul overlap, byte-identical output; +8% on top of DIRECT
(PIPE_WORKERS untouched at 8 - 4/8/16 swept flat on Gen5).
- PILOT_REAL=1: real cross-layer prefetch, the only working prefetch on
Windows; +11% and expert hit rate +19 points.
Full ladder methodology and numbers: 96-token greedy REPLAY, one lever per
step, medians of 3-4 runs (see the fork tuning doc referenced in the PR).
tests/test_env_defaults.py covers the defaults, explicit-override-wins,
the kill-switch scope, and the non-win32 no-op.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Expert count 21,504 -> 19,456 (75x256 + MTP head), matching the rest of the repo.
- Replace phantom IO_THREADS with PIPE_WORKERS (default 8); state the pool only
engages under PIPE=1.
- Drop rotting precise line-count claims (glm.c ~2,400; web ~390) — keep the prose.
- Add the desktop/ directory to the repo-layout section.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
profile_print now emits 'expert-disk N.NNNs service / N.NNNs wait' but
the harness regex still expected the old single expert-disk number, so
every run died with 'benchmark output missing'. Accept both formats and
report disk = service + wait.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
os.cpu_count() returns logical processors, so on SMT machines the plan
sets OMP_NUM_THREADS to 2 threads/core, which thrashes the AVX-512 units
during expert matmul (9950X3D: 32 logical vs 16 physical). Count
RelationProcessorCore records instead, with the existing lscpu/cpu_count
fallbacks intact.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One-line gate change: relax the pipe2 call-site gate from S>=8 to S>=1,
allowing the existing resident-pipeline (pipe_layer_sparse) to run during
single-token decode, not just prefill.
The S>=8 gate was a performance heuristic (prefill-only), not a correctness
constraint — pipe_layer_sparse is fully S-general. At decode it keeps the
residual stream x on the GPU device across all 78 layers, running rmsnorm,
residual adds, and shared-expert matmuls on-device. This eliminates the
~12.5k GPU sync interruptions per decode that caused the expert-matmul
regression (13.3s -> 9.2s), and moves the untracked 'other' CPU work
(rmsnorms, residual adds, routing) onto the GPU (29s -> 17.6s).
Measured (GLM-5.2 744B int4, RTX 5070 Ti, 32GB RAM, budget=4 + full disk stack):
tok/s: 0.72 -> 1.07 (+49%)
decode: 44.5s -> 29.9s (-33%)
expert-matmul: 13.3s -> 9.2s (regression fixed)
'other': 29s -> 17.6s (-39%)
Correctness: 32/32 oracle positions (3 consecutive runs).
Configuration: COLI_CUDA=1 CUDA_DENSE=1 COLI_CUDA_ATTN=1 COLI_CUDA_PIPE=2
CUDA_EXPERT_GB=0 EXPERT_BUDGET=4 PIPE=1 RAM_GB=28 PILOT_REAL=1 DIRECT=1
vmmlaq_s32 computes a 2x2 int32 tile (2 weight rows x 2 activation
rows) per instruction on 8-deep segments. Tile o and s in pairs,
halving weight traffic and doubling per-instruction work at S>=2.
Four independent accumulators over a 64-deep unroll keep the loop
throughput-bound (a single chained accumulator measures no better
than SDOT: latency-bound). S=1 and all tails (odd o, odd s, I not a
multiple of 16/32) keep the existing SDOT/scalar code, and scales
apply in the same order, so results are bit-identical.
Compile-time gated on __ARM_FEATURE_MATMUL_INT8. The default Darwin
build passes no -mcpu and is byte-identical (still SDOT, IDOT_KERNEL
"neon"). Opt in with ARCH=native (new Darwin Makefile knob, appends
-mcpu=<arch>), which reports IDOT_KERNEL "neon-i8mm". The same gate
lights up on any aarch64 with i8mm (Graviton3+, Grace).
test_idot grows a driver-level exactness check through matmul_qt_ex:
fmt 1 and 2, S in {2,3,4,5,8}, O in {1,2,3,64,65}, I in {16,17,100,
1408}, bitwise float equality against a plain-C reference. Green on
both build flavors.
Measured on an M5 Pro (18 threads, matmul_qt_ex microbenchmark at
GLM-5.2 expert shapes, best of 3 process runs, vs the SDOT baseline):
gateup int4 S=8 499.7 -> 1076.6 GF/s (+115%)
gateup int8 S=8 512.9 -> 1166.3 GF/s (+127%)
down int4 S=8 696.9 -> 1186.0 GF/s (+70%)
S=1 decode rows unchanged (SDOT path untouched)
The decode/prefill PROFILE line splits expert-disk time into service
(overlapped async dispatch) and wait (blocking stalls), but every
accumulation site wrote t_edisk and nothing ever wrote t_ewait, so the
wait column always printed 0.000s. Worse, the accounted sum used the
dead t_ewait instead of t_edisk, so the entire disk-read stall was
excluded from accounted and silently fell into the other bucket.
On a disk-streaming MoE the effect is large: other reads as ~60% of
decode when it is really the expert-load stall. Route the three
blocking sites (non-PIPE parallel load, the Metal drain barrier, and
the per-expert pipe_wait in the CPU matmul loop) to t_ewait, keep the
async dispatch in t_edisk, and include both in accounted.
Profiling-only; no behavioural change. Before, on a 168-expert model:
expert-disk 25.077s service / 0.000s wait | ... | other 28.401s
After:
expert-disk 0.111s service / 26.948s wait | ... | other 3.390s
other now holds only the genuinely-unbucketed work (router, norms).