Rename glm.c → colibri.c and extract four self-contained modules
into header-only files (same pattern as st.h/tier.h/grammar.h):
quant.h (672 lines) — SIMD matmul kernels, quantization
sample.h (143 lines) — RNG, top-p sampling, stop-set
kv_persist.h (121 lines) — .coli_kv disk persistence
telemetry.h (189 lines) — dashboard protocol, stats, usage
Main engine file shrinks from 6588 to 5396 lines (−18%).
Build system: primary target is now colibri$(EXE); `make glm`
remains as a phony alias for backward compat. CI, setup.sh,
coli CLI, and all 10 test files that include the engine are
updated. make check passes (C + Python, 73 tests, zero warnings).
@KingIcyCreamProjects caught a real artifact on the 9950X3D (#357 thread): the
nk=8192 cell reported ~75x, far above the real ~13-40x algorithm. Root cause was
two compounding bench bugs, both verified against the code:
1. ONE frozen input per cell. partial_select_desc's median-of-three pivot is
fully deterministic (no RNG), and bench_dsa_select froze the input then ran
2000 reps on that identical array -- so all 2000 reps hit the exact same
pivot sequence. A single lucky input spiked one nk row (stable across re-runs
because it's deterministic, not because it's real).
2. brng was a static global, initialized once and NEVER reset between cells.
So each (shape,nk) cell's input depended on every prior cell's brand() draws
-- reordering nks[] or adding a shape silently shifted all later inputs.
Fix:
- brng_seed() reseeds per (shape, nk, seed) so every cell is reproducible and
independent of cell ordering.
- Report the MEDIAN of N_SEEDS=11 independent inputs, each itself a median over
N_REPEAT=2000 timing reps. A lucky pivot now moves one of 11 samples, not the
reported number.
Measured on Intel Core Ultra 9 185H (this fix, median of 11 seeds):
realistic@8192 10.97x (was 13.57x single-seed on this box)
uniform@8192 9.39x (was 7.48x)
plateau@8192 16.11x (plateau ignores the RNG, unchanged as expected)
band: 6-26x across all cells, no anomalous spikes.
Correctness is unaffected (test_dsa_select, 129 cases, unchanged). This is a bench
fidelity fix only -- no engine code touched.
The DSA lightning indexer selects the top-index_topk (2048) context keys to
attend to by finding the threshold = keep-th largest attention score. It
previously full-qsorted all nk scores per layer per token (O(nk log nk)) just
to read one pivot value, then scanned the original array in position order to
build the kept set.
Replace the qsort with partial_select_desc (Hoare quickselect, median-of-three,
descending): O(nk) average to partition the keep largest into a[0..k), then the
threshold is min of that block. The two position-order scans (>thr then ==thr)
are UNCHANGED, so the kept-position set is bit-identical -- a stronger contract
than #335's sampling heap (which was multiset-only because the heap was unstable
and changed accumulation order). The quickselect pivot IS by definition the
keep-th largest, so the new threshold equals old tmp[keep-1] exactly.
Measured (bench_dsa_select, keep=2048, median of 2000 reps):
nk=2049: 119us -> 5.7us (21x)
nk=8192: 626us -> 43us (15x)
nk=32768: 2.8ms -> 0.28ms (10x)
nk=65536: 6.6ms -> 0.47ms (14x)
The gap widens with context (linear vs n-log-n). DSA only fires past index_topk,
so this is precisely the long-conversation regime where decode latency matters.
Adds test_dsa_select (in TEST_BINS): 129 cases asserting element-wise identical
kept-set vs an independent qsort reference across shapes (random, peaked,
sorted, reverse-sorted, tie-plateau, all-equal) and edges (keep==1, keep==nk).
Also directly checks the partition invariant.
Adds bench_dsa_select (on-demand, NOT a gate): reproduces the table above.