Commit Graph

438 Commits

Author SHA1 Message Date
JustVugg d34461c8d8 Merge remote-tracking branch 'origin/dev' into pr270
# Conflicts:
#	c/Makefile
2026-07-20 17:51:50 +02:00
Vincenzo Fornaro d3362303f5 Merge pull request #424 from ZacharyZcR/feat/website
site: official website — animated demo, expert brain, 3-D atlas + GitHub Pages deploy
2026-07-20 17:42:08 +02:00
Vincenzo Fornaro ff38be0fee Merge pull request #330 from nbeerbower/tok-o200k
tok: o200k pre-tokenizer support, auto-detected, with download-free test coverage
2026-07-20 17:41:40 +02:00
JustVugg 48d4af691d Merge remote-tracking branch 'origin/dev' into pr330
# Conflicts:
#	c/Makefile
2026-07-20 17:41:10 +02:00
Vincenzo Fornaro f8e0612f26 Merge pull request #363 from woolcoxm/fix/win32-coli-chat-cuda-autoenable
win32: auto-enable the GPU in bare coli chat
2026-07-20 17:39:09 +02:00
JustVugg b347092e5c Merge remote-tracking branch 'origin/main' into dev 2026-07-20 17:29:44 +02:00
Vincenzo Fornaro 85e7dffb85 Merge pull request #460 from JustVugg/docs-windows-exe
docs: explain the prebuilt Windows binary — what the .exe is and how to run it (#450)
2026-07-20 17:24:15 +02:00
JustVugg 7de49fa02d docs: explain the prebuilt Windows binary — what the .exe is and how to run it (#450)
The release ships colibri-<ver>-windows-x86_64.zip but nothing said what the
.exe is for or how to use it. The quickstart's Windows Option A now lists the
zip contents (engine .exe / coli launcher / Python support), and gives the two
missing steps: rename the engine to glm.exe so the coli launcher finds it, and
install Python 3 for the launcher/gateway. README's run section gets a short
Windows-prebuilt pointer to the same. Docs-only.

Closes #450

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 17:23:21 +02:00
Vincenzo Fornaro 505a0f69aa Merge pull request #440 from okuvshynov/test_macpro_2019
fix: old x86 Mac scalar -> vnni
2026-07-20 17:20:31 +02:00
Vincenzo Fornaro 86766bf989 Merge pull request #427 from DavidePapero/main
Added Dockerfile and instructions for using Colibrì through a docker container
2026-07-20 17:20:26 +02:00
Vincenzo Fornaro 0337697023 Merge pull request #332 from woolcoxm/fix/auto-tier-single-core
resource_plan: fix --auto-tier pinning decode to one core (#325)
2026-07-20 17:19:50 +02:00
Vincenzo Fornaro af23f314fd Merge pull request #360 from woolcoxm/test/efficiency-suite
tests: efficiency suite — tiny-model regression gates + full-model optimization dossier (#359)
2026-07-20 17:09:21 +02:00
Vincenzo Fornaro 1765ed80ac Merge pull request #453 from ZacharyZcR/feat/ablation-iq3-scheme
tools: IQ3_XXS-codebook scheme in the quant ablation — settles #452's codebook decision
2026-07-20 17:09:06 +02:00
Vincenzo Fornaro bb2bae8904 Merge pull request #456 from monotophic/convert/resume-guard
convert: mirror the --indir resume params guard onto the --repo download loops
2026-07-20 17:08:50 +02:00
Vincenzo Fornaro 3ddbc655ef Merge pull request #437 from 4ny3l/fix-tool-calling-401
fix(serve): lift non-EOS stop tokens in C and fix redundant tools parsing in Python (#401)
2026-07-20 17:07:34 +02:00
JustVugg 3455eeea0a Merge remote-tracking branch 'origin/dev' into pr437
# Conflicts:
#	c/colibri.c
2026-07-20 17:04:18 +02:00
woolcoxm 0b05f6a76b win32: auto-enable the GPU in bare coli chat
On Windows a bare 'coli chat' (no --gpu/--vram/--auto-tier) ALWAYS ran
CPU-only, even on a CUDA build with a GPU present. Two defects:

1. cuda_binary() returned False on Windows. It detects CUDA by running
   'ldd glm | grep libcudart', which is Linux-only (no ldd on win32) and
   meaningless anyway because the Windows engine links cudart only inside a
   runtime-loaded coli_cuda.dll, not as a libcudart symbol in glm.exe. So
   the --gpu/--vram/--auto-tier gates (which call cuda_binary()) never opened.

2. Even with detection fixed, bare 'coli chat' set no CUDA env: env_for's
   else-branch only enables CUDA when --gpu/--vram is passed. Nothing
   auto-enabled the GPU.

Now: cuda_binary() on non-Linux returns True iff coli_cuda.dll exists next
to glm.exe — the exact file backend_loader.c loads from the engine's own
directory, so its presence is a faithful, cheap, DLL-hijack-safe proxy for a
CUDA-capable build. And env_for, scoped to win32 (Linux keeps its working
explicit-flag UX), auto-enables CUDA when a bare chat detects a CUDA build
plus a GPU via nvidia-smi, sizing the expert-tier VRAM budget from real free
VRAM via the existing build_plan/environment_for_plan machinery (same as
--auto-tier, no guessed budget). If nvidia-smi is missing it falls back to
CPU with a clear warning; --gpu none still forces CPU; explicit --vram/--gpu
still win. CUDA_DENSE stays an explicit opt-in (matches --auto-tier).

Verified on a Windows + RTX 5070 Ti box: bare 'coli chat --model <g64>' now
prints '[GPU] auto-enabled CUDA ... 13.0 GB expert tier' and emits
COLI_CUDA=1 / COLI_GPUS=0 / CUDA_EXPERT_GB=13.044 (was: all unset, CPU-only).

Tests: 4 new cases (auto-enable, nvidia-smi-missing fallback, CPU-build
silent, Linux-unchanged) plus the 4 existing default-I/O tests guarded to
mock cuda_binary() so they stay host-independent. Full python suite green
(env_defaults 8, resource_plan 10, doctor 8, makefile_platform 3, cli_output 3).

Out of scope: doctor.cuda_linkage is also POSIX-only and mis-reports on
Windows — separate follow-up.
2026-07-20 11:01:50 -04:00
woolcoxm 171aa69fcb fix(resource_plan): take last two lscpu fields, not [1]/[2] (#332 review)
JustVugg caught a regression by running the PR: the code asked
`lscpu -p=CPU,Core,Socket` and read fields[1]/fields[2], but the comment's
claim was inverted -- `lscpu -p=<list>` emits EXACTLY the requested columns
(no CPU prefix), while bare `lscpu -p` prepends CPU.

On machines where the requested list is short or lscpu collapses to two
columns, every line was skipped by the `< 3` guard, cores stayed empty, and
the code fell through to os.cpu_count() -- the LOGICAL count. Result: 6
physical cores reported as 12 (SMT over-subscription), the opposite of the
fix.

Correct per review:
  - ask lscpu for exactly 'core,socket'
  - take fields[-2], fields[-1] (correct whether or not CPU is prepended)
  - keep the warning scaffolding + the _resolve_physical_cores clamp
    (those are the actual #325 fix; only the indexing was wrong)

Test now exercises BOTH the 2-column (-p=core,socket) and 3-column
(bare -p, CPU prefix) layouts, asserting 12 physical cores each. The
2-column case is the one that regressed: old parser returns 24, new
returns 12.
2026-07-20 10:56:54 -04:00
woolcoxm bd3efb1c97 resource_plan: stop setting OMP_PROC_BIND/OMP_PLACES from --auto-tier (#325)
The core-count fix was necessary but not sufficient: @liangstein confirmed
physical_cpu_count() now returns 64 on his box, yet --auto-tier STILL pinned
decode to one core. A complete env diff between the plain (working) and
--auto-tier (broken) paths showed exactly three keys the plan adds:

  OMP_NUM_THREADS = 64   (correct, verified)
  OMP_PROC_BIND   = spread
  OMP_PLACES      = cores

Since OMP_NUM_THREADS was already correct, the culprit is the affinity pair.
The mechanism: environment_for_plan() sets OMP_PROC_BIND=spread + OMP_PLACES=cores
in the launcher's env. The engine's hot-thread tuning (glm.c main, the
COLI_OMP_TUNED self-exec) then tries setenv("OMP_PROC_BIND","close", overwrite=0)
-- but overwrite=0 cannot replace an already-set var, so the plan's "spread"
wins. On the reporter's libgomp + multi-socket topology, spread + places=cores
collapsed the team to a single CPU even with 64 threads configured.

Fix: don't set affinity from the plan at all. The engine deliberately chose
"close" for cache locality (the tiny back-to-back per-expert matmuls want
adjacent cores), and the plain path already leaves affinity to the engine.
Removing the plan's spread/places makes --auto-tier match the working plain
path; a user wanting a specific policy can still set OMP_PROC_BIND/OMP_PLACES
in their own environment (environment_for_plan only setdefaults OMP_NUM_THREADS).

Verified via env diff: after the fix, --auto-tier adds ONLY OMP_NUM_THREADS
beyond the plain env -- the engine's own close/bind tuning now wins on both
paths identically.

Note: this could not be reproduced on Windows (MinGW libgomp prints "Affinity
not supported on this configuration" and ignores the vars entirely); it is
Linux-libgomp-specific, matching the reporter's Rocky 9 box.

Tests: rewrite test_applies_plan_without_overriding_explicit_settings to assert
the plan sets NO affinity vars on any platform (the old test encoded the buggy
platform-dependent spread/cores contract). Add
test_plan_does_not_set_omp_affinity_vars as a focused regression. 78/78 pass.
2026-07-20 10:56:54 -04:00
woolcoxm f27a89ea82 resource_plan: fix --auto-tier pinning decode to one core (#325)
physical_cpu_count() silently returned 1 in two situations, and that value
flowed through build_plan -> OMP_NUM_THREADS to pin every matmul region to a
single thread under --auto-tier (reported on Rocky Linux 9, 512 GB RAM).

Two root causes:

1. The lscpu parse counted the wrong thing. `lscpu -p=core,socket` prepends the
   CPU column, so the output is actually CPU,Core,Socket; the old set
   comprehension collected (CPU,core,socket) tuples that were unique per logical
   CPU. Now parse CPU,Core,Socket and dedupe on (core, socket) to get true
   physical cores (the SMT-doubling the surrounding comments warn against was
   the actual behavior).

2. Any probe failure fell through to `os.cpu_count() or 1`. On a cgroup'd or
   otherwise constrained box os.cpu_count() can be 1 (or None), silently
   capping the run. Skip offline core/socket fields ("-" instead of raising
   ValueError) so a single offline row no longer discards the whole parse, and
   replace the silent `or 1` fallback with os.cpu_count() (logical) plus an
   explicit warning. Only return 1 when nothing at all is detected, and warn.

Also harden the win32 branch: declare argtypes/restype on
GetLogicalProcessorInformationEx (an undeclared 64-bit WinAPI returns c_int and
takes c_int pointers, so the probe could silently fail), and warn on its
fallbacks. Replace the silent max(1, int()) clamp in build_plan with
_resolve_physical_cores() that clamps to 1 with a warning instead of masking.

Tests: the existing OMP_NUM_THREADS test passed physical_cpus=24 explicitly, so
it never exercised the real probe -- that is why this regressed. Add regression
coverage for the lscpu physical-core dedup, offline fields, lscpu-missing
fallback, zero-cores degenerate case, and an end-to-end build_plan +
environment_for_plan check that OMP_NUM_THREADS reflects physical (not logical)
cores.
2026-07-20 10:53:26 -04:00
woolcoxm cf126cbcf9 fix(tests): skip efficiency tests when glm_tiny fixture is absent (CI #360)
CI runs 'make check' = dependency-free tests, no model downloads (by design,
#140). glm_tiny/ is a gitignored generated fixture, so test_inefficiency.py
hard-failed on the Windows/macOS/Linux runners with 'config.json: No such file
or directory' instead of skipping.

_engine_present() now requires BOTH glm.exe AND glm_tiny/config.json, and
_skip_reason() names exactly which prerequisite is missing so the skip is
actionable. Verified: 8 skipped (0 failed) with the fixture absent; 5 pass +
3 CUDA-skip with it present.
2026-07-20 10:49:07 -04:00
woolcoxm d43b54534f tests: efficiency suite — tiny-model regression gates + full-model optimization dossier
Two layers of efficiency coverage for the engine, both parsing the telemetry
glm.c already emits (REPLAY/PROFILE/[PROF]/CUDA-tier) but nothing previously
asserted on:

1. test_inefficiency.py — tiny-model asserted regression tests (8 tests, run
   in make test via test-python). Gate on: throughput floor, PROFILE phase
   accounting sanity, disk-wait not dominant on a resident model, CPU greedy
   determinism, and (when a CUDA build is present) CUDA init, dense VRAM
   upload, and CPU-vs-CUDA argmax agreement >= 70%. CUDA tests auto-skip with
   a clear build hint on CPU-only binaries.

2. test_efficiency_report.py — opt-in optimization dossier for a real model.
   Turns on every instrumentation flag (PROF, COLI_CUDA_PROFILE, CACHE_ROUTE,
   DISK_SPLIT, LOOKA) and prints 9 sections (provenance, throughput + tail
   latency, where-time-goes, attention breakdown, expert cache, disk I/O +
   phase split, routing quality + predictability, speculation, GPU tiers),
   each flagging inefficiency with the concrete knob to move tok/s. Never
   fails CI.

tools/efficiency.py is the shared harness: parse_run() captures every signal,
run_engine() wraps the subprocess. Reuses PROFILE_RE/SPEED_RE from
tools/benchmark_cuda_fixture.py and extends the tok/s regex to also catch the
run_text (parenthesized) format the full-model PROMPT path uses.

Makefile adds: efficiency / efficiency-cuda / efficiency-report targets.

Verified end-to-end on the full glm52_i4_g64 model (CPU + CUDA).
2026-07-20 10:49:07 -04:00
monotophic 42a5417c13 convert: mirror the --indir resume params guard onto --repo download loops
Upstream courtesy fix, found while auditing the conversion recipe for this
branch -- pre-existing in dev, not introduced by int3-g64, but directly
relevant to anyone converting for real (a #383-class gap: same failure
family as the resume/manifest work already done for --indir, and the
silent-mixing mechanism issue #355 fixed for a narrower case).

The --indir path already refuses to resume with different conversion
parameters on the same --outdir (a manifest records ebits/xbits/io_bits/
group_size/n_layers/bits_map and compares on every resume). The --repo
streaming download loops (main model, --mtp, --indexer) never got the
same guard: each shard's resume check is just `if os.path.exists(outp):
continue` -- true whether or not THIS run's flags match the flags that
produced that shard. A --repo conversion resumed with changed bits
(--xbits 3 -> 4 mid-run, say, after an interruption) would silently mix
bit-widths across shards in the same container, with no error and no log
line distinguishing it from a normal resume.

Fix: check_or_record_params(), a small shared helper mirroring the
--indir manifest's refuse-on-mismatch logic but without needing its
per-shard bookkeeping (the --repo loops already track shard completion
correctly via out-NNNNN.safetensors existence, since shard index maps
directly to output filename there -- only whether the params used SO FAR
still match needed adding). Applied to all three --repo loops with
per-mode sidecar files (.out-mtp-params.json / .out-idx-params.json /
.out-params.json) so a --mtp and a main-model conversion into the same
--outdir don't cross-check each other's parameters.

Also includes PROJ_BITS (the per-projection expert bit overrides) in the
tracked params dict on BOTH paths -- it was missing from --indir's
existing manifest too, so a resume with a changed --up-bits/--gate-bits/
--down-bits would have passed the existing guard silently.

Verified directly (no real HF downloads; --repo network paths can't be
exercised under this task's constraints): unit-tested
check_or_record_params() standalone -- fresh outdir accepts and records,
a same-params resume accepts, a changed --xbits is refused, and a
proj_bits-only change (nothing else different) is refused. Re-ran the
existing --indir dry-run end to end (convert, resume, resume-with-changed-
xbits) to confirm the manifest-based path still works correctly with
proj_bits added to its params dict.

Gates: make test-c (20/20) and make test-python (85/85) both pass.
2026-07-20 09:53:19 -04:00
ZacharyZcR ae4e31a15a tools: IQ3_XXS-codebook scheme in the quant ablation (#452 step 1)
Adds `-iq3` to the ablation harness: a faithful torch model of llama.cpp's
deployed 3.06-bpw IQ3_XXS format — 256-entry 4-dim magnitude grid
(extracted from ggml-common.h, MIT), signs factored per 8 weights with
the odd-parity constraint priced in (a violating block flips its
smallest-magnitude sign), fp16 super-scale per 256 + 4-bit sub-scale per
32 searched over all 16 codes. Nearest-grid search runs as chunked
matmul-argmin (|g|^2 - 2 q.g) — a full cdist materializes tens of GB on
a 100M-param tensor and OOMed the first run.

Measured (OLMoE-1B-7B, n=200 x hellaswag/arc/mmlu; the first four rows
reproduce the published ablations exactly):

  fp16                58.0%
  int4 per-row        48.7%   (-9.3pp, the shipped container's scheme)
  int3-g64            50.5%   (-7.5pp)
  int3-g64-e8-rot     51.5%   (-6.5pp, simulated rate-scaled ball)
  int3-iq3            49.3%   (-8.7pp)
  int3-iq3-rot        51.5%   (-6.5pp)

The deployable IQ3 codebook plus rotation exactly ties the simulated E8
ball — that settles #452's codebook decision toward the IQ3-style block
structure, with rotation mandatory (worth 2.2pp on this codebook).
2026-07-20 12:55:44 +08:00
Vincenzo Fornaro e9b36141a4 Merge pull request #449 from cdhdt/fix/rss-guard-uaf
rss_guard: free the evicted expert slab under g_pilot_mx (fixes a use-after-free with the pilot prefetcher)
2026-07-20 02:29:06 +02:00
cdhdt 9c5ab39b62 rss_guard: free the evicted expert slab under g_pilot_mx (use-after-free)
rss_guard marked the victim slot eid=-1 under the lock, unlocked, and only then
freed s->slab. In that window the slot reads as {eid=-1, slab still valid} --
exactly the state pilot_realload's victim scan reuses first -- so a pilot worker
could claim it and pread into the slab while it was being freed: use-after-free,
or a double-free if the loader took its own realloc path.

Keep the free and the pointer/capacity NULLing inside the critical section, so
'slab valid' and 'slot reusable' are never simultaneously observable. The lock is
held across a free() (microseconds); workers only take it briefly for scan+reserve.

Reproduces on dev with a SINGLE pilot worker (rss_guard runs on the main thread
while the pilot worker runs in the background), whenever the RAM guard is active
(RSS_GUARD_GB, or any resolved g_ram_budget_gb).

make check 83/83, native + portable builds 0 warnings.
2026-07-20 01:27:28 +02:00
Vincenzo Fornaro 4b704823db Merge pull request #445 from ZacharyZcR/fix/pin-budget-release-host
pin: with CUDA_RELEASE_HOST the VRAM prefix must not consume the RAM pin budget — +72% on 6×5090
2026-07-20 00:03:13 +02:00
ZacharyZcR 453d1401ba pin: with CUDA_RELEASE_HOST the VRAM prefix must not consume the RAM pin budget
pin_load computed npin from the RAM budget FIRST, then carved the
VRAM-ranked prefix out of those npin slots. With CUDA_RELEASE_HOST the
prefix's host slabs are freed right after upload — so on a multi-GPU
host the top-ranked experts consumed the RAM budget without occupying
RAM, and the CPU tier pinned only the leftovers. Measured on 6x RTX 5090
(251 GB): 9,280 VRAM + only 1,721 RAM pins (32.5 GB warm) on a box whose
RAM fits ~10k more — the cold tail then paid disk on every token and the
hit rate ceilinged at 99.0% forever.

Move the VRAM budget estimate above the npin finalization and make the
release-destined prefix ADDITIVE to the RAM-derived count. Same box,
same env plus the fix:

    [PIN] placement: 9,280 VRAM + 10,176 RAM (192.5 GB warm)
    expert hit rate 100.0% (pin 100.0% + lru 0.0%)  — disk 0

    1,024-token greedy decode: 3.62 -> 6.21 tok/s  (+72%)
    warm late segment (t=768-1024): 5.11 -> 5.73 tok/s  (+12%)
    prefill: 10.4 -> 8.8 s

The whole-run gain is the LRU warmup penalty disappearing (full
residency from the first token); the late-segment gain is the steady
~0.9%-miss disk residue recovered. Non-release configs are untouched:
prefix_est stays 0 and the arithmetic reduces to today's exactly.
2026-07-20 05:15:45 +08:00
Vincenzo Fornaro 17fbdfa8f8 Merge pull request #393 from ZacharyZcR/feat/auto-tune
plan: auto-tune heuristics — MTP/PIPE/NUMA decisions from bottleneck classification
2026-07-19 22:40:53 +02:00
Vincenzo Fornaro a7ef058fc0 Merge pull request #438 from monotophic/e7/disk-class-instr
PROF: DISK-CLASS — per-load cold/warm disk instrumentation
2026-07-19 22:40:22 +02:00
Oleksandr Kuvshynov 5c84b25645 fix: old x86 Mac scalar -> vnni
Older macs still need -march, not -mcpu.

Before this fix, no ymm (or zmm, vnni, etc.).

Standalone repro:

```
% clang -O3  -mcpu=native glm.c -o /tmp/glm
clang: error: unsupported option '-mcpu=' for target 'x86_64-apple-darwin25.4.0'
% clang -O3  -mcpu=native glm.c -o /tmp/glm -lm
% otool -tv /tmp/glm | grep -c ymm
0
```

I'm not entirely sure why `-lm` would ignore `-mcpu=` rather than error, but either way we don't get vector instructions:

colibri itself:

```
% ulimit -l unlimited; OMP_NUM_THREADS=16 RAM_GB=700 PIN=auto PIN_GB=all PIN_FILL=1 ./coli run --ram 700 --model ~/projects/llms/glm5p2-i4 --ngen 256 "Explain the difference between concurrency and parallelism"
...
== GLM C engine (glm_moe_dsa), cache=8 experts/layer | experts@8-bit dense@8-bit | idot: scalar ==
```

After the fix, we'll get it:

```
% clang -O3  -march=native glm.c -o /tmp/glm -lm
% otool -tv /tmp/glm | grep -c ymm
2621
```

colibri run:
```
% ulimit -l unlimited; OMP_NUM_THREADS=16 RAM_GB=700 PIN=auto PIN_GB=all PIN_FILL=1 ./coli run --ram 700 --model ~/projects/llms/glm5p2-i4 --ngen 256 "Explain the difference between concurrency and parallelism"
...
== GLM C engine (glm_moe_dsa), cache=8 experts/layer | experts@8-bit dense@8-bit | idot: avx512-vnni ==
```
2026-07-19 16:40:09 -04:00
Vincenzo Fornaro fbabaa4544 Merge pull request #422 from nbeerbower/fix-mtp-probe-expert-count
fix: MTP-head probe uses last expert by index, not a hardcoded 255
2026-07-19 22:39:44 +02:00
monotophic 1f00142b25 PROF: DISK-CLASS — per-load cold/warm disk instrumentation (re-derived onto #391 split)
Re-derivation of e7/disk-class-instr @ de6dd6d (base caa49f7) onto
origin/dev @ 61004dc, after PR #391 split c/glm.c into c/colibri.c plus
quant.h/sample.h/kv_persist.h/telemetry.h/grammar.h. Semantic equivalence,
not a copy: same events, same accounting, re-sited onto the new tree.

Placement: colibri.c, unchanged from before the split. telemetry.h (#391)
is the dashboard/stats module (HWINFO/TIERS/EMAP/HITS protocol lines +
usage persistence) — prof_report(), expert_load_impl(), moe(), g_prof_io
and g_edisk_ns all stayed in colibri.c, so DISK-CLASS's aggregation and
printing follow them there. No relocation needed, no DEVIATION.

Read path: #362 (prefetcher-v3, 22509fc) turned out to be testbed-only
scope per its own merge message ("glm.c untouched") — confirmed zero
c/glm.c changes in that merge's diff. The seven expert_load() call sites
this patch touches (pipe_worker, expert_host_ensure, moe()'s OMP miss
loop, pilot_realload, repin_pass_limit, pin_load x2) are structurally
identical to the pre-split tree; the demand flag re-attaches at the same
sites with the same semantics (1 only at moe()'s own PIPE/OMP miss path,
0 everywhere else), so DISK-CLASS still counts demand loads only.

One real drift, unrelated to #362: #417 (cfcc742) fixed the exact "Metal
pre-routed FASE A never bumps the real elast/eaccess_clock" defect this
feature's comments described as a documented, deliberately-unfixed
upstream issue — the real clock now ticks in FASE A too. The private
elast_dc/eaccess_clock_dc clock is kept anyway: its job was never only
to route around that freeze, it also snapshots pre-bump state so a
call's own routing bump can't contaminate its own classification, and
keeping DISK-CLASS's bookkeeping fully separate from stock elast state
is what makes "byte-identical with PROF=0" provable by construction
instead of by argument. Code comments referencing the old defect are
updated to reflect the fix (historical note + #417/cfcc742 pointer)
rather than describing a bug that no longer exists.

Gates: make glm METAL=0 and METAL=1 both clean, zero warnings (matches
stock 61004dc, also built clean with zero warnings for comparison).
make test-c: 0 failures. make test-python: 77 tests, OK.

Authored by Fable 5 in Claude Code, analysis in partnership with
@Monotophic.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 15:42:57 -04:00
Colibri Developer 26bd8b403a fix(engine): filter non-EOS stop tokens in serve mode to prevent premature halt on tool calls
GLM-5.2's config defines three stop tokens: <|endoftext|>, <|user|>, and
<|observation|>. In serve mode, when the model generates <tool_call> blocks,
int4-quantized logit noise can cause argmax to pick a <|user|> or
<|observation|> token ID, immediately stopping generation.

The <|user|> and <|observation|> tokens are role markers handled by the
Python API server, not the C engine. Filter them out in stops_arm() when
SERVE=1, keeping only <|endoftext|> as the stop token.

- Add SERVE-mode guard in stops_arm() that retains only the EOS token
- Log the number of filtered tokens for diagnostics
2026-07-19 13:39:11 -06:00
Colibri Developer 70fb5b00f3 fix(api): pass tools and tool_choice as parameters to generation() to prevent duplicate extraction
The generation() method was re-extracting tools from the raw request body,
bypassing the filtering that render_chat() applies when tool_choice is a
function object (forced function call). This caused parse_tool_calls() to
receive the unfiltered tool list, producing incorrect or missing tool_calls
in the response.

- Add tools and tool_choice parameters to generation() signature
- Pass them from chat_completion() after render_chat() processing
- Add early structural validation of tools in generation_options()
  for clear HTTP 400 errors on malformed input
2026-07-19 13:38:57 -06:00
ZacharyZcR 22560ad15b README: vision, the JIT-for-weights framing, roadmap, acknowledgements
- The vision: open the model up — run it, study it, improve it
- The idea: explain the core algorithm as a JIT for weights — parameters
  as data staged across a heterogeneous hierarchy, learned from routing
- What's next: active placement/scheduling research; Kimi K2, Qwen3 MoE,
  MiniMax on the model roadmap
- Acknowledgements: Z.ai, Moonshot AI, Alibaba Qwen, MiniMax, Allen AI
2026-07-20 00:08:28 +08:00
ZacharyZcR bca4e95d92 site: redesign — measured atlas everywhere, vision, algorithm, models
- hero: the measured expert atlas as a full-screen slowly-turning backdrop
- atlas rebuilt on real data: canonical atlas v1 (721 canonical + 637
  gate-sensitive specialists, 10 measured categories) embedded inline;
  position IS the measured affinity vector, colour = top topic
- demo: third panel 'the atlas, live' — token routing flashes measured
  specialists for the active topic in both the brain grid and the galaxy;
  added SQL and Chinese-poetry turns so cluster shifts are visible
- profiles/ladder updated to current community numbers (#82 NUMA 9.0-9.2,
  #389 Xeon 1TB 5.42, #387 M5 Max 2.0, #161 GB10 3.33, #120 1.23);
  unpublished TTFT/hit values shown as em-dash, never invented
- new sections: vision manifesto (run/study/improve), 'A JIT, but for
  weights' algorithm explainer with tier stack, models roadmap
  (Kimi K2 / Qwen3 MoE / MiniMax planned) + open-weights acknowledgements,
  contribute cards
- visual pass: numbered sections, gradient type, glass panels, fixed nav
2026-07-20 00:08:28 +08:00
ZacharyZcR 3e45074bf2 site: official website — animated demo, expert brain, 3-D atlas, Pages deploy
A zero-build static site under site/, deployed to GitHub Pages by Actions:

- hero with the pixel hummingbird, key numbers, CTA
- 'watch it think': a chat replay paced at measured decode speeds
  (6x5090 / 128GB CPU / 5070 Ti / 25GB floor), with a live tok/s meter
  and the full 19,456-expert grid — colour = tier, brightness = heat,
  routed experts flash white per token
- the expert atlas as a draggable 3-D galaxy (measured-affinity clusters)
- three-tier explainer and the measured hardware ladder
- single HTML file, no dependencies, no build step; custom domain later
  is just a site/CNAME + DNS
2026-07-19 23:31:37 +08:00
Nicholas Beerbower 4586d33c60 fix: MTP-head probe uses the last expert by index, not a hardcoded 255
The has_mtp completeness probe checked for `mlp.experts.255.down_proj.weight`,
which only exists when n_routed_experts == 256. REAP-pruned checkpoints (and any
MoE with a different expert count) have fewer experts, so the probe spuriously
reported has_mtp=0 and disabled MTP speculative decode even though the head was
present and complete. Probe `mlp.experts.<n_experts-1>` instead.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 10:54:01 -04:00
Davide Quack 27b4c4263a Added Dockerfile and instructions for using Colibrì through a docker container 2026-07-19 16:49:29 +02:00
Vincenzo Fornaro 61004dcb84 Merge pull request #391 from ZacharyZcR/refactor/split-colibri
refactor: split glm.c → colibri.c + 4 header modules (−18%)
2026-07-19 15:29:40 +02:00
ZacharyZcR 8a9a0fca4d plan: auto-tune heuristics — bottleneck classification + parameter decisions
Extend resource_plan to classify the hardware into bottleneck regimes
(disk / memory / mixed / compute) and derive tuning knobs automatically:

  MTP:   off when compute-bound (42% loss at full residency, #389)
         or disk-bound with <90% hit (union growth adds reads)
  PIPE:  COLI_CUDA_PIPE=1 single-GPU, =2 multi-GPU, PIPE=1 CPU disk
  NUMA:  selective interleave for GPU hosts, blanket hint for CPU-only
  PIN:   PIN_GB=all when fully resident + no GPU
  OMP:   COLI_NO_OMP_TUNE=1 for Metal (spin steals GPU power)

`coli plan` now shows an auto-tune section with each knob and its
reason. `environment_for_plan()` applies them via setdefault so
explicit user settings always win.

plan version stays at 2 (additive fields: bottleneck_class,
projected_hit_rate, tune). 7 new tests covering all regimes.
2026-07-19 21:20:00 +08:00
ZacharyZcR 083fda5b0a fix: update test_logit_nan to include colibri.c instead of glm.c 2026-07-19 21:18:33 +08:00
ZacharyZcR bc69a9a6d0 fix: remove duplicate argmax_v — use NaN-safe version from sample.h 2026-07-19 21:15:54 +08:00
ZacharyZcR 93b4a8e78e ci: fix Windows CUDA DLL check — glm.exe → colibri.exe 2026-07-19 21:12:16 +08:00
ZacharyZcR e486574442 web: i18n support — English, 简体中文, 繁體中文, Italiano
Lightweight i18n without react-i18next: a LocaleProvider context +
useLocale() hook with {{var}} interpolation, auto-detect from
navigator.language, persisted to localStorage.

98 translation keys across 4 locale files (en/zh-CN/zh-TW/it).
All user-visible strings in App/Brain/Profiling/ErrorBoundary are
now t() calls. Language switcher added to the sidebar footer.

Build clean (tsc + vite), 18 tests pass.
2026-07-19 21:12:12 +08:00
ZacharyZcR 420a0720c3 docs: add simplified Chinese and Italian README, update language nav
New files:
  README.zh-CN.md — simplified Chinese (大陆用词)
  README.it.md    — Italian (the project's "mother tongue")

All four READMEs now link to each other in a consistent nav bar.
Updated zh-TW to reflect glm.c → colibri.c rename and new headers.
2026-07-19 21:12:11 +08:00
ZacharyZcR f853ea8a0b refactor: split glm.c into colibri.c + 4 header modules
Rename glm.c → colibri.c and extract four self-contained modules
into header-only files (same pattern as st.h/tier.h/grammar.h):

  quant.h      (672 lines) — SIMD matmul kernels, quantization
  sample.h     (143 lines) — RNG, top-p sampling, stop-set
  kv_persist.h (121 lines) — .coli_kv disk persistence
  telemetry.h  (189 lines) — dashboard protocol, stats, usage

Main engine file shrinks from 6588 to 5396 lines (−18%).

Build system: primary target is now colibri$(EXE); `make glm`
remains as a phony alias for backward compat. CI, setup.sh,
coli CLI, and all 10 test files that include the engine are
updated. make check passes (C + Python, 73 tests, zero warnings).
2026-07-19 21:12:04 +08:00
Vincenzo Fornaro 8f33bd153b Merge pull request #420 from JustVugg/p417-metal-lfru
metal: advance the LFRU recency clock on the GPU-prerouted decode path (#417)
2026-07-19 15:02:41 +02:00
JustVugg cfcc742591 metal: advance the LFRU recency clock on the GPU-prerouted decode path (#417)
On Metal, when routing is precomputed on the GPU (g_pre_idx), the moe fast
path bumps eusage/ehit/eheat for the selected experts but skips the one thing
the full CPU router does at the equivalent site: elast[layer][e] =
++eaccess_clock. So the session-local recency clock advances during prefill
(full router) but freezes the moment GPU-prerouted decode starts, and REPIN's
tier_pick_lfru() tie-breaker then runs on stale recency for the rest of the
run. Mirror the exact update the non-Metal path already does. Inside
#ifdef COLI_METAL, so CPU/CUDA are untouched; elast only feeds the LFRU
eviction heuristic, so this cannot affect output, only which experts REPIN
keeps warm.

Found and reported by @monotophic with a source-level trace repro
(ELAST_TRACE). Fix is inspection-verified against line ~3055; needs a
Metal build to exercise end-to-end.

Closes #417

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 14:58:32 +02:00