- 'See it running' right after the intro: the live dashboard (744B
answering at 4+ tok/s end-to-end with metrics and tier bar) and the
Brain cortex (19,456 experts, tier colour, routing heat, per-turn
firing) — three seconds to understand what colibri is.
- A short table of contents for the long read below.
- 'Web dashboard' section: coli web one-liner, what each panel shows,
link to the Expert Atlas (#175). Existing content untouched.
Canonical write-up of the GRAMMAR= draft source: mechanism, why it pays in a
disk-streaming MoE specifically, usage/knobs, measured expectations by workload
shape (span-density dependence, from the #146 A/Bs incl. corrections), lossless
guarantees, bench discipline, prior art. Linked from the README feature bullet.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Docs-only. Documents that the OMP active-spin steals SoC power from the Metal GPU on Apple Silicon (default regresses -39%); COLI_NO_OMP_TUNE=1 + PIPE=1 recovers and beats the pre-rebase branch (2.24 vs 2.06 tok/s). Flags a follow-up: Metal builds should default to passive OMP wait.
Records the July 12-13 lab findings on one 6x RTX 5090 machine:
- vLLM-Moet comparison and why per-rank expert residency dominates
- colibri full-resident placement: 2.30 -> 6.28-6.84 tok/s, with the
tested-and-rejected directions documented
- MTP speculation: broken int4 head identified (issue #8), int8 head
reaches 69-79% acceptance but MoE verify batches scale expert cost
with S, so speculation loses at every depth; revisit after GPU
grouped GEMM
- AVX-512 int4 kernel qualification: numerically better than the old
order, quality-neutral (ppl delta 0.24%), +4-7% on CPU-heavy routing
README links the record from the honest-numbers section.