05bba7994c
- c/version.py: single source of truth (__version__ = "1.0.0") - coli: reads version.py, banner shows dynamic version, --version flag - .github/workflows/release.yml: tag push triggers cross-platform build (Linux x86_64, macOS ARM64, Windows x86_64) and creates a GitHub Release with packaged binaries + changelog notes - CHANGELOG.md: v1.0.0 baseline documenting all shipped features To cut a release: 1. bump c/version.py 2. add a CHANGELOG section 3. git tag v1.0.0 && git push --tags
2.4 KiB
2.4 KiB
Changelog
All notable changes to colibrì are documented here. Format follows Keep a Changelog.
[1.0.0] — 2026-07-19
First tagged release. The engine has been running in production since late June 2026; this tag marks the baseline for semantic versioning going forward.
Highlights
- GLM-5.2 (744B MoE) runs on ~25 GB RAM in pure C, streaming experts from disk
- Three-tier placement: VRAM (hot) / RAM (warm) / NVMe (cold), with a learning cache that pins your workload's hottest experts automatically
- CUDA backend: multi-GPU expert tier, dense tensor distribution, batched
ragged attention, resident pipeline (
COLI_CUDA_PIPE=2) - Metal backend (Apple Silicon): batched expert SwiGLU + fused decode attention on unified memory GPU
- MTP speculation: native GLM-5.2 draft heads, grammar-forced drafts, kernel-
pinned verification (
SPEC_PIN=1) - OpenAI-compatible API:
coli servewith SSE streaming, KV slots, bounded queue, web dashboard (coli web) - Web UI: chat with live metrics, expert cortex brain page, profiling breakdown, expert atlas 3-D galaxy
- Cross-platform: Linux, macOS, Windows 11 (native MinGW), PowerPC; CI on all three
- Auto-tune:
coli plan --auto-tierclassifies the bottleneck and derives MTP/PIPE/NUMA/PIN settings with explanations
Engine
- Token-exact validation against
transformersoracle (teacher-forcing 32/32) - Compressed MLA KV cache (576 floats/token, 57× smaller), persisted across
restarts (
.coli_kv, zero re-prefill) - DSA sparse attention (lightning indexer), faithfully implemented
- Router-lookahead prefetch (
PILOT=1, 71.6% predictive) - Async expert I/O pool (
PIPE=1), io_uring batching (URING=1) - NUMA-aware expert placement (
COLI_NUMA=1, +13–40% on multi-socket) - AVX2 / AVX-512 / AVX-VNNI / ARM NEON / NEON-i8mm / POWER VSX kernels
- int4 / int8 / int2 / grouped-int4 (fmt=4) quantization formats
Tools
coli convert— FP8→int4 one-shard-at-a-time convertercoli doctor— read-only setup diagnosticscoli plan— resource planner with auto-tune prescriptioncoli bench— MMLU / HellaSwag / ARC quality benchmarks- Expert atlas (
tools/analyze.py --web) — measured topic affinity for 19,456 experts
Community
- 30+ hardware datapoints in the benchmark tracker
- Contributions from 20+ authors across engine, docs, tooling, and ports