The release ships colibri-<ver>-windows-x86_64.zip but nothing said what the .exe is for or how to use it. The quickstart's Windows Option A now lists the zip contents (engine .exe / coli launcher / Python support), and gives the two missing steps: rename the engine to glm.exe so the coli launcher finds it, and install Python 3 for the launcher/gateway. README's run section gets a short Windows-prebuilt pointer to the same. Docs-only. Closes #450 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
5.7 KiB
Quick Start — from zero to a running model
A step-by-step guide for first-time users on Linux, Windows, and macOS.
No prior experience with C, CUDA, or model conversion is assumed. If you get
stuck, ./coli doctor (below) tells you exactly what's missing.
What you're setting up: colibrì runs a very large Mixture-of-Experts model (e.g. GLM-5.2, 744B parameters) on a normal machine by streaming the model's experts from disk instead of needing them all in RAM. The engine is a single C program; Python is only used once, to prepare the model files.
0. What you need first (prerequisites)
| Minimum | Recommended | |
|---|---|---|
| RAM | ~16 GB | 24 GB+ |
| Free disk | ~380 GB for the int4 model | a fast NVMe SSD (streaming speed = your token speed) |
| OS | Linux, Windows 10/11, or macOS | any |
| Tools | a C compiler + make + git + python3 |
— |
You do not need a GPU. A GPU only helps if you have one; the engine runs CPU-only by default.
1. Install the build tools
Linux (Ubuntu / Debian)
sudo apt update
sudo apt install -y build-essential git python3
build-essential gives you gcc, make, and OpenMP (libgomp) — everything the
engine needs.
Windows
You have two options.
Option A — download a prebuilt binary (no compiler needed).
Grab colibri-<version>-windows-x86_64.zip from the
Releases page and unzip it.
Inside you'll find:
| File | What it is |
|---|---|
colibri-<version>-windows-x86_64.exe |
the engine — the C program that actually runs the model |
coli |
the command-line launcher (chat, serve, convert, doctor, …) |
openai_server.py, resource_plan.py, doctor.py |
Python support for the API server and placement planner |
Two setup steps:
- Rename the engine to
glm.exeso the launcher can find it (it looks for a binary namedglm):Rename-Item colibri-*-windows-x86_64.exe glm.exe - Install Python 3 from python.org — the
colilauncher and the API gateway are Python scripts (the engine itself is pure C and needs nothing).
Then continue to step 3. Prefer to skip the launcher? You can
run the engine directly — .\glm.exe reads the model path from the SNAP
environment variable (see docs/windows.md) — but coli chat is the
easy path.
Option B — build from source with MSYS2. Install MSYS2, open the UCRT64 shell, and run:
pacman -S --needed mingw-w64-ucrt-x86_64-gcc make git python
macOS
xcode-select --install # C compiler (clang)
brew install libomp git python # OpenMP for multithreading
2. Get the code and build the engine
git clone https://github.com/JustVugg/colibri.git
cd colibri/c
./setup.sh
setup.sh checks your compiler and OpenMP, builds the engine, and runs a tiny
self-test. When it prints:
engine self-test: 32/32 (expected 32/32)
the engine is working correctly. (On Windows Option A you already have the binary — you can skip this step.)
3. Get the model
You have two paths.
Easiest — download a ready-made int4 container
A pre-converted GLM-5.2 int4 model is on Hugging Face. Use the version with the int8 MTP heads (the plain int4 heads disable speculative decoding — see #8):
https://huggingface.co/mateogrgic/GLM-5.2-colibri-int4-with-int8-mtp
Download it into a folder on a fast disk, e.g. /nvme/glm52_i4 (Linux/macOS) or
D:\glm52_i4 (Windows). It is about 372 GB, so make sure you have the space.
Or convert it yourself from the FP8 source
One resumable command downloads and converts the model shard by shard, so it never needs the full ~756 GB on disk at once:
./coli convert --model /nvme/glm52_i4
This step uses Python and runs only once. Safe to interrupt and re-run — it resumes where it left off.
4. Run it
Point COLI_MODEL at the folder from step 3 and start chatting:
# Linux / macOS
COLI_MODEL=/nvme/glm52_i4 ./coli chat
# Windows (UCRT64 shell)
COLI_MODEL=/d/glm52_i4 ./coli chat
Useful first commands:
COLI_MODEL=/nvme/glm52_i4 ./coli doctor # read-only check: is everything ready?
COLI_MODEL=/nvme/glm52_i4 ./coli plan # shows where the model will live (RAM/disk/GPU)
COLI_MODEL=/nvme/glm52_i4 ./coli chat --topp 0.85 # faster: reads less from disk, same quality
Tip:
--topp 0.85is worth adding on a disk-bound machine — it reads fewer expert bytes per token with no quality loss, which directly means more tokens per second.
5. What to expect
- First launch loads the resident weights (~10 GB) — this takes a moment.
- Speed depends on your disk. The experts stream from storage, so a fast NVMe SSD is the single biggest factor in tokens/second. On a slow or shared disk, generation can be well under 1 token/second — that's expected, and it's the honest cost of running a 744B model on a small machine.
- It's still the full model. Placement only changes speed, never the model's answers or precision.
If something doesn't work, run ./coli doctor — it reports exactly what's
missing (compiler, model files, permissions) and how to fix it.
Where to go next
| Topic | Doc |
|---|---|
| Windows native build (and CUDA DLL) | docs/windows.md |
| Tuning: cache, prefetch, speculation | docs/tuning.md |
| OpenAI-compatible API + web dashboard | docs/api.md |
| Every environment variable | docs/ENVIRONMENT.md |