Files
JustVugg ec087f6bca docs: bring the Quick Start guide (incl. Windows prebuilt steps) to main
main shipped v1.0.0 with prebuilt Windows binaries but no docs on main
explaining what the .exe is or how to run it (#450) — the quickstart lived
only on dev. Bring docs/quickstart.md to main and add a Windows-prebuilt
pointer to the README's Get started section, so users on the default branch
find the steps. Docs-only.

Refs #450

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 17:36:06 +02:00

5.7 KiB

Quick Start — from zero to a running model

A step-by-step guide for first-time users on Linux, Windows, and macOS. No prior experience with C, CUDA, or model conversion is assumed. If you get stuck, ./coli doctor (below) tells you exactly what's missing.

What you're setting up: colibrì runs a very large Mixture-of-Experts model (e.g. GLM-5.2, 744B parameters) on a normal machine by streaming the model's experts from disk instead of needing them all in RAM. The engine is a single C program; Python is only used once, to prepare the model files.


0. What you need first (prerequisites)

Minimum Recommended
RAM ~16 GB 24 GB+
Free disk ~380 GB for the int4 model a fast NVMe SSD (streaming speed = your token speed)
OS Linux, Windows 10/11, or macOS any
Tools a C compiler + make + git + python3

You do not need a GPU. A GPU only helps if you have one; the engine runs CPU-only by default.


1. Install the build tools

Linux (Ubuntu / Debian)

sudo apt update
sudo apt install -y build-essential git python3

build-essential gives you gcc, make, and OpenMP (libgomp) — everything the engine needs.

Windows

You have two options.

Option A — download a prebuilt binary (no compiler needed). Grab colibri-<version>-windows-x86_64.zip from the Releases page and unzip it. Inside you'll find:

File What it is
colibri-<version>-windows-x86_64.exe the engine — the C program that actually runs the model
coli the command-line launcher (chat, serve, convert, doctor, …)
openai_server.py, resource_plan.py, doctor.py Python support for the API server and placement planner

Two setup steps:

  1. Rename the engine to glm.exe so the launcher can find it (it looks for a binary named glm):
    Rename-Item colibri-*-windows-x86_64.exe glm.exe
    
  2. Install Python 3 from python.org — the coli launcher and the API gateway are Python scripts (the engine itself is pure C and needs nothing).

Then continue to step 3. Prefer to skip the launcher? You can run the engine directly — .\glm.exe reads the model path from the SNAP environment variable (see docs/windows.md) — but coli chat is the easy path.

Option B — build from source with MSYS2. Install MSYS2, open the UCRT64 shell, and run:

pacman -S --needed mingw-w64-ucrt-x86_64-gcc make git python

macOS

xcode-select --install          # C compiler (clang)
brew install libomp git python  # OpenMP for multithreading

2. Get the code and build the engine

git clone https://github.com/JustVugg/colibri.git
cd colibri/c
./setup.sh

setup.sh checks your compiler and OpenMP, builds the engine, and runs a tiny self-test. When it prints:

engine self-test: 32/32  (expected 32/32)

the engine is working correctly. (On Windows Option A you already have the binary — you can skip this step.)


3. Get the model

You have two paths.

Easiest — download a ready-made int4 container

A pre-converted GLM-5.2 int4 model is on Hugging Face. Use the version with the int8 MTP heads (the plain int4 heads disable speculative decoding — see #8):

https://huggingface.co/mateogrgic/GLM-5.2-colibri-int4-with-int8-mtp

Download it into a folder on a fast disk, e.g. /nvme/glm52_i4 (Linux/macOS) or D:\glm52_i4 (Windows). It is about 372 GB, so make sure you have the space.

Or convert it yourself from the FP8 source

One resumable command downloads and converts the model shard by shard, so it never needs the full ~756 GB on disk at once:

./coli convert --model /nvme/glm52_i4

This step uses Python and runs only once. Safe to interrupt and re-run — it resumes where it left off.


4. Run it

Point COLI_MODEL at the folder from step 3 and start chatting:

# Linux / macOS
COLI_MODEL=/nvme/glm52_i4 ./coli chat

# Windows (UCRT64 shell)
COLI_MODEL=/d/glm52_i4 ./coli chat

Useful first commands:

COLI_MODEL=/nvme/glm52_i4 ./coli doctor   # read-only check: is everything ready?
COLI_MODEL=/nvme/glm52_i4 ./coli plan     # shows where the model will live (RAM/disk/GPU)
COLI_MODEL=/nvme/glm52_i4 ./coli chat --topp 0.85   # faster: reads less from disk, same quality

Tip: --topp 0.85 is worth adding on a disk-bound machine — it reads fewer expert bytes per token with no quality loss, which directly means more tokens per second.


5. What to expect

  • First launch loads the resident weights (~10 GB) — this takes a moment.
  • Speed depends on your disk. The experts stream from storage, so a fast NVMe SSD is the single biggest factor in tokens/second. On a slow or shared disk, generation can be well under 1 token/second — that's expected, and it's the honest cost of running a 744B model on a small machine.
  • It's still the full model. Placement only changes speed, never the model's answers or precision.

If something doesn't work, run ./coli doctor — it reports exactly what's missing (compiler, model files, permissions) and how to fix it.


Where to go next

Topic Doc
Windows native build (and CUDA DLL) docs/windows.md
Tuning: cache, prefetch, speculation docs/tuning.md
OpenAI-compatible API + web dashboard docs/api.md
Every environment variable docs/ENVIRONMENT.md