Files
colibri/docs/cuda.md
T

5.1 KiB
Raw Blame History

CUDA backend (Linux)

colibrì includes an opt-in CUDA backend for model-resident tensors. Streaming experts deliberately remain on the original CPU path: copying an expert from NVMe to the GPU on every use would only replace the disk bottleneck with a PCIe bottleneck. Resident quantized tensors are uploaded lazily once and reused.

cd c
make cuda-test CUDA=1                  # q8/q4/q2/f32 kernel correctness
make CUDA=1
# optional dense-path experiment (hot experts are configured below)
COLI_CUDA=1 COLI_GPU=0 CUDA_DENSE=1 SNAP=/nvme/glm52_i4 ./glm 64 4 4

Requirements: Linux, an NVIDIA driver, and a CUDA Toolkit under /usr/local/cuda (override with CUDA_HOME=/path/to/cuda). CUDA_ARCH=native builds for the GPU in the current machine. Requesting CUDA with a CPU-only binary, an invalid device, or an unavailable runtime fails at startup instead of silently falling back. For Windows, see windows.md (runtime DLL path).

The VRAM expert tier

A measured PIN profile promotes its hottest experts into a persistent VRAM tier while keeping the rest in RAM:

STATS=stats.txt SNAP=/nvme/glm52_i4 ./glm 64 4 4   # collect routing frequencies first
COLI_CUDA=1 COLI_GPU=0 CUDA_EXPERT_GB=16 \
PIN=stats.txt PIN_GB=160 SNAP=/nvme/glm52_i4 ./glm 64 4 4

# multi-GPU expert tier, 150 GB total budget across six 32 GB devices
COLI_CUDA=1 COLI_GPUS=0,1,2,3,4,5 CUDA_EXPERT_GB=150 \
CUDA_DENSE=1 PIN=stats.txt PIN_GB=300 RAM_GB=226 \
SNAP=/nvme/glm52_i4 ./glm 64 4 4

# large-RAM host: fill safe VRAM, then keep every remaining expert in RAM
COLI_CUDA=1 COLI_GPUS=0,1,2,3,4,5 CUDA_EXPERT_GB=auto \
CUDA_DENSE=1 COLI_CUDA_ATTN=1 PIN=stats.txt PIN_GB=all RAM_GB=auto \
SNAP=/nvme/glm52_i4 ./glm 64 4 4

Selected experts are uploaded during startup, so capacity failures occur before inference. The budget is clamped against free VRAM after reserving the projected dense resident set and 2 GB of runtime headroom per device. With COLI_GPUS, CUDA_EXPERT_GB is a total budget across the device set; experts are assigned whole to the least-loaded device that can hold them. Multi-GPU runs default to PIN_FILL=1 (measured hot set first, then unused VRAM filled with zero-heat experts) and CUDA_RELEASE_HOST=1 (RAM copy released after upload, reloaded from disk only if CUDA later fails).

CUDA_EXPERT_GB=auto fills each device up to measured free memory minus projected dense tensors and headroom. PIN_GB=all then loads the remaining routed experts into RAM up to the --ram budget (it clamps — #229), eliminating decode-time disk misses when capacity permits. This mode is intended for dedicated high-memory inference hosts.

Full-residency reference result (6× RTX 5090, 251 GiB host)

CUDA_EXPERT_GB=auto PIN_GB=all selected a 176.7 GB VRAM tier + 191.3 GB RAM tier (all 19,456 experts resident), adapting the VRAM tier every 16 tokens. With the GPU-resident pipeline (COLI_CUDA_PIPE=2) and Tensor-Core W4A16 dispatch (COLI_CUDA_TC_W4A16=1), 96-token greedy decode measured 5.86.8 tok/s (TTFT ~13 s; 1571-token prefill ~122 s then 4.2 tok/s). Full experiment log: experiments/glm52-6x5090-2026-07-12.md. These are host-specific capacity results, not portable defaults.

The GPU-resident pipeline (COLI_CUDA_PIPE)

COLI_CUDA_PIPE=2 keeps the residual stream on-device across layers: rmsnorms, residual adds, router GEMMs and the shared expert run on the GPU while the CPU expert loop runs uninterrupted, with batched attention and grouped expert uploads at prefill. On a single-GPU host this also pays at decode (S=1): +49% measured on a 5070 Ti (#273/#274); on multi-GPU hosts the per-layer P2P hops cancel the gain, so the decode gate is device-count aware. COLI_CUDA_TC_W4A16=1 enables Tensor-Core int4×fp16 mixed dispatch for batched rows (pays at ≥16 rows).

Notes and limitations

  • Text-mode timing reports prefill separately from decode.
  • MTP speculation defaults off on CUDA (cold draft routes increase expert traffic); explicit DRAFT=n overrides. Since #294, SPEC_PIN=1 keeps draft/verify kernels consistent when speculation is on.
  • Devices use independent contexts; a single expert is not sharded. Kernels are correctness-first custom kernels.
  • Profile quality matters more than raw VRAM capacity: the same 150 GB tier measured 0.941.64 tok/s hot-first vs 0.29 tok/s filled without routing heat.
  • The GPU tier earns its VRAM only when the CPU is the weak link — a tuned AVX-512 CPU can match a 5090 on expert matmul (#101).

Reproducible backend A/B without the full checkpoint

cd c
python tools/make_glm_bench_model.py --output /nvme/colibri-bench-medium --device cuda
python tools/benchmark_cuda_fixture.py --model /nvme/colibri-bench-medium --gpu 0

The 313M-parameter fixture has random weights and is not a language model. It preserves the real MLA/MoE/streaming shapes to compare CPU streaming, dense-only CUDA, CPU hot-store, and CUDA hot-expert execution with identical replay tokens.