Files
colibri/c
NeuralNotwerk 97fa52698f glm: make MLOCK=1 actually wire pinned experts under COLI_MMAP
pin_wire() only locked the legacy private-slab path (s->slab / s->fslab),
which stay NULL under COLI_MMAP -- so MLOCK=1 on an mmap build silently
wired 0 bytes and the 'pinned' set was just warm page cache, evictable
under memory pressure (measured: hundreds of MB/s of re-reads from disk
mid-generation on a 503 GB host once the cache tightened).

Three pieces:
- qt_wire_mmap(): mlock each pinned expert's weight + scale ranges inside
  the file mappings. Skips cuda_eligible QTs: VRAM-tier experts compute
  from device memory, and expert_host_release() early-returns for mmap
  experts (no slab) WITHOUT nulling q8/q4, so a host-pointer check alone
  wires the whole VRAM tier too (~137 GB of never-touched locked pages on
  a 6x32GB-class rig -- enough to starve the kernel into thrashing).
- qt_unwire_mmap(): REPIN gpu_swap promotions drop the promoted expert's
  host lock instead of leaking it on every swap.
- expert_load() deliberately does NOT wire (it also runs for the transient
  VRAM-staging pass); wiring happens once in pin_wire() on the final set.

Measured on GLM-5.2 int4, 4x RTX 5090 + 1x RTX 4090, 503 GB RAM,
PIN_GB=all: wired goes 0 -> 226 GB (exactly the RAM tier), decode-time
disk reads drop to zero once warm.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 22:12:31 +00:00
..