97fa52698f
pin_wire() only locked the legacy private-slab path (s->slab / s->fslab), which stay NULL under COLI_MMAP -- so MLOCK=1 on an mmap build silently wired 0 bytes and the 'pinned' set was just warm page cache, evictable under memory pressure (measured: hundreds of MB/s of re-reads from disk mid-generation on a 503 GB host once the cache tightened). Three pieces: - qt_wire_mmap(): mlock each pinned expert's weight + scale ranges inside the file mappings. Skips cuda_eligible QTs: VRAM-tier experts compute from device memory, and expert_host_release() early-returns for mmap experts (no slab) WITHOUT nulling q8/q4, so a host-pointer check alone wires the whole VRAM tier too (~137 GB of never-touched locked pages on a 6x32GB-class rig -- enough to starve the kernel into thrashing). - qt_unwire_mmap(): REPIN gpu_swap promotions drop the promoted expert's host lock instead of leaking it on every swap. - expert_load() deliberately does NOT wire (it also runs for the transient VRAM-staging pass); wiring happens once in pin_wire() on the final set. Measured on GLM-5.2 int4, 4x RTX 5090 + 1x RTX 4090, 503 GB RAM, PIN_GB=all: wired goes 0 -> 226 GB (exactly the RAM tier), decode-time disk reads drop to zero once warm. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>