536d8bfd1a
Root cause of #419's 'OOM slab': every per-slab mbind carries its own memory policy, so bound regions cannot merge — measured ~2 VMAs per slab, with or without MPOL_MF_MOVE. A PIN_GB=all load (19,456 experts x slab+fslab) creates ~78k VMAs and crosses the default vm.max_map_count=65530: posix_memalign dies with terabytes free. The fix binds the pinned hot-store as ONE arena per layer. Experts of a layer share a tensor shape, so a layer's pins pack at a fixed stride into two arenas (weights + scales): 2 mbinds and a handful of VMAs per layer instead of ~500. Slices are pre-attached to the slots before the load loop — slab_cap covers expert_load's realloc check, so its alloc branch never fires and expert_load itself is untouched. aslab marks arena ownership: expert_host_release detaches instead of freeing (a REPIN gpu-swap promotion must not free() an interior pointer), and expert_host_ensure re-attaches the slice before reloading. Per-slab mbind remains for the bounded allocations (dense qalloc, LRU ecache, GPU-tier staging), now without MPOL_MF_MOVE: every bind lands before the pread that first-touches the pages, so there is nothing to migrate. numa_init also gains a capability probe done right: one page-aligned page (mbind rejects unaligned addresses with EINVAL), disabling only on errno==EPERM — so a constrained container degrades with a message instead of crashing later, and an EINVAL can never masquerade as a missing capability. The arena path activates only when interleave is actually on (g_numa_nodes>=2, Linux, non-mmap): default builds stay byte-identical.