d0971ff7c1
The cap_for_ram reserve for the 64-slot expert working set (ws_b = 64 × eb = 1.21 GB) is overcounted when EXPERT_BUDGET is active. At budget=4 only ws[0..3] are populated (not all 64), so the actual working set is 4 × eb = 76 MB — 16x less than reserved. The excess 1.06 GB was starving the LRU cache, capping it at 3 when budget=4 needs cap>=4. Fix: clamp ws_b to (budget+4) × eb when EXPERT_BUDGET < 64. This raises cap from 3 to 4, matching the budget. The cache can now hold all experts a token needs, eliminating the LRU thrashing that caused excessive disk re-reads (the SSD hammering). Measured (pipe2 + full stack, budget=4, RAM_GB=28): tok/s: 0.85 -> 1.03 (+21%) hit rate: 57% -> 73% (+28%) expert-disk: 18.2s -> 12.4s (-32%) decode: 37.7s -> 31.0s (-18%) Correctness: 32/32 oracle positions. Same fix applied to expert_avail() (the mirror function for pin budgeting).