413b370fbf
Add EXPERT_BUDGET env var that caps the number of distinct experts loaded per layer across the batch-union. When the union exceeds the budget, keeps only the highest-aggregate-gate-weight experts and drops the rest from idxs[] so they're never loaded from disk. Complementary to TOPP (per-position) — this trims the cross-position union that multiplies under MTP/prefill. Based on MoE-Spec (arXiv 2602.16052): 'top 32 of 64 experts capture 93% of routing weight.' Measurements (GLM-5.2 744B, 24GB RAM, cap=2, MTP=0, 32 tokens): Baseline (budget=0): 0.18 tok/s, 9.3% hit, 176s decode, 39s prefill EXPERT_BUDGET=12: 0.19 tok/s, 14.0% hit, 171s decode, 12s prefill EXPERT_BUDGET=6: 0.26 tok/s, 21.0% hit, 123s decode, 7s prefill EXPERT_BUDGET=4: 0.33 tok/s, 16.4% hit, 97s decode, 9s prefill Budget=4 nearly doubles decode speed (+83%) and 4x's prefill speed by halving disk reads per layer. Default OFF (EXPERT_BUDGET=0).