Files
colibri/issue_budget.md
T
woolcoxm e71d4fbe29 windows-dev: CUDA build script, expert budget + grouped quant research
- build_cuda.bat: one-shot nvcc DLL compilation for sm_120 (RTX 5070 Ti)
- issue_budget.md: EXPERT_BUDGET research (cap distinct experts/layer, 2x tok/s)
- issue_grouped_quant.md: root cause of int4 incoherence (per-row vs group-128 scales)
- glm_fp8_emit.py: FP8 e4m3 test weight generator for converter validation
2026-07-15 02:32:12 -04:00

8.7 KiB
Raw Blame History

TL;DR

New EXPERT_BUDGET=N env var that caps the number of distinct experts loaded per layer across the batch-union. When the union exceeds the budget, keeps only the highest-aggregate-gate-weight experts and drops the rest — they're never loaded from disk. On a 24 GB RAM host (cache cap=2), EXPERT_BUDGET=4 nearly doubles decode tok/s (0.18 → 0.33) and 4x's prefill speed (38.7s → 8.9s). Based on MoE-Spec (arXiv 2602.16052): "top 32 of 64 experts capture 93% of routing weight."

Branch: experiment/expert-budget (based on latest dev at 62419af)


The problem

Every expert miss costs ~19 MB of disk I/O. On low-RAM hosts where the LRU cache cap is tiny (e.g. cap=2 on 24 GB RAM), nearly every routed expert is a miss. With topk=8 and 75 sparse layers, a single-token decode reads ~8.5 GB of experts from disk. Under MTP with S=4, the batch-union can produce 20-32 distinct experts per layer — almost all misses — multiplying disk reads further.

The README documents this directly:

"on a cold cache each verified draft routes to extra experts (~660 → ~1100 expert-loads/token)"

The existing TOPP env var trims experts within a single position's top-K. But it cannot reduce the cross-position union — the deduplication across batch positions (prefill, MTP verification) that multiplies disk loads. That's the gap this fills.

How it works

The batch-union in moe() (line ~2306 of glm.c) deduplicates routed experts across all S positions into uniq[0..nu). After the union is built but before the cache-resolve/load loop, the budget cap kicks in:

if(EXPERT_BUDGET > 0 && nu > EXPERT_BUDGET):
    1. Compute aggregate gate weight per unique expert
       (sum of ws[s*K+kk] across all positions that route to it)
    2. Sort experts by descending aggregate weight
    3. Keep top EXPERT_BUDGET, mark rest as dropped
    4. Remove dropped experts from each position's idxs[]/keff[]
       (decrement keff, renormalize remaining weights if norm_topk)
    5. Compact uniq[] to kept experts only

Dropped experts are removed from idxs[] entirely, so they're never resolved, never loaded from disk, never computed. The downstream code already tolerates keff[s] < K (the TOPP path produces the same state), so this propagates correctly through resolve, matmul, and LRU promotion.

Relationship to existing features:

  • TOPP trims within one position's top-K (per-position)
  • EXPERT_BUDGET trims across the cross-position union (per-layer)
  • They compose: TOPP=0.8 EXPERT_BUDGET=12 first trims each position to its top-p mass, then caps the union

Code change

c/glm.c — 54 lines added, single file:

  1. Global + counter (near existing g_topk/g_topp):
static int g_expert_budget=0; /* EXPERT_BUDGET=N */
static int64_t g_budget_dropped=0; /* total experts dropped */
  1. Env var parsing (near existing TOPK/TOPP):
g_expert_budget = getenv("EXPERT_BUDGET")?atoi(getenv("EXPERT_BUDGET")):0;
  1. Budget cap in moe() batch-union (after uniq[] is built, before resolve loop) — full logic described above.

  2. Stats reporting alongside existing TOPK/TOPP line:

EXPERT_BUDGET=4 (dropped 13613 experts, ~257.3 GB I/O saved)

Measurements

Test setup: GLM-5.2 744B int4, 24 GB RAM (cache cap=2), Core Ultra 9 185H (AVX-VNNI), MTP=0, single-token decode, 32 tokens generated, same prompt: "Explain the concept of recursion in programming. Provide a simple example."

Config tok/s vs baseline hit rate prefill decode experts dropped I/O saved
Baseline (budget=0) 0.18 9.3% 38.7s 176.3s 0 0
EXPERT_BUDGET=12 0.19 +5% 14.0% 12.3s 171.3s 3,313 62.6 GB
EXPERT_BUDGET=6 0.26 +44% 21.0% 7.5s 122.5s 8,543 161.5 GB
EXPERT_BUDGET=4 0.33 +83% 16.4% 8.9s 97.4s 13,613 257.3 GB

Profile breakdown (prefill)

Metric Baseline Budget=4
expert-disk service 28.4s 3.5s (8x less)
expert-matmul 7.1s 1.4s
attention 2.3s 2.8s

Profile breakdown (decode, 32 tokens)

Metric Baseline Budget=4
expert-disk service 133.4s proportional ~65s
expert-matmul 23.9s ~12s
decode total 176.3s 97.4s

Key observations

  1. Prefill is the biggest winner. With S=14 positions, the batch-union has 50+ distinct experts per layer. Budget=12 cuts that to 12, saving 40+ × 19 MB × 75 layers of disk reads. Prefill goes from 38.7s to 8.9s — 4.4x faster.

  2. Decode improvement scales with budget tightness. Budget=12 barely constrains single-token decode (S=1 → max 8 experts < 12), so decode is barely affected. Budget=4 actually halves the decode load (8 → 4 experts/layer), nearly doubling tok/s.

  3. Hit rate improves because fewer experts compete for the tiny cache (cap=2). Budget=6 hit 21% vs baseline 9.3% — more than doubled.

  4. No crashes or instability at any budget value.

Quality impact — honest assessment

Budget=4 keeps only the top-4 of 8 routed experts per layer. The dropped 4 experts had the lowest gate weights — they contributed the least to the output. But this IS a quality trade-off.

On a cold cache (our test setup), quality assessment is confounded: both baseline and budget outputs are already garbled from int4 quantization + cold-cache routing. Sample comparison:

  • Baseline output: "Hire Some To Take Object-Oriented Programming Assignment"
  • Budget=4 output: "The world is a dangerous place, not so much because of the small percentage of people who are doing to do the little of people who are doing to"

Both are incoherent — the cold cache at int4 on 24 GB RAM produces poor output regardless of budget. The budget=4 output is not dramatically worse than the already-poor baseline — they're both garbled, just differently garbled. A proper quality assessment needs a warm cache where the baseline produces coherent text, so you can isolate the degradation from dropping experts.

Expected quality on warm cache: With norm_topk=1 (GLM-5.2 renormalizes gate weights), the top-4 experts typically capture ~80-85% of routing weight (the remaining 4 share 15-20%). Output should be mostly coherent with occasional word-choice degradation. Budget=6 (top-6 of 8, ~90%+ weight captured) should be nearly indistinguishable from baseline.

Recommendation for users:

  • EXPERT_BUDGET=6-8 on cold/low-RAM hosts — good speedup, minimal quality loss
  • EXPERT_BUDGET=4 on very-low-RAM hosts where speed matters more than quality
  • Leave OFF on high-RAM hosts where all experts are resident anyway (budget never triggers)

Safety

  • Default OFF (EXPERT_BUDGET=0) — zero behavior change unless explicitly set
  • Opt-in quality trade-off — same design philosophy as TOPP: the user explicitly chooses speed over quality
  • No output corruption — dropped experts' contributions are omitted and remaining weights renormalized (identical math to TOPP)
  • Compatible with all features — PILOT, MTP, CUDA, CACHE_ROUTE, PIPE, TOPP all work unchanged; they just see fewer experts in the union
  • No crash risk — no new memory allocation patterns, no new I/O, purely a filter on existing data structures

Reproduce

make ARCH=native

# Baseline:
SNAP=<model_dir> MTP=0 PROMPT="Explain recursion in programming." NGEN=32 ./glm 64

# With budget:
SNAP=<model_dir> MTP=0 EXPERT_BUDGET=4 PROMPT="Explain recursion in programming." NGEN=32 ./glm 64

# With MTP (where budget saves the most):
SNAP=<model_dir> MTP=1 EXPERT_BUDGET=16 PROMPT="Explain recursion in programming." NGEN=32 ./glm 64

The stats line prints EXPERT_BUDGET=N (dropped X experts, ~Y GB I/O saved).

What's needed before merge

  1. Warm-cache quality A/B test on a host with enough RAM (cap>=16) to produce coherent baseline output, then compare budget=4/6/8 text quality
  2. MTP interaction test — verify that budget + MTP composes correctly and doesn't crash when draft verification routes to budgeted-away experts
  3. Decide on defaults — should this auto-activate on low-RAM hosts (like cap auto-lowering)? My recommendation: no, leave it OFF and document the recommended value per RAM tier

Prior art

MoE-Spec (arXiv 2602.16052) — "Expert Budgeting for Efficient Speculative Decoding": Training-free expert budgeting at verification time. Key finding: "top 32 of 64 experts capture 93% of routing weight." Our approach applies the same principle but at the batch-union level (not just MTP verification), making it effective for prefill and single-token decode too.