GLM-5.2 MLA uses interleaved (DeepSeek-style) RoPE, which the C engine
implements. transformers < 5.11.0 applied split-half (Llama-style) RoPE in
GlmMoeDsa* instead; an oracle built on those versions silently drifts and the
engine scores 25/32 instead of the documented 32/32 (#281). Weights come out
identical across versions -- only the forward pass differs -- so a too-old
transformers produces an invalid ref_glm.json with no warning.
Add a version gate at the top of make_glm_oracle.py: hard sys.exit with an
actionable message citing the issue and the upgrade command. Reads the version
from importlib.metadata (authoritative installed-dist version) rather than the
mutable transformers.__version__ attribute -- the latter gets reset by the lazy
model-class import (from transformers import GlmMoeDsaForCausalLM), so reading
it after that import is unreliable. The gate runs before the heavy import and
falls back to the attribute only if the dist metadata lookup fails (editable/
src installs).
Validated end-to-end on transformers 5.13.1: script runs, ref_glm.json and
model.safetensors are byte-identical to the shipped versions, engine scores
32/32. With the floor raised to (5,14) the gate blocks with the expected
message.
make_glm_oracle.py wrote glm_tiny/model.safetensors and glm_tiny/config.json
without creating the directory first. safetensors.save_file writes a temp file
inside the target dir before the atomic rename, so on a clean checkout (no
pre-existing glm_tiny/) it aborts with an opaque error:
SafetensorError: Error while serializing: I/O error:
The system cannot find the path specified. (os error 3)
at path ".../glm_tiny/.tmpXXXXXX"
The directory only ever existed because it was left over from a previous run,
so the first-ever `python tools/make_glm_oracle.py` fails for every new user
following the README's verify step.
Create glm_tiny/ with Path.mkdir(parents=True, exist_ok=True) before the save
branch — covers the fp8 path, the bf16 path, and config.json. Path is already
imported; no new dependency, no change to the CPU build.
Two independent fixes validated end-to-end on fresh fixtures:
1. KV cache disk I/O (issue_diskio.md opportunities #1 + #4):
- kv_disk_append: fopen/fclose every turn -> persistent FILE* kept open
for the engine lifetime, lazy open on first append, closed in
serve_ctx_free. Eliminates per-turn handle creation overhead.
- kv_disk_append: ~157 small fwrites per position -> one contiguous
record memcpy'd into a staging buffer then a single fwrite per
position. The staging buffer grows on demand via realloc.
- kv_disk_truncate: closes the persistent handle before truncating
so the file actually shrinks on disc, then reopens lazily.
- KVState gains disk_fp, disk_buf, disk_buf_cap fields.
- Verified: serve-mode round-trip, write 11 tokens then reload and
resume with no re-prefill, then append 8 more and reload to 19.
2. Expert weight unfusing in test-model generators:
- The real GLM-5.2-FP8 checkpoint stores routed experts UNFUSED as
per-expert 2-D tensors, each with its own _scale_inv. HF fuses
gate+up into a single 3-D gate_up_proj for compute efficiency.
- The converter and C engine both expect the unfused layout. The
fused 3-D tensors were silently skipped by the converter, and the
engine crashed with missing-tensor errors.
- New unfuse_experts in glm_fp8_emit.py splits gate_up_proj and
down_proj into per-expert 2-D tensors. Called after reference
generation but before saving, in both generators, both FP8 and bf16.
- Also fixed: make_glm_oracle.py FP8 round-trip guard used p.dim()<2
which let 3-D fused experts through and crashed fp8_block_quantize.
Changed to p.dim()!=2 to match the converter ndim!=2 guard.
Validated full chain on fresh fixtures:
generator --fp8 -> 570 e4m3 tensors + 629 scale_inv, was 90 when fused
converter --group-size 0 -> per-row int4 fmt=2, engine loads clean
converter --group-size 128 -> grouped int4 fmt=4, 8-16x more scales,
engine loads clean, fmt=4 auto-detected in both mmap and slab paths
dequant error: grouped 1.14-1.22x lower than per-row vs FP8 source