fa821a15a2
GLM-5.2 MLA uses interleaved (DeepSeek-style) RoPE, which the C engine implements. transformers < 5.11.0 applied split-half (Llama-style) RoPE in GlmMoeDsa* instead; an oracle built on those versions silently drifts and the engine scores 25/32 instead of the documented 32/32 (#281). Weights come out identical across versions -- only the forward pass differs -- so a too-old transformers produces an invalid ref_glm.json with no warning. Add a version gate at the top of make_glm_oracle.py: hard sys.exit with an actionable message citing the issue and the upgrade command. Reads the version from importlib.metadata (authoritative installed-dist version) rather than the mutable transformers.__version__ attribute -- the latter gets reset by the lazy model-class import (from transformers import GlmMoeDsaForCausalLM), so reading it after that import is unreliable. The gate runs before the heavy import and falls back to the attribute only if the dist metadata lookup fails (editable/ src installs). Validated end-to-end on transformers 5.13.1: script runs, ref_glm.json and model.safetensors are byte-identical to the shipped versions, engine scores 32/32. With the floor raised to (5,14) the gate blocks with the expected message.
Tools
These scripts support model preparation and offline engineering work. They are not runtime dependencies of the C engine.
convert_fp8_to_int4.py,download_glm52.py: model preparationmake_glm_oracle.py,make_glm_bench_model.py: deterministic fixturesbenchmark_cuda_fixture.py,eval_glm.py,fetch_benchmarks.py: benchmarksgen_unicode.py: tokenizer table generation
Run them from c/, for example:
python3 tools/convert_fp8_to_int4.py --selftest
python3 tools/make_glm_bench_model.py --output /tmp/colibri-bench