5230717cf4
Scoring raw completions without GLM's training-time prefix runs the model out-of-distribution: scores drop and A/B sensitivity distorts (#108). Detect GLM via config.json model_type and prepend automatically, with a stderr notice. EVAL_PREFIX (including empty) still overrides for research use. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Tools
These scripts support model preparation and offline engineering work. They are not runtime dependencies of the C engine.
convert_fp8_to_int4.py,download_glm52.py: model preparationmake_glm_oracle.py,make_glm_bench_model.py: deterministic fixturesbenchmark_cuda_fixture.py,eval_glm.py,fetch_benchmarks.py: benchmarksgen_unicode.py: tokenizer table generation
Run them from c/, for example:
python3 tools/convert_fp8_to_int4.py --selftest
python3 tools/make_glm_bench_model.py --output /tmp/colibri-bench