ae4e31a15a
Adds `-iq3` to the ablation harness: a faithful torch model of llama.cpp's deployed 3.06-bpw IQ3_XXS format — 256-entry 4-dim magnitude grid (extracted from ggml-common.h, MIT), signs factored per 8 weights with the odd-parity constraint priced in (a violating block flips its smallest-magnitude sign), fp16 super-scale per 256 + 4-bit sub-scale per 32 searched over all 16 codes. Nearest-grid search runs as chunked matmul-argmin (|g|^2 - 2 q.g) — a full cdist materializes tens of GB on a 100M-param tensor and OOMed the first run. Measured (OLMoE-1B-7B, n=200 x hellaswag/arc/mmlu; the first four rows reproduce the published ablations exactly): fp16 58.0% int4 per-row 48.7% (-9.3pp, the shipped container's scheme) int3-g64 50.5% (-7.5pp) int3-g64-e8-rot 51.5% (-6.5pp, simulated rate-scaled ball) int3-iq3 49.3% (-8.7pp) int3-iq3-rot 51.5% (-6.5pp) The deployable IQ3 codebook plus rotation exactly ties the simulated E8 ball — that settles #452's codebook decision toward the IQ3-style block structure, with rotation mandatory (worth 2.2pp on this codebook).
Tools
These scripts support model preparation and offline engineering work. They are not runtime dependencies of the C engine.
convert_fp8_to_int4.py,download_glm52.py: model preparationmake_glm_oracle.py,make_glm_bench_model.py: deterministic fixturesbenchmark_cuda_fixture.py,eval_glm.py,fetch_benchmarks.py: benchmarksgen_unicode.py: tokenizer table generation
Run them from c/, for example:
python3 tools/convert_fp8_to_int4.py --selftest
python3 tools/make_glm_bench_model.py --output /tmp/colibri-bench