e141db047d
Split the resident weight classification into 5 sub-types so each can get different precision: sh = shared expert (highest sensitivity, fires every token) o = o_proj (reconstructs output, biggest attn tensor) kvb = kv_b_proj (reconstructs KV cache on every decode) attn = q_a/q_b/kv_a (other attention projections) dmlp = dense MLP (first 3 layers) New args: --shared-bits, --o-bits, --kvb-bits, --attn-bits, --dmlp-bits Each defaults to ebits (backward compat). When set, the converter applies that precision to just that tensor type. Research-backed plan: put the 3 compounding tensors (shared expert, o_proj, kv_b_proj) at int8 and everything else at grouped int4. Extra RAM cost: only +5.3 GB (those tensors are small vs the 372 GB expert pool on disk).
Tools
These scripts support model preparation and offline engineering work. They are not runtime dependencies of the C engine.
convert_fp8_to_int4.py,download_glm52.py: model preparationmake_glm_oracle.py,make_glm_bench_model.py: deterministic fixturesbenchmark_cuda_fixture.py,eval_glm.py,fetch_benchmarks.py: benchmarksgen_unicode.py: tokenizer table generation
Run them from c/, for example:
python3 tools/convert_fp8_to_int4.py --selftest
python3 tools/make_glm_bench_model.py --output /tmp/colibri-bench