4268e00fa1
A downloaded (supply-chain) model file was fully trusted by the loader. Three memory-safety holes, all reachable by pointing the engine at a crafted shard — demonstrated crashing on pre-fix, now rejected fail-closed: st.h (safetensors): - header length `hlen` (u64 from the file) was unbounded before malloc(hlen+1): a crafted value overflows (malloc(0) then hdr[hlen]=0 OOB) or forces a giant allocation. Now bounded to the file size and a 512 MB cap; malloc NULL-checked. - json_get() returns NULL for missing/mistyped fields, but dtype/data_offsets/ shape were dereferenced blind (off->kids[0]) — a header omitting data_offsets SIGSEGV'd (verified). Now type/arity-checked before use. - data_offsets [a0,b0] were trusted: b0<a0 gave a negative nbytes -> malloc((size_t)) giant and an oversized memcpy into the caller's buffer in st_read_f32 (heap overflow); off could point outside the file. Now validated 0<=a0<=b0 and data_start+b0<=filesize. json.h: j_parse_val recursed with no depth limit -> stack overflow on nested input like [[[[...]]]]. Added J_MAX_DEPTH=1024 (headers are ~3 deep); wide-but- flat objects like the GLM header are unaffected (depth is decremented per return). eval_glm.py: tempfile.mktemp() -> mkstemp() — closes the TOCTOU/symlink race on a shared tmp dir (CWE-377). Network path (openai_server.py + serve SUBMIT parser) audited separately and is already sound: hmac.compare_digest auth, MAX_BODY cap, resolve()+relative_to traversal guard, list-form subprocess, bounded/validated SUBMIT header. All 62 tests pass; valid GLM/OLMoE shards load unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Tools
These scripts support model preparation and offline engineering work. They are not runtime dependencies of the C engine.
convert_fp8_to_int4.py,download_glm52.py: model preparationmake_glm_oracle.py,make_glm_bench_model.py: deterministic fixturesbenchmark_cuda_fixture.py,eval_glm.py,fetch_benchmarks.py: benchmarksgen_unicode.py: tokenizer table generation
Run them from c/, for example:
python3 tools/convert_fp8_to_int4.py --selftest
python3 tools/make_glm_bench_model.py --output /tmp/colibri-bench