c4d14017df
SCORE mode scored raw token streams: GLM sees [gMASK]<sop> at the start of every training sequence, so unprefixed requests run the model out-of- distribution and silently distort logprobs (#108). The eval harness got the text-level fix in #194; contributors driving SCORE directly (e.g. the perplexity work in #153) were still exposed. Same detection rule as tools/eval_glm.py: config.json model_type contains "glm" (case-insensitive). The two ids are looked up in the snapshot's tokenizer.json ([gMASK], <sop>) rather than hardcoded. Requests that already carry the prefix pass through untouched — the patched eval_glm.py sends prefixed streams, so no double-prefixing. SCORE_PREFIX=0 restores stock behavior. One [SCORE] stderr notice when active. Validated on GLM-5.2 744B (M4 Pro, Metal): default run flips the '2 + 2 =' smoke question back to ' 4' (-1.006 vs ' 5' -2.950); SCORE_PREFIX=0 reproduces stock output; pre-prefixed requests give bit-identical logprobs to auto-prefixed ones. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>