26bd8b403a
GLM-5.2's config defines three stop tokens: <|endoftext|>, <|user|>, and <|observation|>. In serve mode, when the model generates <tool_call> blocks, int4-quantized logit noise can cause argmax to pick a <|user|> or <|observation|> token ID, immediately stopping generation. The <|user|> and <|observation|> tokens are role markers handled by the Python API server, not the C engine. Filter them out in stops_arm() when SERVE=1, keeping only <|endoftext|> as the stop token. - Add SERVE-mode guard in stops_arm() that retains only the EOS token - Log the number of filtered tokens for diagnostics