0811730845
With METAL=1 + COLI_METAL=1, run_serve_mux (SERVE_BATCH=1) truncated every completion to exactly 1 token: step_decode_batch passes per-row kvs[]/positions[] with pos_base=0, but the two Metal decode fast paths (attention_rows and the FULL-LAYER CB in layer_forward_rows) ignored them and dispatched coli_metal_attn_decode/coli_metal_layer_decode with the model-bound Lc/Rc and the hardcoded pos_base. The kernels' contract is one sequence, row s at pos_base+s: ragged rows got roped at position 0 and attended over a T=1 window of the wrong cache, so greedy decode hit a stop token on the first batched step (DONE ... STAT 1). Gate both fast paths on !kvs. Ragged mux rows now take the CPU absorb path, which already reads kvs[s]/positions[s]/ks->kv_start per row; plain serve, chat/run, prefill and MTP verification (kvs==NULL) keep the fused GPU kernels unchanged. Verified on GLM-5.2 int4 (M5 Max): mux with COLI_METAL=1 went from 1-token DONE to full 16/16-token greedy completions, byte-identical to the plain-serve comparator on both test prompts; CPU-only mux was already correct (bisection); test-c and metal-test pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>