63a6824881
run_serve_mux (SERVE_BATCH, used by openai_server.py / coli web) completes requests in mux_done, so the run_serve per-turn hook never fired there. Snapshot the window where hits0 is taken, record batched-forward latency around step_decode_batch, report on stderr at DONE. With KV_SLOTS>1 the window shares batched forwards across slots — same convention as the existing STAT hit%. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TCBNCciBaHea41QmidLUMn