09bf17f001
Answers 'where does it slow down on THIS machine with THIS config' so users can tune RAM_GB/PIPE/DIRECT/PIN for their hardware without folklore: - startup header: CPU/cores/RAM/backend + effective knobs (cache cap, pin, DRAFT/PIPE/DIRECT/MMAP/IDOT/DSA/PILOT/CACHE_ROUTE) — every saved log is self-describing when comparing runs across configs or machines - per-forward decode latency ring (32k) -> p50/p90/p99/max, plus a tail diagnosis when p99 >> p50 (cold-cache expert loads) - expert I/O accounting at the pread/mmap-touch sites: GB fetched, MB/token, GB/s, hit rate, loads/token, pinned/LRU tier fill - phase shares of wall time and a plain-language verdict naming the knob most likely to move tok/s (I/O-bound vs compute-bound vs attention-bound) - reports after REPLAY / PROMPT / oracle runs (stdout) and per turn in serve mode (stderr; stdout stays the framed protocol) Additive only: with PROF unset every mode's output is byte-identical. hwinfo_emit's /proc probe is factored into hw_probe() and shared. Validated end-to-end on a tiny-random unquantized fixture (REPLAY, PROF on/ off, RAM_GB squeeze flips hit 96.9%->26.6% and the verdict follows); make check clean, 0 warnings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TCBNCciBaHea41QmidLUMn