Two independent fixes validated end-to-end on fresh fixtures:
1. KV cache disk I/O (issue_diskio.md opportunities #1 + #4):
- kv_disk_append: fopen/fclose every turn -> persistent FILE* kept open
for the engine lifetime, lazy open on first append, closed in
serve_ctx_free. Eliminates per-turn handle creation overhead.
- kv_disk_append: ~157 small fwrites per position -> one contiguous
record memcpy'd into a staging buffer then a single fwrite per
position. The staging buffer grows on demand via realloc.
- kv_disk_truncate: closes the persistent handle before truncating
so the file actually shrinks on disc, then reopens lazily.
- KVState gains disk_fp, disk_buf, disk_buf_cap fields.
- Verified: serve-mode round-trip, write 11 tokens then reload and
resume with no re-prefill, then append 8 more and reload to 19.
2. Expert weight unfusing in test-model generators:
- The real GLM-5.2-FP8 checkpoint stores routed experts UNFUSED as
per-expert 2-D tensors, each with its own _scale_inv. HF fuses
gate+up into a single 3-D gate_up_proj for compute efficiency.
- The converter and C engine both expect the unfused layout. The
fused 3-D tensors were silently skipped by the converter, and the
engine crashed with missing-tensor errors.
- New unfuse_experts in glm_fp8_emit.py splits gate_up_proj and
down_proj into per-expert 2-D tensors. Called after reference
generation but before saving, in both generators, both FP8 and bf16.
- Also fixed: make_glm_oracle.py FP8 round-trip guard used p.dim()<2
which let 3-D fused experts through and crashed fp8_block_quantize.
Changed to p.dim()!=2 to match the converter ndim!=2 guard.
Validated full chain on fresh fixtures:
generator --fp8 -> 570 e4m3 tensors + 629 scale_inv, was 90 when fused
converter --group-size 0 -> per-row int4 fmt=2, engine loads clean
converter --group-size 128 -> grouped int4 fmt=4, 8-16x more scales,
engine loads clean, fmt=4 auto-detected in both mmap and slab paths
dequant error: grouped 1.14-1.22x lower than per-row vs FP8 source