tests: efficiency suite — tiny-model regression gates + full-model optimization dossier

Two layers of efficiency coverage for the engine, both parsing the telemetry
glm.c already emits (REPLAY/PROFILE/[PROF]/CUDA-tier) but nothing previously
asserted on:

1. test_inefficiency.py — tiny-model asserted regression tests (8 tests, run
   in make test via test-python). Gate on: throughput floor, PROFILE phase
   accounting sanity, disk-wait not dominant on a resident model, CPU greedy
   determinism, and (when a CUDA build is present) CUDA init, dense VRAM
   upload, and CPU-vs-CUDA argmax agreement >= 70%. CUDA tests auto-skip with
   a clear build hint on CPU-only binaries.

2. test_efficiency_report.py — opt-in optimization dossier for a real model.
   Turns on every instrumentation flag (PROF, COLI_CUDA_PROFILE, CACHE_ROUTE,
   DISK_SPLIT, LOOKA) and prints 9 sections (provenance, throughput + tail
   latency, where-time-goes, attention breakdown, expert cache, disk I/O +
   phase split, routing quality + predictability, speculation, GPU tiers),
   each flagging inefficiency with the concrete knob to move tok/s. Never
   fails CI.

tools/efficiency.py is the shared harness: parse_run() captures every signal,
run_engine() wraps the subprocess. Reuses PROFILE_RE/SPEED_RE from
tools/benchmark_cuda_fixture.py and extends the tok/s regex to also catch the
run_text (parenthesized) format the full-model PROMPT path uses.

Makefile adds: efficiency / efficiency-cuda / efficiency-report targets.

Verified end-to-end on the full glm52_i4_g64 model (CPU + CUDA).
This commit is contained in:
woolcoxm
2026-07-17 10:52:41 -04:00
parent e9b36141a4
commit d43b54534f
5 changed files with 1133 additions and 0 deletions
+23
View File
@@ -363,6 +363,29 @@ test-python:
test: test-c test-python
# --- Efficiency / regression suite (issue: "test the program for inefficiencies") ---
# The tiny-model assertions live in test_inefficiency.py and run as part of
# test-python (they're discovered by the test_*.py glob). These targets are
# convenience entry points; the opt-in full-model report is NEVER in `make test`.
#
# make efficiency tiny-model asserted regression tests (CPU; part of test-python)
# make efficiency-cuda the CUDA-path tests (requires a CUDA build — see below)
# make efficiency-report opt-in full-model 🟢/🔴 diagnostic, never fails CI
#
# CUDA build (Windows): the CUDA tests need a host built with -DCOLI_CUDA plus
# the runtime DLL. Do this FIRST — note CUDA_DLL=1 on BOTH the host and the
# rule below, or `make glm.exe` will rebuild a CPU-only host and overwrite it:
# make clean && make glm.exe CUDA_DLL=1 && make cuda-dll
# The tests auto-skip with a clear message if the host is CPU-only.
efficiency: test-python
$(PYTHON) -m unittest tests.test_inefficiency -v
efficiency-cuda:
$(PYTHON) -m unittest tests.test_inefficiency.TinyCudaEfficiencyTest -v
efficiency-report:
$(PYTHON) tests/test_efficiency_report.py
# Local validation: one portable CPU build and dependency-free tests.
check:
$(MAKE) clean