6d3a6168e40ef2d490f708e20baacd0908b67d92
6 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6afffbcbf2 |
Profiling page: per-turn phase timings, live in the web dashboard
The engine already tracks where each turn's wall time goes (expert-disk service, I/O wait, expert matmul, attention, lm_head) — it just only spoke at exit or under PROF=1. Stream it instead: - glm.c: mux serve emits a per-turn "PROF" protocol line next to TIERS/HITS (window deltas per request, same convention as the STAT hit%); the phase window base is now always captured (a few loads per request). - openai_server.py: parses PROF into a 120-turn rolling window and serves it at /profile (read-only, same trust level as /health). - web: new Profiling tab — stat tiles (tok/s, wall, tokens/forward, disk service), wall-time composition bars for the last turn and the window, per-turn throughput and stacked phase columns with hover readouts, and a table of recent turns. Disk service is shown apart from the stack: it overlaps with compute, so only the I/O wait the compute thread felt counts inside wall time. Phase colours are a CVD-validated set with gaps + legend + table so identity never rides on colour alone. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WhTmF8yvEBgSkUKSVfZF7P |
||
|
|
869e80ee2d |
web: fallback when crypto.randomUUID() is unavailable
Some browsers and insecure contexts (plain HTTP) throw on crypto.randomUUID(), which silently breaks the chat UI. Fall back to a plain UUID v4 generator when the native call fails - same behaviour on modern browsers, works everywhere else. |
||
|
|
ec89136029 |
GPU resident pipeline: batch CUDA attention, head-sharded kv_b, prefill expert groups, W4A16 mixed dispatch (#111)
* Fuse CUDA expert MLP execution * Group CUDA expert transfers by device * Instrument grouped CUDA expert execution * Bound grouped CUDA decode scratch * Execute expert groups across GPUs in parallel * Release host backing for multi-GPU experts * Define quality-preserving memory policies * Overlap cold expert loading with resident compute * Adapt expert placement with session LFRU * Fuse q4 expert gate and up dispatch * Plan CPU work on physical cores * Batch grouped expert CUDA kernels * Separate VRAM and RAM expert placement * Add ragged multi-sequence decode forward * feat(runtime): add continuous decode scheduler * Route concurrent API requests through batch scheduler * Harden multiplex request lifecycle and framing * Cancel disconnected multiplex requests * Bind API port before starting the engine * fix automatic KV slot allocation * add native int4 Tensor Core grouped GEMM * add Tensor Core throughput benchmark * optimize packed int4 low-row kernels * add asynchronous CUDA staging streams * document validated six-GPU dense acceleration * tune six-GPU expert hot set * raise validated expert hot-set target * add CUDA MLA absorption core * fuse grouped expert gate and up projections * Warn for explicit lossy routing flags * Add full-resident expert placement mode * Adapt VRAM expert slots to live routes * Accelerate int4 matvec on AVX-512 * Reduce AVX-512 and RoPE decode overhead * Seed every GPU expert layer after prefill * Limit live GPU swaps during decode * CUDA batch MLA attention, kv_b head-sharding, fused o_proj, expert-group dispatch, W4A16 kernels Lab-qualified on the 6x RTX 5090 machine (914-token request benchmark): - batch MLA absorption kernel (COLI_CUDA_ATTN=1): whole-batch attention on device, 154.8s -> 102.4s - attention -> o_proj fusion on the layer device: -> 97.4s - kv_b head-sharding across cards (COLI_CUDA_ATTN_SHARD=1), no weight duplication: -> 94.05s - per-device expert-group dispatch with pinned-buffer async transfers, W4A16 tensor-core kernels for the shared expert, OMP hot-thread tuning Negative results (reverted, kept out): GPU-side weighted scatter-add (atomics + per-layer D2H lose 43.8%), shared-expert fused small-batch kernel (-38.8%), W4A4 grouped tensor cores (int4 activations corrupt output). Details in the lab research log. * GPU resident pipeline: device-resident prefill attention chain, GPU expert groups in prefill, batched router, W4A16 mixed dispatch COLI_CUDA_PIPE=1 keeps the prefill data plane on the layer home device; control flow (routing, cache/pin management) stays on CPU. Any CUDA failure falls back to the unchanged CPU path. - Device primitives + unit tests (tests/test_pipe_cuda.cu): rmsnorm (strided), interleaved RoPE, silu-mul, residual add, fixed-order row merge (no atomics), device-input GEMM, persistent per-device scratch. All verified against the engine's CPU math on SM120 (worst 1.2e-5). - attn_pipe_prefill: q_a -> norm -> q_b -> rope -> kv_a -> norm -> rope -> batch attention -> o_proj in one device chain (q_a/q_b/kv_a colocated with kv_b); only the final [S,D] and the new KV rows return to host. Attention 41.2s -> 30.8s on the 1571-token benchmark. - Prefill batch-union now uses the GPU expert groups (previously gated to S<=64, leaving all VRAM-resident experts idle during prefill - measured 21ms of GPU expert time in a 148s prefill). Expert phase 78.9s -> 69.0s. - Router computed as one batched matmul instead of S sequential rows (bit-identical math). - W4A16 tensor-core path for expert groups (COLI_CUDA_TC_W4A16=1) with row-count mixed dispatch: >=16 rows per expert use tensor cores, smaller batches keep the naive kernel (tensor cores measured negative below ~16 rows). Expert phase 69.0s -> 64.3s, decode unaffected. Net on the 1571-token prefill benchmark: 148.8s -> 114.3-126.8s (component timings stable across runs; wall drifts +-3-5s because .coli_usage placement learning shifts the expert tiers between runs). PROFILO now also prints the prefill-phase breakdown. * Skip OMP hot-thread tuning when CUDA is enabled The active-spin worker team measured 66.9s->20.9s on the CPU-only Zen5 build, but on the six-GPU full-residency workload the spinning workers contend with the CUDA dispatch threads: ~4x slower prefill with the process stuck near 1.8 cores. Gate the tuning on COLI_CUDA so each configuration keeps the behavior it was measured to prefer. * Inc.2a: sparse layers fully resident on the layer device, residual hops cards at layer boundaries COLI_CUDA_PIPE=2 keeps the residual stream on the layer home device for consecutive sparse layers (cudaMemcpyPeer at boundaries): in/post norms, attention chain, both residual adds and the shared-expert MLP run on device. Per layer only the post-norm activations (router + CPU-tier experts + group gather), the new KV rows and, on DSA indexer layers, the pre-attention norm leave the card. Per-layer transfers drop from ~130MB to ~70MB. A device-side snapshot at layer entry makes any mid-layer CUDA failure fall back to the unchanged CPU path idempotently. 1571-token prefill: 127.1s (PIPE=1 control) -> 117.6/118.9s, components attention 30.8->26.1, other 31.8->22.5-24.5; output verified coherent against the control. * Head-sharded attention inside the pipe: negative on PCIe star topology, gated opt-in Slicing q per card from the home device and collecting ctx back serializes ~95MB/layer through the home card's PCIe link: attention 26.1s -> 41.4/44.4s on the 1571-token benchmark (two repeats), wall 117.6 -> 135-138s. The standalone host-path sharding won because six cards uploaded from host RAM in parallel; a home-device star has no such parallelism without NVLink. Kept behind COLI_CUDA_PIPE_SHARD=1 for interconnects where peer bandwidth does not share one root port. * Inc.3: device-resident KV shadow for decode attention Decode re-uploaded the whole latent+rope window per layer per token (~300MB/token at 1571 context). Each layer now keeps a device shadow of the compressed KV on its kv_b card, bulk-synced when behind and appended incrementally; the host cache stays canonical. Invalidation on kv_bind (slot switch), kv_alloc (resize) and on any overwrite of mirrored rows, with the legacy full-upload path as fallback. Measured (COLI_CUDA_PIPE gate): short-context decode 5.48 -> 5.59/5.87 tok/s, 1571-context decode 4.14 -> 4.22 tok/s. Decode remains CPU-expert bound; the shadow removes the transfer tax, not the compute. * tools: unified user-experience benchmark (bench_ux.sh) Two fixed scenarios (short chat, long-document QA), TTFT + decode tok/s + first-line drift check, TEMP=0 DRAFT=0 enforced, medians over REPS runs. Encodes the measurement discipline from the lab record: same binary per comparison, judge medians because .coli_usage placement learning drifts wall times between runs. * tools: bench_ux.sh executable bit * gitignore compiled test binaries * tools: expert_atlas.py — measure per-expert topic affinity (#175) Diffs .coli_usage across 10 themed probe batches (code/math/chinese/ prose/science/law/poetry/structured/translation/casual, 3 prompts each) driven through a running API server — one engine load total. Every touched expert gets a topic-affinity vector, entropy, and a specialist/ generalist label; output experts.json feeds the Brain page hover. * serve: persist .coli_usage after every turn in mux mode, not only at exit run_serve_mux saved the learning cache once at shutdown; a crash lost the whole session's routing history, and live consumers of the file (expert_atlas.py diffs it between probe batches) saw a frozen snapshot. Now saved per turn like the interactive path (165KB write, negligible). * web: Brain hover shows measured expert atlas when published If /experts.json (from tools/expert_atlas.py, #175) is served next to the app, the tooltip upgrades from the depth heuristic to measured data: specialist/generalist label, entropy, and the top-3 topic affinities. Row index maps to real layer (row+3, last row = MTP 78). Falls back to the heuristic when no atlas is published. --------- Co-authored-by: JustVugg <JustVugg@users.noreply.github.com> |
||
|
|
2878c60987 |
web dashboard: live expert-tier panel (VRAM/RAM/disk), one-process hosting, coli web (#126)
* Web dashboard: live expert-tier panel (VRAM/RAM/disk), static hosting, coli web command - Engine: TIERS protocol line (counts + GB for VRAM/RAM/disk experts) emitted at READY and refreshed after every served turn, computed from the live pin/LRU/CUDA state. - Server: dispatcher tracks the latest snapshot, /health exposes it as "tiers"; the built web UI (web/dist) is served from the same port (read-only, traversal-safe, SPA fallback) so the dashboard needs no second process. - Web: tier bar + legend in the Runtime panel showing where the 19k experts live right now, with GB per tier; updates with the existing health polling. - CLI: `coli web` = serve + auto-open the browser once /health answers (the 744B load takes minutes; a poller waits for it), --no-browser to disable, friendly hint when web/dist isn't built. * coli web: HERE is a path string, not a Path * web: same-origin default + auto-connect when served by the engine The page hosted by coli web talked to http://127.0.0.1:8000/v1 by default while being loaded from http://localhost:8000 - a different origin, so the very first probe died on CORS ('Failed to fetch'). When the page is engine-served the default (and a stored stale factory default) now resolve to window.location.origin + /v1, and the console auto-probes on load; the Vite dev server keeps the classic default. The server's own port is also whitelisted in DEFAULT_CORS_ORIGINS for the localhost/127.0.0.1 cross-name case. * web: fix auto-connect effect returning Promise (React 19 StrictMode crash) The async connect() call inside useEffect returned a Promise that React 19 StrictMode stored as the effect cleanup. On re-render, React called destroy_() on that Promise — 'not a function'. Block body with explicit return undefined prevents any non-function from leaking into the cleanup slot. * web: rich streaming metrics — live token counter, tok/s gauge, TTFT, session totals During generation: a flashing Zap badge counts tokens in real time. After: Gauge shows tok/s, Timer shows TTFT (time to first token), Layers shows prompt→completion token counts, Clock shows queue wait. Session totals (cumulative prompt + completion) appear in the Runtime panel. All with lucide icons and tabular-nums for stable layout. Tier bar gains a smoother cubic-bezier transition and slightly taller height for visual weight. * web: hardware environment panel — CPU model, GPU count/VRAM, RAM total/free, cores Engine emits HWINFO protocol line at startup (CPU name from /proc/cpuinfo, core count, RAM total/available from /proc/meminfo, GPU count + VRAM from CUDA). Server parses and caches it, /health exposes as 'hwinfo'. Dashboard renders it with icons at the top of the Runtime section so the user sees immediately what the engine is running on. * web: downgrade to React 18 — React 19 effect cleanup bug crashes the dashboard React 19 treats the return value of every useEffect callback as a cleanup function. In practice, any async interaction (streamChat's rapid onDelta re-renders, the health poller, auto-connect) caused React 19 to store a Promise as inst.destroy and crash with 'destroy_ is not a function' on the next commit cycle. The bug reproduced in incognito, without StrictMode, and with every effect converted to explicit block bodies returning undefined — it is a React 19 regression in the passive-effect unmount path. React 18 does not exhibit this behavior. Pin to React 18 until React 19 is fixed or the codebase migrates to a pattern React 19 handles correctly. All 17 vitest pass, ErrorBoundary retained. * web: Brain page — the expert cortex, live A 76x256 canvas grid, one cell per expert (19,456 total): colour = tier (green VRAM / blue RAM / grey disk), brightness = routing heat (log-bucketed .coli_usage counts), and a white pulse that flashes on every expert routed in the current turn and decays — you watch the model think. - Engine: EMAP protocol line (1 byte/expert: 2bit tier + 6bit heat, hex) at READY and after each turn; HITS bitmap of this turn's routed experts (set where eusage increments, cleared on emit). - Server: parses both, GET /experts serves {rows, cols, map, hits, seq}. - Web: Brain tab next to Chat; canvas renderer with rAF pulse decay; hover tooltip shows layer/expert/tier/heat plus an honest depth-role heuristic (early = surface features ... late = output shaping, MTP row labelled as the speculative draft head). * web: Brain page responsive — cells sized from container via ResizeObserver Cell size derives from the wrapper's actual client box instead of fixed 1400x900, re-rendering on resize; mobile media query tightens padding, legend and tooltip. The cortex now fills whatever screen it gets. |
||
|
|
57730f6196 |
Web UI: runtime console + KV session panel, with runtime/storage libs and tests (#32)
* Add web runtime console and KV sessions * test(web): cover runtime and KV session behavior * Fix runtime helper formatting |
||
|
|
a2942b2172 | Web interface (React/TS, shadcn, pure OpenAI-API client) under web/ (#23) |