web dashboard: live expert-tier panel (VRAM/RAM/disk), one-process hosting, coli web (#126)
* Web dashboard: live expert-tier panel (VRAM/RAM/disk), static hosting, coli web command - Engine: TIERS protocol line (counts + GB for VRAM/RAM/disk experts) emitted at READY and refreshed after every served turn, computed from the live pin/LRU/CUDA state. - Server: dispatcher tracks the latest snapshot, /health exposes it as "tiers"; the built web UI (web/dist) is served from the same port (read-only, traversal-safe, SPA fallback) so the dashboard needs no second process. - Web: tier bar + legend in the Runtime panel showing where the 19k experts live right now, with GB per tier; updates with the existing health polling. - CLI: `coli web` = serve + auto-open the browser once /health answers (the 744B load takes minutes; a poller waits for it), --no-browser to disable, friendly hint when web/dist isn't built. * coli web: HERE is a path string, not a Path * web: same-origin default + auto-connect when served by the engine The page hosted by coli web talked to http://127.0.0.1:8000/v1 by default while being loaded from http://localhost:8000 - a different origin, so the very first probe died on CORS ('Failed to fetch'). When the page is engine-served the default (and a stored stale factory default) now resolve to window.location.origin + /v1, and the console auto-probes on load; the Vite dev server keeps the classic default. The server's own port is also whitelisted in DEFAULT_CORS_ORIGINS for the localhost/127.0.0.1 cross-name case. * web: fix auto-connect effect returning Promise (React 19 StrictMode crash) The async connect() call inside useEffect returned a Promise that React 19 StrictMode stored as the effect cleanup. On re-render, React called destroy_() on that Promise — 'not a function'. Block body with explicit return undefined prevents any non-function from leaking into the cleanup slot. * web: rich streaming metrics — live token counter, tok/s gauge, TTFT, session totals During generation: a flashing Zap badge counts tokens in real time. After: Gauge shows tok/s, Timer shows TTFT (time to first token), Layers shows prompt→completion token counts, Clock shows queue wait. Session totals (cumulative prompt + completion) appear in the Runtime panel. All with lucide icons and tabular-nums for stable layout. Tier bar gains a smoother cubic-bezier transition and slightly taller height for visual weight. * web: hardware environment panel — CPU model, GPU count/VRAM, RAM total/free, cores Engine emits HWINFO protocol line at startup (CPU name from /proc/cpuinfo, core count, RAM total/available from /proc/meminfo, GPU count + VRAM from CUDA). Server parses and caches it, /health exposes as 'hwinfo'. Dashboard renders it with icons at the top of the Runtime section so the user sees immediately what the engine is running on. * web: downgrade to React 18 — React 19 effect cleanup bug crashes the dashboard React 19 treats the return value of every useEffect callback as a cleanup function. In practice, any async interaction (streamChat's rapid onDelta re-renders, the health poller, auto-connect) caused React 19 to store a Promise as inst.destroy and crash with 'destroy_ is not a function' on the next commit cycle. The bug reproduced in incognito, without StrictMode, and with every effect converted to explicit block bodies returning undefined — it is a React 19 regression in the passive-effect unmount path. React 18 does not exhibit this behavior. Pin to React 18 until React 19 is fixed or the codebase migrates to a pattern React 19 handles correctly. All 17 vitest pass, ErrorBoundary retained. * web: Brain page — the expert cortex, live A 76x256 canvas grid, one cell per expert (19,456 total): colour = tier (green VRAM / blue RAM / grey disk), brightness = routing heat (log-bucketed .coli_usage counts), and a white pulse that flashes on every expert routed in the current turn and decays — you watch the model think. - Engine: EMAP protocol line (1 byte/expert: 2bit tier + 6bit heat, hex) at READY and after each turn; HITS bitmap of this turn's routed experts (set where eusage increments, cleared on emit). - Server: parses both, GET /experts serves {rows, cols, map, hits, seq}. - Web: Brain tab next to Chat; canvas renderer with rAF pulse decay; hover tooltip shows layer/expert/tier/heat plus an honest depth-role heuristic (early = surface features ... late = output shaping, MTP row labelled as the speculative draft head). * web: Brain page responsive — cells sized from container via ResizeObserver Cell size derives from the wrapper's actual client box instead of fixed 1400x900, re-rendering on resize; mobile media query tightens padding, legend and tooltip. The cortex now fills whatever screen it gets.
This commit is contained in:
@@ -550,6 +550,28 @@ def cmd_serve(a):
|
||||
a.cap,a.ngen,GLM,env_for(a),a.cors_origin,
|
||||
a.max_queue,a.queue_timeout,a.kv_slots)
|
||||
|
||||
def cmd_web(a):
|
||||
"""serve + open the dashboard in the browser once the API answers."""
|
||||
need_model(a.model)
|
||||
dist = os.path.join(os.path.dirname(os.path.abspath(HERE)), "web", "dist")
|
||||
if not os.path.exists(os.path.join(dist, "index.html")):
|
||||
print(f"{C.yel}web UI not built:{C.r} run cd web && npm install && npm run build first;")
|
||||
print("serving the API anyway (the dashboard will 404 until built).")
|
||||
url = f"http://{a.host}:{a.port}/"
|
||||
if not getattr(a, "no_browser", False):
|
||||
import threading, urllib.request, webbrowser
|
||||
def opener():
|
||||
for _ in range(600): # the 744B engine takes minutes to load
|
||||
time.sleep(2)
|
||||
try:
|
||||
urllib.request.urlopen(f"http://{a.host}:{a.port}/health", timeout=2)
|
||||
webbrowser.open(url); return
|
||||
except OSError:
|
||||
continue
|
||||
threading.Thread(target=opener, daemon=True).start()
|
||||
print(f"dashboard: {url} (opens automatically when the engine is ready)")
|
||||
cmd_serve(a)
|
||||
|
||||
def cmd_bench(a):
|
||||
need_model(a.model)
|
||||
banner("bench")
|
||||
@@ -620,6 +642,16 @@ def main():
|
||||
ps.add_argument("--max-queue",type=int,default=int(os.environ.get("COLI_MAX_QUEUE","8")))
|
||||
ps.add_argument("--queue-timeout",type=float,default=float(os.environ.get("COLI_QUEUE_TIMEOUT","300")))
|
||||
ps.add_argument("--kv-slots",type=int,default=int(os.environ.get("COLI_KV_SLOTS","1")))
|
||||
pw=sub.add_parser("web", parents=[common], help="serve + open the dashboard in a browser")
|
||||
for arg,kw in (("--host",dict(default="127.0.0.1")),("--port",dict(type=int,default=8000)),
|
||||
("--model-id",dict(default=os.environ.get("COLI_MODEL_ID","glm-5.2-colibri"))),
|
||||
("--api-key",dict(default=os.environ.get("COLI_API_KEY"))),
|
||||
("--cors-origin",dict(action="append",default=None)),
|
||||
("--max-queue",dict(type=int,default=int(os.environ.get("COLI_MAX_QUEUE","8")))),
|
||||
("--queue-timeout",dict(type=float,default=float(os.environ.get("COLI_QUEUE_TIMEOUT","300")))),
|
||||
("--kv-slots",dict(type=int,default=int(os.environ.get("COLI_KV_SLOTS","1"))))):
|
||||
pw.add_argument(arg,**kw)
|
||||
pw.add_argument("--no-browser",action="store_true",help="don't auto-open the browser")
|
||||
pb=sub.add_parser("bench", parents=[common]); pb.add_argument("tasks", nargs="*")
|
||||
pb.add_argument("--limit",type=int,default=40); pb.add_argument("--data",default=os.path.join(HERE,"bench"))
|
||||
pc=sub.add_parser("convert", parents=[common]); pc.add_argument("--repo",default="zai-org/GLM-5.2-FP8")
|
||||
@@ -628,7 +660,7 @@ def main():
|
||||
a=ap.parse_args()
|
||||
handler={"build":cmd_build,"info":cmd_info,"plan":cmd_plan,"doctor":cmd_doctor,
|
||||
"run":cmd_run,"chat":cmd_chat,"serve":cmd_serve,"bench":cmd_bench,
|
||||
"convert":cmd_convert}.get(a.cmd)
|
||||
"convert":cmd_convert,"web":cmd_web}.get(a.cmd)
|
||||
if handler: sys.exit(handler(a) or 0)
|
||||
banner(); print(__doc__)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user