Commit Graph

314 Commits

Author SHA1 Message Date
JustVugg 8b36736e5d Merge branch 'p357' into trial357
# Conflicts:
#	c/Makefile
2026-07-17 20:47:21 +02:00
JustVugg 70a58799d6 Merge branch 'p354' into trialsamp
# Conflicts:
#	c/Makefile
2026-07-17 20:45:53 +02:00
Vincenzo 1a956b0119 Merge pull request #331 from nbeerbower/st-pread-full
st: chunked pread with EINTR retry and honest short-read errors
2026-07-17 20:42:42 +02:00
Vincenzo 19f267794f Merge pull request #329 from mohamedmastouri2000-boop/docs/windows-native
docs: Windows 11 native install walkthrough (toolchain, Smart App Control, CUDA DLL, failure index)
2026-07-17 20:42:38 +02:00
Vincenzo d2226e222f Merge pull request #346 from bm1016bm-svg/docs/readme-zh-tw
docs: add Traditional Chinese README
2026-07-17 20:42:35 +02:00
Vincenzo d6d28f0c5c Merge pull request #353 from woolcoxm/docs/environment-refresh-dev
docs(env): refresh ENVIRONMENT.md — document all 34 drifted env vars
2026-07-17 20:42:31 +02:00
Vincenzo c1e709eb96 Merge pull request #351 from LordMZTE/nix-flake
fix(flake): create flake.lock, use --set-default for COLI_ENGINE
2026-07-17 20:42:28 +02:00
Vincenzo 3d0a28b7bf Merge pull request #350 from woolcoxm/fix/build-warnings-dev
fix(warnings): silence two -Wall/-Wextra warnings on the MinGW build
2026-07-17 20:42:25 +02:00
Vincenzo 59d74ae435 Merge pull request #340 from tonuonu/fix/fs-detect-9p-statfs
glm.c: detect 9p via statfs f_type, not the /mnt/ path prefix
2026-07-17 20:42:22 +02:00
Vincenzo 1af9435760 Merge pull request #328 from mohamedmastouri2000-boop/fix/olmoe-ebits
convert_olmoe.py: wire --ebits through to quantization (fixes #323)
2026-07-17 20:42:18 +02:00
Vincenzo 1552db2323 Merge pull request #347 from ZacharyZcR/tools/e8-rate-scaled
tools: rate-scale the E8 lattice ball by bit-width (#81) — int3-e8 is now real int3
2026-07-17 19:17:33 +02:00
JustVugg 211d4488c3 win32: stop erasing the real ReadFile error — pread failures name their WinErr (#307)
compat_pread mapped every non-EOF ReadFile failure to EIO, so the field report
was always the same three words: "Input/output error". #307 then burned three
rounds of guessing across three people — storage? contention? alignment? —
because the actual GetLastError code never appeared anywhere.

compat_pread now stashes the code in a per-thread slot and pread_full appends
it to the failure line:

    pread qs: Input/output error (off 780592, 0/8192 bytes, WinErr=1450)

That one number is the difference between "insufficient system resources"
(1450 -> memory pressure), "sharing violation" (32 -> AV interference),
"device not ready" (21 -> the T: drive itself) and a dozen other distinct
diagnoses. Diagnostic only: no behavior change, no retry policy — that
discussion lives in #361 and should be settled AFTER the first report tells
us which error we are actually retrying.

Linux path untouched byte-for-byte; the Windows compile is certified by the
check.yml MSYS2 job.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 18:49:28 +02:00
woolcoxm 5d16368817 sampling: partial top-keep select in attention_rows DSA — O(nk) quickselect, not O(nk log nk) qsort (#356)
The DSA lightning indexer selects the top-index_topk (2048) context keys to
attend to by finding the threshold = keep-th largest attention score. It
previously full-qsorted all nk scores per layer per token (O(nk log nk)) just
to read one pivot value, then scanned the original array in position order to
build the kept set.

Replace the qsort with partial_select_desc (Hoare quickselect, median-of-three,
descending): O(nk) average to partition the keep largest into a[0..k), then the
threshold is min of that block. The two position-order scans (>thr then ==thr)
are UNCHANGED, so the kept-position set is bit-identical -- a stronger contract
than #335's sampling heap (which was multiset-only because the heap was unstable
and changed accumulation order). The quickselect pivot IS by definition the
keep-th largest, so the new threshold equals old tmp[keep-1] exactly.

Measured (bench_dsa_select, keep=2048, median of 2000 reps):
  nk=2049:   119us -> 5.7us   (21x)
  nk=8192:   626us -> 43us    (15x)
  nk=32768:  2.8ms -> 0.28ms  (10x)
  nk=65536:  6.6ms -> 0.47ms  (14x)
The gap widens with context (linear vs n-log-n). DSA only fires past index_topk,
so this is precisely the long-conversation regime where decode latency matters.

Adds test_dsa_select (in TEST_BINS): 129 cases asserting element-wise identical
kept-set vs an independent qsort reference across shapes (random, peaked,
sorted, reverse-sorted, tie-plateau, all-equal) and edges (keep==1, keep==nk).
Also directly checks the partition invariant.

Adds bench_dsa_select (on-demand, NOT a gate): reproduces the table above.
2026-07-17 09:54:40 -04:00
woolcoxm 37c96ee07a bench: add bench_topp -- head-to-head old-qsort vs new-heap timing on V=151936 (#335)
test_topp proves correctness; bench_topp measures the latency claim. It re-implements
the OLD dist_build (full-vocab qsort) inline on a private buffer and times it against
the REAL new dist_build over identical inputs in one process: V=151936, temp=0.7,
3 shapes (realistic / uniform / plateau) x 4 nuc values (0.5/0.9/0.95/0.99), 2000
timed reps each, median reported. Deliberately NOT in TEST_BINS -- it's a microbench,
not a gate. Build on demand: make tests/bench_topp && ./tests/bench_topp
2026-07-17 09:22:26 -04:00
woolcoxm ba00889fb9 sampling: partial top-p select in dist_build — O(V) heapify + k pops, not O(V log V) qsort (#335)
dist_build() sorted the entire 151936-entry vocab by probability (qsort) on every
sampled token whenever 0 < g_nuc < 1 — the serve default — and again per draft
position under rejection sampling. Measured cost: 5.6-8.0 ms/call; the actual work
is finding the few-hundred-token head whose cumulative mass reaches g_nuc.

Replace the full qsort + linear scan with a max-heap partial select:
  - Floyd heapify g_pidx over V by descending g_pbuf prob  (O(V), cache-friendly)
  - pop winners to the array's high end until cum >= g_nuc  (k * O(log V))
  - the remaining heap prefix IS the tail -> zero it, renormalize the head

Winners land in g_pidx[out..V-1] in descending order, so s2 accumulates in the
same order as before -> head is unchanged on tie-free shapes (ties were already
unspecified under the unstable qsort). All four dist_build/dist_sample contract
properties hold: g_pbuf stays id-indexed, g_pidx stays internal, the tail is
fully zeroed, the head renormalizes to 1.

No API change, no caller change, no new globals.

c/tests/test_topp.c (new): drives the REAL dist_build via the include-glm.c
pattern against an independent double-precision reimplementation of the OLD
algorithm. 123 cases: 6 sizes (1..1519) x 5 shapes (uniform/peaked/geometric/
plateau/sharptail) x 4 nuc values, plus the g_nuc>=1 guard-off paths and V=1.
Tie-free shapes compare head values within 1e-6 rel (float vs double renorm
noise); tie shapes compare value-multisets. No scratch files -> builds clean on
Windows MinGW without the unmerged mkdtemp shim (#352).
2026-07-17 09:01:23 -04:00
woolcoxm 099e0c13c4 docs(env): refresh ENVIRONMENT.md — document all 34 drifted env vars
ENVIRONMENT.md says it is "generated from source by scanning every
getenv() site", but it had drifted far behind the code. Reconciling the
doc against the C sources (the MAINTAINING-DOCS.md procedure) found 34
engine env vars read by the code but undocumented, and nothing stale.

Added (defaults/effects taken from source, per the maintenance rules -
nothing invented):

Performance/tuning (11): COLI_NUMA, PILOT_TWO, COUPLE/COUPLE_K/COUPLE_D,
  ROUTE_TRACE, COLI_NO_FUSED_PAIR, DISK_SPLIT, I4S, SPEC_PIN,
  COLI_RAM_OVERCOMMIT

CUDA (16): COLI_CUDA_ATTN_SHARD, COLI_CUDA_PIPE/_PIPE_SHARD/_PIPE_S_MIN,
  COLI_CUDA_MTP, COLI_CUDA_ASYNC, COLI_CUDA_DUAL_PROJ, COLI_CUDA_W4_PACKED,
  COLI_CUDA_TC_INT4/_TC_MIN_ROWS/_TC_W4A16/_TC_W4A16_MIN,
  COLI_CUDA_SHARED_W4A16/_SHARED_W4A16_MIN_ROWS, COLI_METAL_UNTRACKED

Advanced/debug (7): SCHEMA, EXPERT_BUDGET/_EXPERIMENTAL, TOKENS,
  SCORE_PREFIX, REPIN_VERBOSE, PPL (olmoe-only), COLI_PROMPT (CLI section)

Also bumped the "Generated from" line to dev @ d5327e2 and noted the scan
now covers olmoe.c, backend_cuda.cu, backend_metal.mm (not just glm.c).

Verified: the code-vs-doc diff is now empty - all 111 distinct C env vars
are documented. The reverse diff (doc vars not in C code) is 14 entries,
all legitimate: 11 Python-side vars correctly in the Server/CLI section,
plus 3 prose constants (IOSQE_ASYNC, O_DIRECT, the VAR format word).
2026-07-17 08:28:41 -04:00
woolcoxm 411f237f94 fix(warnings): silence two -Wall/-Wextra warnings on the MinGW build
A clean 'make glm' on MinGW emitted exactly two warnings, both real:

1. compat.h:240 - ignoring pragma comment [-Wunknown-pragmas]
   #pragma comment(lib, "psapi.lib") is an MSVC directive; MinGW/GCC
   warns about it. Guarded with ifdef _MSC_VER - MinGW links psapi via
   -lpsapi (already in the Makefile), MSVC keeps the pragma.

2. glm.c:1210 - g_numa_nodes defined but not used [-Wunused-variable]
   g_numa_nodes is only read/written inside ifdef __linux__ blocks, so on
   every non-Linux build it is a static that is never used. Moved the
   definition under the same __linux__ guard; nothing references it off-Linux.

Verified: rm -f *.o glm.exe && make glm -> 0 warnings, 0 errors.
2026-07-17 08:26:32 -04:00
LordMZTE 03e643f006 fix(flake): create flake.lock, use --set-default for COLI_ENGINE 2026-07-17 14:02:30 +02:00
JustVugg d5327e2252 omp: cap LLVM libomp idle spin — a parked engine must not burn 3000% CPU (#341)
The hot-thread tuning sets OMP_WAIT_POLICY=active and caps libgomp's spin
with GOMP_SPINCOUNT=200000. LLVM libomp reads NEITHER: under active policy it
sets KMP_BLOCKTIME=infinite, so once the answer ends and the engine parks on
stdin, the whole OMP team spins forever — ~100% per thread, the 3000% figure
reported on FreeBSD 15.1 in #341 (clang/libomp is the default toolchain
neighborhood there; the same applies to macOS builds).

KMP_BLOCKTIME=200 keeps the team hot for 200 ms after each parallel region —
plenty to bridge back-to-back per-expert matmuls — and then sleeps it.
setenv with overwrite=0, so a user-provided KMP_BLOCKTIME still wins, and
libgomp builds ignore the variable entirely: on GCC toolchains this commit
is a no-op by construction.

Not reproduced locally (this box is gcc/libgomp and idles clean); shipped as
the mechanism-matching candidate fix with a request on the issue for the
reporter to verify on dev and to name their OpenMP runtime (ldd | grep omp).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 13:53:20 +02:00
Vincenzo 617ae51631 Merge pull request #348 from woolcoxm/fix/nix-flake-packaging-345
fix(flake): package engine, support modules, and python for Nix builds
2026-07-17 13:52:06 +02:00
JustVugg 7eb239328d coli: serve pidfile + coli stop — one command to shut engine and server down
No more manual pkill: cmd_serve writes a pidfile, cmd_stop finds the server
(pidfile, then /proc cmdline) and its engine (comm glm/exe/olmoe with SERVE=1
in environ — the engine re-execs for OMP tuning so its comm is 'exe', which is
why every 'pkill -x glm' in history killed nothing). SIGTERM, wait, SIGKILL.
--dry-run lists targets without acting. First real use took down a live
serve+engine pair cleanly and released 16 GB.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 41a872c331a2a0a8655699e0171c68dd2bcda186)
2026-07-17 13:27:02 +02:00
JustVugg 8bf4cb9a98 coli: chat attaches to a running serve — the engine survives the chat (local)
The cold-chat cost, measured on this box: every `coli chat` spawns a private
engine (34-136 s of resident load) and starts with an empty expert cache (hit
4% cold vs 55% warm). Quitting throws both away.

`coli chat` now probes localhost:8000 first (~1 ms when nothing listens) and,
if a `coli serve` answers, runs the REPL over plain OpenAI SSE against it:
stdlib urllib only, engine byte-protocol untouched. --attach [URL] forces it,
--no-attach restores a private engine. reasoning_content keepalive pings are
filtered; :reset starts a new conversation client-side (the server's KV slots
reuse prefixes per conversation on their own).

Verified against a mock SSE server (pings ignored, markdown rendered, :reset,
clean exit) and against the real 744B model: two consecutive sessions, second
attach instant with zero reload — the engine stayed resident at 15.7 GB across
both. Honest limit: warmth carry-over BETWEEN different conversations is small
here because cap=3 slots/layer is a short memory; the structural wins are the
load never being repaid and same-conversation continuation.

LOCAL ONLY for now — not pushed, per the current working rule.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit aa406ccab8a4501b924bc3a9f4725dd1d18a685d)
2026-07-17 13:26:59 +02:00
woolcoxm 528b3cff37 fix(flake): package engine, support modules, and python for Nix builds (#345)
Four bugs in the Nix flake, all addressed:

1. checkPhase failed with "make: python3: No such file or directory"
   because `make test-c` shells out to python3 (c/Makefile: PYTHON ?=
   python3) but python3 wasn't in nativeBuildInputs. Add pkgs.python3.

2. `coli serve` reported "engine is not built" even though glm was
   packaged, because the engine landed in $out/bin while coli looks for
   it beside itself, in libexec/colibri, or via $COLI_ENGINE (c/coli:57-
   66). The wrapper now sets COLI_ENGINE to the bundled engine.

3. `coli serve` failed with ModuleNotFoundError: openai_server because
   the support modules (openai_server.py, resource_plan.py, doctor.py)
   were not installed. Install them and put their dir on PYTHONPATH.

4. Rebuilt the install layout into a self-contained $out/lib/colibri/
   that mirrors the source tree coli expects, with $out/bin/{glm,coli}
   as user-facing entry points (glm symlinked; coli wrapped).

NOTE: flake.lock is intentionally not included — it must be generated
on a Nix host via `nix flake lock` and committed separately, since it
requires narHash values that only the nix tooling can compute.
2026-07-17 07:18:24 -04:00
ZacharyZcR 3f239dbb92 tools: rate-scale the E8 lattice ball by bit-width — int3-e8 is now a real int3 codebook, not a fixed 2-bit ball 2026-07-17 18:39:53 +08:00
BM Cho 721bfcb325 docs: add Traditional Chinese README 2026-07-17 17:03:49 +08:00
Tonu Samuel 81a56777df glm.c: detect 9p by statfs f_type, not the "/mnt/" path prefix
The slow-filesystem warning fired for any model path under /mnt/, which
false-positives on native-Linux mounts (ZFS/ext4/xfs/NFS) that commonly
live there. Check the actual filesystem type via statfs() against the 9p
superblock magic (0x01021997) instead, so it only warns for a genuine
WSL 9p mount. Linux-only (statfs); no behavior change on other platforms.
2026-07-17 08:00:26 +03:00
Nicholas Beerbower ed8dab4da4 test_st_pread: Windows-compatible — relative tmpdir, fork subtest gated POSIX
Same fix pattern as test_stops: mkdtemp with a CWD-relative template
(MinGW resolves Windows paths; /tmp is not one), and the fork/pipe/
truncate-based truncation subtest compiles only where those exist —
Windows still runs the cross-platform chunk-loop content check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 16:32:43 -04:00
Nicholas Beerbower 6507817d9f st: chunked pread with EINTR retry and honest short-read errors
Two latent bugs in every st.h reader, both hit in the field:

 - a single pread caps at ~2^31 bytes on Linux, so any tensor past
   2.1 GB (bf16 embed/unembed tensors of large models qualify) came
   back silently truncated with perror printing '... : Success'
   (errno untouched by a short read) — the same misleading-error
   symptom glm.c fixed for its own reads in #236;
 - no EINTR retry.

st_pread_full loops in ST_PREAD_CHUNK pieces (1 GB default, override
for tests), retries EINTR, and reports offset + progress on failure.
All five read sites converted; behavior on well-formed files is
byte-identical (GLM oracle re-verified on this branch: 32/32).

tests/test_st_pread builds with -DST_PREAD_CHUNK=7 so a 96-byte tensor
exercises the multi-chunk loop, and forks a child against a shard
truncated after st_init (init's static bounds check means the pread
path only fires when a file shrinks under a live handle) asserting
exit(1) with a 'short read' message and no 'Success'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 15:20:43 -04:00
JustVugg d45cfa268a Merge branch 't321' into trial321b
# Conflicts:
#	README.md
2026-07-16 21:11:23 +02:00
Mohamed Mastouri 69284facb8 docs: Windows 11 native install walkthrough - toolchain, SAC, CUDA DLL, failure index (#198) 2026-07-16 21:59:20 +03:00
Mohamed Mastouri 80f886cd22 convert_olmoe.py: wire --ebits through to quantization (#323) 2026-07-16 21:59:16 +03:00
Vincenzo cead3d8bf3 Merge pull request #324 from dawnfield-institute/feat/olmoe-ppl-mode
olmoe: PPL=1 teacher-forced perplexity mode (loss meter for throughput experiments)
2026-07-16 20:47:37 +02:00
Vincenzo 57030ab231 Merge pull request #327 from lEWFkRAD/ci/windows-cuda-build
ci: build the Windows CUDA path (nvcc + MSVC host) — the one toolchain no job covers
2026-07-16 20:47:33 +02:00
lEWFkRAD b75422d605 ci: build the Windows CUDA path (nvcc + MSVC host) — the one toolchain no job covers
engine-cuda-syntax compiles backend_cuda.cu with nvcc's GCC host on Linux, and
check.yml's windows job is MinGW/UCRT64 CPU-only by design (#140). Neither
exercises nvcc's MSVC host — the only host compiler nvcc accepts on Windows — so
`make cuda-dll` and `make glm CUDA_DLL=1` are currently built by nothing.

The gap has already cost real bugs, all compile-time, all catchable without a GPU:

  - #158: MSVC rejects the GCC-style -Xcompiler=-Wall,-Wextra ("D8021 invalid
    numeric argument '/Wextra'"), CUDA_HOME/NVCC defaults were unresolvable, and
    the kernel test used POSIX setenv. `make cuda-dll` was unbuildable as shipped.
  - #314: CUDA_HOME containing spaces — the CUDA installer's default layout. This
    job installs to exactly that path, so it reproduces the condition by default.

Two Windows-specific facts the job encodes, both verified by making it fail first:

  - sub-packages needs "cudart", not just "nvcc": cuda-dll *links* the runtime, and
    on Windows the headers/import lib are a separate installer component. With
    '["nvcc"]' alone it dies at `#include <cuda_runtime.h>` — the Linux syntax job
    never notices because it only compiles.
  - runs-on is pinned to windows-2022: windows-latest now ships Visual Studio 18
    (MSVC 14.5x) and CUDA's crt/host_config.h hard-errors on any host newer than
    VS 2022. That is a real constraint on every Windows CUDA user today, so the pin
    tracks what the toolkit supports rather than papering over it with
    -allow-unsupported-compiler.

Build-only on purpose: hosted runners have no NVIDIA device, so this proves the
Windows+MSVC CUDA build stays buildable, not that the kernels or the DLL loader
behave on real silicon — that still needs hardware (#157). A check that overclaims
buys the false confidence the engine-cuda-syntax comment already warns about.

Verified green on a fork before proposing:
https://github.com/lEWFkRAD/colibri/actions/runs/29524302993

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 14:35:30 -04:00
JustVugg e7188df16a tests: test_stops must not assume /tmp exists (fixes the windows job)
My test_stops.c called mkdtemp("/tmp/coli_stops_XXXXXX"). These binaries are
built by MinGW into NATIVE Windows .exe files, which resolve Windows paths —
"/tmp" is not one, so mkdtemp failed ENOENT and `make check` went red on the
windows job the moment #143 gave us cross-platform CI.

Now a relative template in the CWD, which is what test_compat_direct.c already
does (`#define TMPF "test_direct.tmp"`). test_uring.c uses /tmp but is
Linux-only by construction (Makefile guards it behind $(LINUX)); I copied the
wrong neighbour.

Worth being precise about what this was, because the red job looked scarier
than it was: Windows is fine. `make check` built the engine, ran the whole C
suite, and passed everything else — test_i4_grouped, kv_alloc, compat_direct —
before tripping on my temp path. The CI caught a bug in the test, not in the
port. That is #143 earning its keep on day one.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 20:08:02 +02:00
Vincenzo ee5d273cd4 Merge pull request #232 from nbeerbower/profiling-upstream
Profiling: PROF=1 opt-in performance profile + live per-turn Profiling page in the web dashboard
2026-07-16 19:58:31 +02:00
JustVugg 98f7c88ca9 numa: parenthesise the page-align arithmetic (#313)
`(uintptr_t)p+n+4095 & ~(uintptr_t)4095` is CORRECT — `+` binds tighter than
`&` in C, so the rounding happens before the mask, which is what was meant.
But gcc -Wall warns (`suggest parentheses around '+' in operand of '&'`), and
this repo builds clean at -Wall -Wextra. Warning-free is a property worth more
than the two characters it costs to keep: it is what makes a NEW warning
visible instead of scrolling past in a wall of noise.

No behaviour change; the emitted code is identical.

Co-Authored-By: ZacharyZcR <#313>

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 19:50:20 +02:00
Vincenzo 3cec90cd45 Merge pull request #313 from ZacharyZcR/feat/numa-expert-interleave
glm: COLI_NUMA=1 — interleave resident weights across NUMA nodes (+13% decode on 2-socket, raw mbind, zero deps) (#82)
2026-07-16 19:49:17 +02:00
Vincenzo 2dca36a8a0 Merge pull request #314 from mohamedmastouri2000-boop/fix/windows-cuda-dll-build
Add build-config stamp: flag changes (e.g. CUDA_DLL=1) force a relink instead of a silent stale binary
2026-07-16 19:49:14 +02:00
Vincenzo 644e0d16c6 Merge pull request #302 from alekseysorokin68/main
Fix: set stdout to O_BINARY on Windows to fix READY sentinel
2026-07-16 19:49:11 +02:00
Vincenzo 59988c41c8 Merge pull request #143 from lEWFkRAD/ci/make-check-matrix
ci: make check on ubuntu / windows (MSYS2 UCRT64) / macos — and it already caught a Windows test bug (#140)
2026-07-16 19:46:27 +02:00
Vincenzo 545d9d0ab5 Merge pull request #317 from ZacharyZcR/docs/serve-protocol
docs: serve protocol reference — the engine⇄server wire format (#310)
2026-07-16 19:46:24 +02:00
Vincenzo b04c2ca50e Merge pull request #322 from woolcoxm/fix/oracle-transformers-pin
oracle: hard-fail on transformers < 5.11.0 (interleaved-RoPE floor, #281)
2026-07-16 19:46:20 +02:00
Vincenzo c686f6bd51 Merge pull request #320 from Magnet-js/docs/winget-python-readme
docs: add winget python one-liner to Windows install block (#310)
2026-07-16 19:46:17 +02:00
Vincenzo 95b059b15b Merge pull request #318 from tt1203/fix/oracle-mkdir-glm-tiny
fix: make_glm_oracle.py creates glm_tiny/ before saving (fresh-checkout failure)
2026-07-16 19:46:13 +02:00
JustVugg 78c77bf6c8 ci: make the CUDA job able to fail, and able to run (#144)
Two independent defects in the CUDA syntax job, both caught on its first run.

1. `cuda: '12.6.3'` — Jimver/cuda-toolkit@v0.2.19's version table stops at
   12.6.2, so the install step died with "Version not available: 12.6.3"
   before nvcc was ever invoked. One digit.

2. `nvcc ... 2>&1 | head -40` — a pipeline exits with the status of its LAST
   command, so head's 0 masked every nvcc error. The job printed "CUDA syntax
   check passed" unconditionally: it could not fail. Defect 1 is why we found
   out, since it broke the step *before* the pipe.

The second one is the one that matters. backend_cuda.cu is compiled by nothing
else in this repo — no local build, no test — so this job is the only thing
standing between a CUDA change and a user's GPU. A check that cannot fail is
worse than no check: it buys false confidence in exactly the file that most
needs the real thing. #298 spent a night debugging CUDA against no oracle at
all; this job is supposed to be that oracle.

It is also, precisely, the disease of the week in YAML form: a signal that
measures its own intention rather than the thing it claims to measure. See the
`~335 GB I/O saved` counter that counted dropped experts instead of bytes not
read (#303), and `route_agree: 95.3%` cited as "quality preserved" when it
only measures which experts coincide.

The other three jobs (engine, web, python) passed on the first run.

Co-Authored-By: ZacharyZcR <#144>

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 19:36:31 +02:00
Vincenzo 5272d19abc Merge pull request #144 from ZacharyZcR/ci/github-actions
ci: GitHub Actions — engine, web, Python, CUDA syntax (free for public repos)
2026-07-16 19:32:18 +02:00
ZacharyZcR 7bc9f6e2e9 docs/media: refreshed dashboard shots (full-residency 4 tok/s, disk 0), new Atlas galaxy, phone + sidebar shots 2026-07-17 00:29:50 +08:00
Peter Groom c5f2027fa1 olmoe: PPL=1 teacher-forced perplexity mode (loss meter for throughput experiments)
Feeds reference tokens through the normal step() decode path and reports
NLL/ppl + hit rate + tok/s + RSS. Inert unless PPL=1; default path
untouched (12/12 vs ref.json verified with patch applied). Cross-checked
vs HF transformers bf16 on identical token ids: 12.11 vs 12.25 ppl (#108).
2026-07-16 13:10:50 -03:00
mohamedmastouri2000-boop c70d94368e Merge branch 'main' into fix/windows-cuda-dll-build 2026-07-16 19:05:43 +03:00