ec89136029
* Fuse CUDA expert MLP execution * Group CUDA expert transfers by device * Instrument grouped CUDA expert execution * Bound grouped CUDA decode scratch * Execute expert groups across GPUs in parallel * Release host backing for multi-GPU experts * Define quality-preserving memory policies * Overlap cold expert loading with resident compute * Adapt expert placement with session LFRU * Fuse q4 expert gate and up dispatch * Plan CPU work on physical cores * Batch grouped expert CUDA kernels * Separate VRAM and RAM expert placement * Add ragged multi-sequence decode forward * feat(runtime): add continuous decode scheduler * Route concurrent API requests through batch scheduler * Harden multiplex request lifecycle and framing * Cancel disconnected multiplex requests * Bind API port before starting the engine * fix automatic KV slot allocation * add native int4 Tensor Core grouped GEMM * add Tensor Core throughput benchmark * optimize packed int4 low-row kernels * add asynchronous CUDA staging streams * document validated six-GPU dense acceleration * tune six-GPU expert hot set * raise validated expert hot-set target * add CUDA MLA absorption core * fuse grouped expert gate and up projections * Warn for explicit lossy routing flags * Add full-resident expert placement mode * Adapt VRAM expert slots to live routes * Accelerate int4 matvec on AVX-512 * Reduce AVX-512 and RoPE decode overhead * Seed every GPU expert layer after prefill * Limit live GPU swaps during decode * CUDA batch MLA attention, kv_b head-sharding, fused o_proj, expert-group dispatch, W4A16 kernels Lab-qualified on the 6x RTX 5090 machine (914-token request benchmark): - batch MLA absorption kernel (COLI_CUDA_ATTN=1): whole-batch attention on device, 154.8s -> 102.4s - attention -> o_proj fusion on the layer device: -> 97.4s - kv_b head-sharding across cards (COLI_CUDA_ATTN_SHARD=1), no weight duplication: -> 94.05s - per-device expert-group dispatch with pinned-buffer async transfers, W4A16 tensor-core kernels for the shared expert, OMP hot-thread tuning Negative results (reverted, kept out): GPU-side weighted scatter-add (atomics + per-layer D2H lose 43.8%), shared-expert fused small-batch kernel (-38.8%), W4A4 grouped tensor cores (int4 activations corrupt output). Details in the lab research log. * GPU resident pipeline: device-resident prefill attention chain, GPU expert groups in prefill, batched router, W4A16 mixed dispatch COLI_CUDA_PIPE=1 keeps the prefill data plane on the layer home device; control flow (routing, cache/pin management) stays on CPU. Any CUDA failure falls back to the unchanged CPU path. - Device primitives + unit tests (tests/test_pipe_cuda.cu): rmsnorm (strided), interleaved RoPE, silu-mul, residual add, fixed-order row merge (no atomics), device-input GEMM, persistent per-device scratch. All verified against the engine's CPU math on SM120 (worst 1.2e-5). - attn_pipe_prefill: q_a -> norm -> q_b -> rope -> kv_a -> norm -> rope -> batch attention -> o_proj in one device chain (q_a/q_b/kv_a colocated with kv_b); only the final [S,D] and the new KV rows return to host. Attention 41.2s -> 30.8s on the 1571-token benchmark. - Prefill batch-union now uses the GPU expert groups (previously gated to S<=64, leaving all VRAM-resident experts idle during prefill - measured 21ms of GPU expert time in a 148s prefill). Expert phase 78.9s -> 69.0s. - Router computed as one batched matmul instead of S sequential rows (bit-identical math). - W4A16 tensor-core path for expert groups (COLI_CUDA_TC_W4A16=1) with row-count mixed dispatch: >=16 rows per expert use tensor cores, smaller batches keep the naive kernel (tensor cores measured negative below ~16 rows). Expert phase 69.0s -> 64.3s, decode unaffected. Net on the 1571-token prefill benchmark: 148.8s -> 114.3-126.8s (component timings stable across runs; wall drifts +-3-5s because .coli_usage placement learning shifts the expert tiers between runs). PROFILO now also prints the prefill-phase breakdown. * Skip OMP hot-thread tuning when CUDA is enabled The active-spin worker team measured 66.9s->20.9s on the CPU-only Zen5 build, but on the six-GPU full-residency workload the spinning workers contend with the CUDA dispatch threads: ~4x slower prefill with the process stuck near 1.8 cores. Gate the tuning on COLI_CUDA so each configuration keeps the behavior it was measured to prefer. * Inc.2a: sparse layers fully resident on the layer device, residual hops cards at layer boundaries COLI_CUDA_PIPE=2 keeps the residual stream on the layer home device for consecutive sparse layers (cudaMemcpyPeer at boundaries): in/post norms, attention chain, both residual adds and the shared-expert MLP run on device. Per layer only the post-norm activations (router + CPU-tier experts + group gather), the new KV rows and, on DSA indexer layers, the pre-attention norm leave the card. Per-layer transfers drop from ~130MB to ~70MB. A device-side snapshot at layer entry makes any mid-layer CUDA failure fall back to the unchanged CPU path idempotently. 1571-token prefill: 127.1s (PIPE=1 control) -> 117.6/118.9s, components attention 30.8->26.1, other 31.8->22.5-24.5; output verified coherent against the control. * Head-sharded attention inside the pipe: negative on PCIe star topology, gated opt-in Slicing q per card from the home device and collecting ctx back serializes ~95MB/layer through the home card's PCIe link: attention 26.1s -> 41.4/44.4s on the 1571-token benchmark (two repeats), wall 117.6 -> 135-138s. The standalone host-path sharding won because six cards uploaded from host RAM in parallel; a home-device star has no such parallelism without NVLink. Kept behind COLI_CUDA_PIPE_SHARD=1 for interconnects where peer bandwidth does not share one root port. * Inc.3: device-resident KV shadow for decode attention Decode re-uploaded the whole latent+rope window per layer per token (~300MB/token at 1571 context). Each layer now keeps a device shadow of the compressed KV on its kv_b card, bulk-synced when behind and appended incrementally; the host cache stays canonical. Invalidation on kv_bind (slot switch), kv_alloc (resize) and on any overwrite of mirrored rows, with the legacy full-upload path as fallback. Measured (COLI_CUDA_PIPE gate): short-context decode 5.48 -> 5.59/5.87 tok/s, 1571-context decode 4.14 -> 4.22 tok/s. Decode remains CPU-expert bound; the shadow removes the transfer tax, not the compute. * tools: unified user-experience benchmark (bench_ux.sh) Two fixed scenarios (short chat, long-document QA), TTFT + decode tok/s + first-line drift check, TEMP=0 DRAFT=0 enforced, medians over REPS runs. Encodes the measurement discipline from the lab record: same binary per comparison, judge medians because .coli_usage placement learning drifts wall times between runs. * tools: bench_ux.sh executable bit * gitignore compiled test binaries * tools: expert_atlas.py — measure per-expert topic affinity (#175) Diffs .coli_usage across 10 themed probe batches (code/math/chinese/ prose/science/law/poetry/structured/translation/casual, 3 prompts each) driven through a running API server — one engine load total. Every touched expert gets a topic-affinity vector, entropy, and a specialist/ generalist label; output experts.json feeds the Brain page hover. * serve: persist .coli_usage after every turn in mux mode, not only at exit run_serve_mux saved the learning cache once at shutdown; a crash lost the whole session's routing history, and live consumers of the file (expert_atlas.py diffs it between probe batches) saw a frozen snapshot. Now saved per turn like the interactive path (165KB write, negligible). * web: Brain hover shows measured expert atlas when published If /experts.json (from tools/expert_atlas.py, #175) is served next to the app, the tooltip upgrades from the depth heuristic to measured data: specialist/generalist label, entropy, and the top-3 topic affinities. Row index maps to real layer (row+3, last row = MTP 78). Falls back to the heuristic when no atlas is published. --------- Co-authored-by: JustVugg <JustVugg@users.noreply.github.com>
147 lines
7.8 KiB
C
147 lines
7.8 KiB
C
#ifndef COLIBRI_BACKEND_CUDA_H
|
|
#define COLIBRI_BACKEND_CUDA_H
|
|
|
|
#include <stddef.h>
|
|
#include <stdint.h>
|
|
|
|
/* COLI_CUDA_DLLEXPORT marks functions exported from coli_cuda.dll on Windows.
|
|
* Define COLI_CUDA_BUILDING_DLL when compiling the .cu into the DLL (so the
|
|
* functions are __declspec(dllexport)); the host loader does NOT include this
|
|
* header's declarations — it resolves symbols at runtime via GetProcAddress. */
|
|
#if defined(_WIN32) && defined(COLI_CUDA_BUILDING_DLL)
|
|
#define COLI_CUDA_DLLEXPORT __declspec(dllexport)
|
|
#else
|
|
#define COLI_CUDA_DLLEXPORT
|
|
#endif
|
|
|
|
#ifdef __cplusplus
|
|
extern "C" {
|
|
#endif
|
|
|
|
#define COLI_CUDA_MAX_DEVICES 16
|
|
|
|
/* Opaque, persistent device copy of one resident quantized tensor. */
|
|
typedef struct ColiCudaTensor ColiCudaTensor;
|
|
|
|
/* Devices are CUDA ordinals, not positions in the input list. */
|
|
COLI_CUDA_DLLEXPORT int coli_cuda_init(const int *devices, int count);
|
|
COLI_CUDA_DLLEXPORT void coli_cuda_shutdown(void);
|
|
COLI_CUDA_DLLEXPORT int coli_cuda_device_count(void);
|
|
COLI_CUDA_DLLEXPORT int coli_cuda_device_at(int index);
|
|
COLI_CUDA_DLLEXPORT int coli_cuda_mem_info(int device, size_t *free_bytes, size_t *total_bytes);
|
|
/* device < 0 returns aggregate statistics for all configured devices. */
|
|
COLI_CUDA_DLLEXPORT void coli_cuda_stats(int device, size_t *tensor_count, size_t *tensor_bytes);
|
|
COLI_CUDA_DLLEXPORT void coli_cuda_group_stats(uint64_t *calls, uint64_t *experts, uint64_t *rows,
|
|
double *h2d_ms, double *kernel_ms, double *d2h_ms);
|
|
|
|
/* Upload without executing, so capacity failures happen during model startup. */
|
|
COLI_CUDA_DLLEXPORT int coli_cuda_tensor_upload(ColiCudaTensor **tensor,
|
|
const void *weights, const float *scales,
|
|
int fmt, int I, int O, int device);
|
|
|
|
/*
|
|
* y[S,O] = x[S,I] @ W[O,I]^T.
|
|
* fmt matches QT in glm.c: 0=f32, 1=int8, 2=int4, 3=int2.
|
|
* The first successful call uploads W and its row scales; later calls reuse it.
|
|
* Returns 1 on success and 0 when CUDA is not initialized or the format is invalid.
|
|
*/
|
|
COLI_CUDA_DLLEXPORT int coli_cuda_matmul(ColiCudaTensor **tensor,
|
|
float *y, const float *x,
|
|
const void *weights, const float *scales,
|
|
int fmt, int S, int I, int O, int device);
|
|
|
|
/* Fused expert pipeline: y = down(silu(gate(x)) * up(x)). All three tensors
|
|
* must already be resident on one device. Activations cross PCIe once in
|
|
* each direction instead of once per matrix. */
|
|
COLI_CUDA_DLLEXPORT int coli_cuda_expert_mlp(ColiCudaTensor *gate, ColiCudaTensor *up,
|
|
ColiCudaTensor *down, float *y, const float *x, int S);
|
|
|
|
/* Prefill-oriented shared expert path. INT4 weights stay packed in global
|
|
* memory, activations are converted to FP16 per tile, and Tensor Cores
|
|
* accumulate into FP32. Unlike COLI_CUDA_TC_INT4 this does not quantize the
|
|
* activation to INT4. */
|
|
int coli_cuda_shared_mlp_w4a16(ColiCudaTensor *gate, ColiCudaTensor *up,
|
|
ColiCudaTensor *down, float *y,
|
|
const float *x, int S);
|
|
|
|
/* Packed group of same-shaped experts. Inputs and outputs contain sum(rows)
|
|
* consecutive [D] rows in call order. */
|
|
COLI_CUDA_DLLEXPORT int coli_cuda_expert_group(ColiCudaTensor *const *gates,
|
|
ColiCudaTensor *const *ups,
|
|
ColiCudaTensor *const *downs,
|
|
const int *rows, int count,
|
|
float *y, const float *x);
|
|
|
|
/* Decode-only MLA weight-absorption core for one token. kv_b is [H*(Q+V),K]. */
|
|
COLI_CUDA_DLLEXPORT int coli_cuda_attention_absorb(ColiCudaTensor *kv_b,float *ctx,const float *q,
|
|
const float *latent,const float *rope,int H,int Q,
|
|
int R,int V,int K,int T,float attention_scale);
|
|
|
|
/* Causal MLA absorption for S contiguous rows from one sequence. The KV
|
|
* arrays contain T rows ending at the final query; query s attends T-S+s+1
|
|
* rows. One transfer and one launch replace S host round-trips. */
|
|
COLI_CUDA_DLLEXPORT int coli_cuda_attention_absorb_batch(ColiCudaTensor *kv_b,float *ctx,const float *q,
|
|
const float *latent,const float *rope,int S,
|
|
int H,int Q,int R,int V,int K,int T,
|
|
float attention_scale);
|
|
|
|
/* Same attention batch followed immediately by resident o_proj on the same
|
|
* device. Only the final [S,D] tensor crosses back to the host. */
|
|
COLI_CUDA_DLLEXPORT int coli_cuda_attention_project_batch(ColiCudaTensor *kv_b,ColiCudaTensor *o_proj,
|
|
float *out,const float *q,const float *latent,
|
|
const float *rope,int S,int H,int Q,int R,
|
|
int V,int K,int T,float attention_scale);
|
|
|
|
COLI_CUDA_DLLEXPORT void coli_cuda_tensor_free(ColiCudaTensor *tensor);
|
|
COLI_CUDA_DLLEXPORT size_t coli_cuda_tensor_bytes(const ColiCudaTensor *tensor);
|
|
COLI_CUDA_DLLEXPORT int coli_cuda_tensor_device(const ColiCudaTensor *tensor);
|
|
|
|
/* Replace a resident tensor's contents without reallocating its device slot. */
|
|
int coli_cuda_tensor_update(ColiCudaTensor *tensor,
|
|
const void *weights, const float *scales);
|
|
|
|
/* ---- resident-pipeline primitives (Inc.0): device-pointer entry points ---- */
|
|
float *coli_cuda_pipe_scratch(int device,int slot,size_t bytes);
|
|
void *coli_cuda_pipe_alloc(int device,size_t bytes);
|
|
void coli_cuda_pipe_free(int device,void *p);
|
|
int coli_cuda_pipe_upload(int device,void *dst,const void *src,size_t bytes);
|
|
int coli_cuda_pipe_download(int device,const void *src,void *dst,size_t bytes);
|
|
int coli_cuda_pipe_rmsnorm(int device,float *y_dev,const float *x_dev,
|
|
const float *w_dev,int S,int D,float eps);
|
|
int coli_cuda_pipe_rope(int device,float *v_dev,const int *pos_dev,int rows,
|
|
int stride,int offset,int R,int heads,float theta);
|
|
int coli_cuda_pipe_silu_mul(int device,float *gate_dev,const float *up_dev,size_t n);
|
|
int coli_cuda_pipe_add(int device,float *x_dev,const float *t_dev,size_t n);
|
|
int coli_cuda_pipe_rows_add(int device,float *x_dev,const float *partial_dev,
|
|
const int *rows_dev,int nrows,int D);
|
|
int coli_cuda_pipe_gemm(ColiCudaTensor *t,float *y_dev,const float *x_dev,int S);
|
|
int coli_cuda_pipe_rmsnorm_s(int device,float *y_dev,const float *x_dev,
|
|
const float *w_dev,int S,int D,float eps,
|
|
int xstride,int ystride);
|
|
int coli_cuda_pipe_rope_base(int device,float *v_dev,int pos_base,int rows,
|
|
int stride,int offset,int R,int heads,float theta);
|
|
int coli_cuda_pipe_copy2d(int device,float *dst,int dpitch,const float *src,
|
|
int spitch,int width,int height);
|
|
int coli_cuda_attention_project_batch_dev(ColiCudaTensor *kv_b,ColiCudaTensor *o_proj,
|
|
float *out,const float *q_dev,const float *latent_dev,const float *rope_dev,
|
|
int S,int H,int Q,int R,int V,int K,int T,float scale);
|
|
int coli_cuda_attention_absorb_batch_dev(ColiCudaTensor *kv_b_shard,float *ctx_dev,
|
|
const float *q_dev,const float *latent_dev,const float *rope_dev,
|
|
int S,int H,int Q,int R,int V,int K,int T,float scale);
|
|
int coli_cuda_attention_absorb_kvdev(ColiCudaTensor *kv_b,float *ctx,const float *q,
|
|
const float *latent_dev,const float *rope_dev,int H,int Q,int R,int V,int K,int T,
|
|
float scale);
|
|
int coli_cuda_pipe_peer_copy(int dst_dev,float *dst,int src_dev,
|
|
const float *src,size_t bytes);
|
|
int coli_cuda_attention_project_batch_dev_out(ColiCudaTensor *kv_b,ColiCudaTensor *o_proj,
|
|
float *out_dev,const float *q_dev,const float *latent_dev,const float *rope_dev,
|
|
int S,int H,int Q,int R,int V,int K,int T,float scale);
|
|
int coli_cuda_pipe_sync(int device);
|
|
|
|
#ifdef __cplusplus
|
|
}
|
|
#endif
|
|
|
|
#endif
|
|
|