From a55cdfd4b800c65f8b12d97c6834bf82952866ff Mon Sep 17 00:00:00 2001 From: woolcoxm <13604288+woolcoxm@users.noreply.github.com> Date: Wed, 15 Jul 2026 15:36:10 -0400 Subject: [PATCH] cuda: gate pipe2 S-threshold on device count (single-GPU S=1, multi-GPU S>=8) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Per ZacharyZcR's 6x5090 A/B on #273: the S=1 resident-pipeline relaxation is +49% on a single GPU (5070 Ti) but a wash on multi-GPU. With layers sharded across N devices, each resident forward at S=1 crosses P2P per layer group and those small hops don't amortize — the same term that killed pipe x head-shard in #111. Multi-GPU decode walls on disk service, which pipe2 can't touch. Make the threshold device-count-dependent: - single GPU (g_cuda_ndev<=1): S>=1 (the breakthrough path) - multi GPU (g_cuda_ndev> 1): S>=8 (the original prefill-only gate) with COLI_CUDA_PIPE_S_MIN as an env override for anyone who wants to measure. The two calibration points bracket the design space: 1x 5070 Ti, modest CPU, "other"-bound decode -> S=1 (+49%) 6x 5090, sharded, disk-service-bound -> S=1 a wash, keep S>=8 Refs #273 (comment) --- c/glm.c | 14 ++++++++++++-- 1 file changed, 12 insertions(+), 2 deletions(-) diff --git a/c/glm.c b/c/glm.c index e5c26d9..c75fcf2 100644 --- a/c/glm.c +++ b/c/glm.c @@ -3198,10 +3198,20 @@ static void layers_forward_rows(Model *m, float *x, int S, int pos_base, float *nrm=falloc((int64_t)S*D), *tmp=falloc((int64_t)S*D); #ifdef COLI_CUDA /* PIPE2 (Inc.2a): il residuo resta sul device del layer, saltando tra le schede - * ai confini di layer. x host diventa STALE finche' la residenza e' attiva. */ + * ai confini di layer. x host diventa STALE finche' la residenza e' attiva. + * + * S threshold is device-count-dependent (#273): on a single GPU the resident + * stream wins at S=1 (evicts the CPU round-trips that dominate small-batch + * decode — +49% on a 5070 Ti). With layers sharded across multiple GPUs each + * resident forward crosses P2P per layer group, and at one token per forward + * those hops don't amortize — A/B on 6x5090 showed S=1 is a wash there. So: + * single-GPU engages at S=1, multi-GPU keeps the original S>=8 prefill gate. + * COLI_CUDA_PIPE_S_MIN overrides for anyone who wants to measure. */ float *x_dev=NULL; int x_dev_on=-1; size_t xb=(size_t)S*(size_t)D*4; - int pipe2 = g_cuda_pipe>=2 && !kvs && S>=1 && g_cuda_enabled && c->kv_lora<=512 && + int pipe_s_min = getenv("COLI_CUDA_PIPE_S_MIN") ? atoi(getenv("COLI_CUDA_PIPE_S_MIN")) + : (g_cuda_ndev<=1 ? 1 : 8); + int pipe2 = g_cuda_pipe>=2 && !kvs && S>=pipe_s_min && g_cuda_enabled && c->kv_lora<=512 && !(m->has_dsa && pos_base+S>c->index_topk); #endif for(int i=0;in_layers;i++){