Files
colibri/c
woolcoxm a55cdfd4b8 cuda: gate pipe2 S-threshold on device count (single-GPU S=1, multi-GPU S>=8)
Per ZacharyZcR's 6x5090 A/B on #273: the S=1 resident-pipeline relaxation is
+49% on a single GPU (5070 Ti) but a wash on multi-GPU. With layers sharded
across N devices, each resident forward at S=1 crosses P2P per layer group and
those small hops don't amortize — the same term that killed pipe x head-shard
in #111. Multi-GPU decode walls on disk service, which pipe2 can't touch.

Make the threshold device-count-dependent:
  - single GPU  (g_cuda_ndev<=1): S>=1  (the breakthrough path)
  - multi  GPU  (g_cuda_ndev> 1): S>=8  (the original prefill-only gate)
with COLI_CUDA_PIPE_S_MIN as an env override for anyone who wants to measure.

The two calibration points bracket the design space:
  1x 5070 Ti, modest CPU, "other"-bound decode  -> S=1 (+49%)
  6x 5090,    sharded,   disk-service-bound      -> S=1 a wash, keep S>=8

Refs #273 (comment)
2026-07-15 15:36:10 -04:00
..