a55cdfd4b8
Per ZacharyZcR's 6x5090 A/B on #273: the S=1 resident-pipeline relaxation is +49% on a single GPU (5070 Ti) but a wash on multi-GPU. With layers sharded across N devices, each resident forward at S=1 crosses P2P per layer group and those small hops don't amortize — the same term that killed pipe x head-shard in #111. Multi-GPU decode walls on disk service, which pipe2 can't touch. Make the threshold device-count-dependent: - single GPU (g_cuda_ndev<=1): S>=1 (the breakthrough path) - multi GPU (g_cuda_ndev> 1): S>=8 (the original prefill-only gate) with COLI_CUDA_PIPE_S_MIN as an env override for anyone who wants to measure. The two calibration points bracket the design space: 1x 5070 Ti, modest CPU, "other"-bound decode -> S=1 (+49%) 6x 5090, sharded, disk-service-bound -> S=1 a wash, keep S>=8 Refs #273 (comment)