113ece3bc7
First increment of the #431 plan (device router -> indirect kernels -> one-graph decode). At decode (S=1) on the pipe2 path, the router runs on the layer's home device: a tiny E x D logits GEMV + sigmoid, then a single-thread selection kernel that clones moe()'s plain routing path verbatim — bias-augmented top-K by choice with strict-> tie-breaking, weights from the raw logit, route-level TOPP truncation, norm_topk, routed_scale. Results pack into one scratch buffer and come back in a single ~68-byte D2H; moe() consumes them through the same pre-routed shortcut the Metal layer-CB uses (g_pre_idx, #417 bookkeeping included), so usage/heat/recency accounting is identical to the CPU router. Structural value: routing becomes available ON the device timeline, which is what PR-B (indirect expert kernels, static topology) and PR-C (whole-decode CUDA Graph) build on. Opt-in, default off. Gated to the plain routing path — CACHE_ROUTE, ROUTE_P and ROUTE_TRACE keep the CPU ranking they need; any upload or launch failure falls back to the CPU router silently. Router weights (E x D f32, ~6.3 MB/layer) upload lazily to the layer's home device. tests/test_router_cuda.cu: kernel-vs-CPU-reference oracle over 200 random trials (mixed TOPP/norm_topk/scale): 200/200 exact selections, zero near-tie flips, zero weight mismatches on a 5090.