af421d6d3a
dot_i8i8 and dot_i4i8 accumulated the whole SDOT reduction into a single int32x4_t. SDOT has ~3-4 cycle latency, so the serial dependency on `acc` capped each core at ~26 GB/s (int8) / ~12 GB/s (int4) of weight throughput regardless of memory bandwidth. Split into 4 independent accumulators (64 values/iter) so the loads become the bottleneck instead of the reduction chain; the original single-acc loop is kept as the tail handler. Measured on an Apple M4 (isolated microbench, expert-shaped 2048x6144): int8*int8 26.0 -> 63.2 GB/s/core (2.4x) int4*int8 12.4 -> 29.9 GB/s/core (2.4x) Output is bit-identical to the previous kernels (verified over random inputs). Non-DOTPROD NEON, AVX2/AVX-512/VNNI and VSX paths are untouched; only the __ARM_FEATURE_DOTPROD branch changed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>