Meet mKernel: A Multi-GPU, Multi-Node Fused Kernel Library for GPU-Driven Communication
GPU communication overhead is a measurable bottleneck in manufacturing AI workloads. According to knowledge cited by the mKernel undertaking, communication can devour 43.6% of the ahead go and 32% of end-to-end coaching time. Across fashionable Mixture-of-Experts (MoE) fashions, inter-device communication can account for as much as 47% of whole execution time. Researchers from UC Berkeley’s…
