MiniMax Sparse Attention (MSA): a Two-Branch Block-Sparse Attention Trained on a 109B-Parameter MoE With a 3T-Token Budget
MiniMax launched MSA (MiniMax Sparse Attention), a sparse consideration methodology constructed instantly on Grouped Query Attention (GQA). It targets one bottleneck: the quadratic value of softmax consideration at lengthy context. The MiniMax analysis workforce examined it inside a 109B-parameter Mixture-of-Experts mannequin skilled with native multimodal knowledge. They additionally open-sourced an inference kernel and shipped a…
