Nous Research Proposes Lighthouse Attention: A Training-Only Selection-Based Hierarchical Attention That Delivers 1.4–1.7× Pretraining Speedup at Long Context
Training massive language fashions on lengthy sequences has a well known drawback: consideration is dear. The scaled dot-product consideration (SDPA) at the core of each transformer scales quadratically Θ(N²) in each compute and reminiscence with sequence size N. FlashAttention addressed this by IO-aware tiling that avoids materializing the complete N×N consideration matrix in high-bandwidth reminiscence,…
