Sparse AttentionA transformer efficiency technique where each token attends only to a relevant subset of other tokens, enabling near-linear scaling with context length.3 Apr 2026architecturetransformersefficiencylong-contextresearch