Sparse AttentionA transformer efficiency technique where each token attends only to a relevant subset of other tokens, enabling near-linear scaling with context length.3 Apr 2026architecturetransformersefficiencylong-contextresearch
Test-Time Training (TTT)A method where a model runs gradient updates during inference, using its current input as training data to adapt on the fly.3 Apr 2026traininginferencememoryefficiencyresearch
AI digest: efficiency breakthroughs and infrastructure reality checksMajor advances in model efficiency clash with infrastructure growing pains as the AI boom hits practical limits.25 Mar 2026ai-newsdigestefficiencyinfrastructure