What comes after attention? This startup says it already knows.

This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).

Subquadratic, a startup that launched earlier this year with a sparse-attention model supporting 12-million token context windows, has published its first model card and benchmarks for SubQ 1.1 Small. CTO Alex Whedon clarifies the company is not purely a sparse-attention company — it is already developing 'zero attention' architectures that drop the attention mechanism entirely, drawing inspiration from Yann LeCun's world model research. The SubQ 1.1 Small model achieves near-perfect scores on long-context retrieval benchmarks (99.12% on RULER at 128K tokens), claims 64.5x less compute than dense attention at 1 million tokens, and runs 56x faster than FlashAttention-2. General capability lands near mid-tier frontier models. The model was built by replacing dense attention in an existing open-weight model with SSA and running ~1 trillion tokens of long-context pretraining. Current access is limited to enterprise design partners with large AI spend, with a limited individual release planned before general availability. The next model is expected to be mid-tier sized rather than frontier-class.

9m read timeFrom thenewstack.io
Post cover image
Table of contents
What the model card showsSmaller, cheaper, built for enterprisesBuilt on an existing modelWhy hybrids don’t go far enoughBeyond sparse attentionThe near-term plan
473 Impressions1 Comment