Why SubQ-1.1-Small Proves We’ve Been Building AI “Backwards”

This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).

SubQ-1.1-Small introduces Subquadratic Sparse Attention (SSA), a content-dependent routing mechanism that reduces attention FLOPs by 64.5x at 1 million tokens compared to dense attention. Trained on 1M-token sequences, the model generalizes to 12M tokens with 98% recall on needle-in-a-haystack tasks while attending to only 0.13% of token pairs. It also scores 85.4% on GPQA Diamond and 89.7% on LiveCodeBench, preserving reasoning capabilities. The author argues this makes complex RAG pipelines, chunking strategies, and vector databases obsolete — developers could simply drop entire codebases into a prompt instead of fragmenting and embedding them.

3m read timeFrom medium.com
Post cover image
Table of contents
The Secret to Efficiency: SSAThe Milestone: Extreme Context GeneralizationGet Pablo jusue’s stories in your inboxBeyond Memory: Real IntelligenceConclusion: The End of “Chunking”
227 Impressions