Revision History
DeepSeek Sparse Attention (DSA) · 5 revisions
Sizes are character counts of the article source. The signed number is the change from the previous revision.
Recent edit summaries
Detailed summaries recorded by editors. Generic maintenance summaries are omitted here; every recorded revision remains below. Dates describe the edit, not necessarily the event it covers.
- Correction: GPQA-Diamond, HLE and HMMT 2025 are not the three largest V3.2-Exp drops (Aider-Polyglot fell 1.6 points, twice GPQA's); the report explains only those three. Minor: release-note quote restored verbatim ('DSA'); 'for the first time' quote sourced to the release README; selection-overhead and 'shift' wording attributed; quality claim attributed to DeepSeek's evaluations; dead ReLU link
Version 5 · Sep 16, 2026, 03:13 PM
- Correction: indexer inputs are hidden-state projections (not MLA latents), benchmark drops attributed to fewer reasoning tokens per the tech report, training details cited to the report; add Related and derivative work (IndexCache, HISA, SparDA, NOSA)
Version 4 · Sep 16, 2026, 02:55 PM
- Added 6 contextual internal links
Version 3 · Jul 23, 2026, 04:02 PM