Revision History
Sparse attention · 9 revisions
Sizes are character counts of the article source. The signed number is the change from the previous revision.
Recent edit summaries
Detailed summaries recorded by editors. Generic maintenance summaries are omitted here; every recorded revision remains below. Dates describe the edit, not necessarily the event it covers.
- Correction: Mixtral 8x7B uses fully dense 32K attention, not sliding-window (its technical report and config); Claude 3.5 Sonnet launched with 200K context, not 128K, and 200K was Claude 3's standard window. Minor: Sparse Transformer fixed pattern attends to the last c positions (not one element); Longformer window schedule per the paper; Atri Rudra is at Buffalo; BigBird path-length wording; NSA authors are Yuan et al.; DSA introduced with V3.2-Exp (Sept 2025); DeepSpeed row dated and cited; Block-Sparse-Attention row per README; 'among the first'; 100,000 squared equals 10 billion; four red links unlinked
Version 9 · Sep 16, 2026, 03:12 PM
- Correction: fix reference 8 (arXiv:2512.07011 is Ohayon et al. 2025, not a 2024 Han-lab paper); add selection-overhead and offloading subsections (DSA, IndexCache, HISA, SparDA, NOSA), table rows, refs 11-16
Version 8 · Sep 16, 2026, 02:52 PM
- Added 6 contextual internal links
Version 7 · Jul 23, 2026, 03:09 PM
- Render formulas with LaTeX math notation
Version 6 · Jul 11, 2026, 07:13 PM