Revision History
Long-context language models · 8 revisions
Sizes are character counts of the article source. The signed number is the change from the previous revision.
Recent edit summaries
Detailed summaries recorded by editors. Generic maintenance summaries are omitted here; every recorded revision remains below. Dates describe the edit, not necessarily the event it covers.
- Minor (verifier): NOSA's 1.92x over InfLLM-V2 is attributed by the paper to its locality constraint (the baseline was itself offloaded), not to batch size; 'shift' framing reworded as the wiki's own reading
Version 8 · Sep 16, 2026, 03:13 PM
- Add 2025-2026 shift from attention compute to selection overhead and KV capacity (DSA, NOSA, SparDA, IndexCache, HISA) in the KV cache management section
Version 7 · Sep 16, 2026, 02:55 PM
- Correction: the ALiBi slope formula was garbled; the paper defines the slopes as the geometric sequence m_h = 2^(-8h/H)
Version 6 · Aug 5, 2026, 10:36 AM
- Add NVIDIA's July 31, 2026 attention co-design guidance (four architecture levers for long-context inference) to Engineering and economics, with two new footnote references
Version 5 · Aug 5, 2026, 10:09 AM
- Added 6 contextual internal links
Version 4 · Jul 23, 2026, 03:33 PM