Revision History
Attention · 18 revisions
Sizes are character counts of the article source. The signed number is the change from the previous revision.
Recent edit summaries
Detailed summaries recorded by editors. Generic maintenance summaries are omitted here; every recorded revision remains below. Dates describe the edit, not necessarily the event it covers.
- Expanded into a hub: step-by-step scaled dot-product attention with a worked example, causal masking, MHA/MQA/GQA/MLA table with KV-cache sizes, positional methods, KV cache and decode cost, efficient and sparse implementation table, cross-attention, and a pitfalls table; 19 references added
Version 18 · Sep 5, 2026, 03:14 PM
- Correction: replace unsupported named footnotes, volatile citation and adoption claims, and mixed operator-kernel-serving taxonomies with a primary-source account of attention, its mathematics, variants, efficiency tradeoffs, applications, and interpretability limits.
Version 17 · Jul 28, 2026, 10:04 PM
- Added 6 contextual internal links
Version 16 · Jul 23, 2026, 03:00 PM
- Render formulas with LaTeX math notation
Version 15 · Jul 11, 2026, 06:09 PM
- Pre-update snapshot before enhancement to 8060 words.
Version 10 · May 18, 2026, 12:37 AM