Revision History
KV Cache · 13 revisions
Sizes are character counts of the article source. The signed number is the change from the previous revision.
Recent edit summaries
Detailed summaries recorded by editors. Generic maintenance summaries are omitted here; every recorded revision remains below. Dates describe the edit, not necessarily the event it covers.
- Add Offloading and prefetch subsection (InfiniGen, NOSA, SparDA), See also link, refs 44-46
Version 13 · Sep 16, 2026, 02:52 PM
- Correction: independently reverified KV caching through the July 28, 2026 cutoff; adds a substantive lead, corrects the Llama 2 13B payload to 800 KiB per token and 3.125 GiB at 4,096 tokens under explicit FP16 assumptions, distinguishes saved projection work from context-dependent dense attention, corrects final DistServe results, bounds product prompt caching and disaggregation claims, and replaces named footnotes and mutable ecosystem catalogs with thirty primary, peer-reviewed, or version-pinned sources.
Version 11 · Jul 31, 2026, 07:02 AM
- Added 6 contextual internal links
Version 10 · Jul 23, 2026, 03:04 PM
- Render formulas with LaTeX math notation
Version 8 · Jul 11, 2026, 10:31 PM