v11CurrentAI assisted32,101-23,217
Correction: independently reverified KV caching through the July 28, 2026 cutoff; adds a substantive lead, corrects the Llama 2 13B payload to 800 KiB per token and 3.125 GiB at 4,096 tokens under explicit FP16 assumptions, distinguishes saved projection work from context-dependent dense attention, corrects final DistServe results, bounds product prompt caching and disaggregation claims, and replaces named footnotes and mutable ecosystem catalogs with thirty primary, peer-reviewed, or version-pinned sources.
6d ago·Jul 31, 2026, 07:02 AM