Revision History
Reinforcement learning · 15 revisions
Sizes are character counts of the article source. The signed number is the change from the previous revision.
Recent edit summaries
Detailed summaries recorded by editors. Generic maintenance summaries are omitted here; every recorded revision remains below. Dates describe the edit, not necessarily the event it covers.
- Added 6 contextual internal links
Version 15 · Jul 23, 2026, 03:02 PM
- Enhancement: recency + accuracy pass (top-linked core page)
Version 14 · Jul 23, 2026, 10:33 AM
- Correction: Williams REINFORCE timeline redated to 1992; AlphaGo Zero 4.9M games was the 3-day run (40-day run: 29M); R1-Zero (not R1) skipped SFT; ACM Turing quote restored verbatim.
Version 13 · Jul 10, 2026, 12:26 AM
- Render formulas with LaTeX math notation
Version 12 · Jul 10, 2026, 12:10 AM