Welcome to the Era of Experience
Silver and Sutton argue the 'era of human data' (pretraining + RLHF) is hitting diminishing returns because high-quality human data in math, code and science is nearly exhausted. They propose agents that live in continuous streams of experience, take grounded actions, learn from environment-grounded (not human-preference) rewards, and continually update world models and plans over lifetimes rather than episodes.
Why it matters here: This is essentially the manifesto for the whole program: continual learning over a lifelong stream, agent-chosen actions (queries, code execution, what to study), and grounded reward signals. It also supplies the motivating claim (human data exhaustion) and warns that reward must be grounded in environment signals rather than human judgment to escape the human-performance ceiling.
Ranker note: The manifesto for the entire program (lifelong stream, agent-chosen actions, grounded reward); read first to frame every design choice.
Reception: Very high-profile 2025 position paper by two of RL's most senior figures; widely discussed, with published commentaries and rebuttals; already a standard framing citation for experience-driven-AI work.