On March 13, 2019, Richard Sutton posted a short essay on his personal website and called it "The Bitter Lesson." Its claim was blunt. After seventy years of AI research, the biggest lesson was that "general methods that leverage computation are ultimately the most effective, and by a large margin." Building what human experts know into a machine felt productive, and it kept losing to systems that simply searched and learned with more computing power.
Six days later, roboticist Rodney Brooks answered with "A Better Lesson." "I think Sutton is wrong for a number of reasons," he wrote. His main point was that the systems Sutton praised were themselves built with human ingenuity, and that saying a solution style removes human ingenuity is "a terribly myopic view of the world."
On May 11, 1997, Garry Kasparov resigned the sixth game of his match with IBM's Deep Blue after 19 moves. During that match he said: "I'm a human being. When I see something that is well beyond my understanding, I'm afraid." In March 2016, in Seoul, Google DeepMind's AlphaGo played its 37th move of game two against Lee Sedol, and Lee left the room for about fifteen minutes. Fan Hui, the European champion, watching the move, said: "It's not a human move. I've never seen a human play this move." David Silver, who led the AlphaGo project, did his PhD under Sutton.
Look at what chess and Go share besides being hard. Every game ends with a winner. A system can try something, get an unambiguous score, and learn from it billions of times, and nobody has to argue about who was right. Search and learning scale with compute partly because the game itself works as the referee.
In September 2026, OpenAI announced a result on the Navier–Stokes equations, one of the Clay Institute's Millennium Prize Problems. Coverage of the announcement reports that about 10,000 AI agents worked for roughly 88 hours, followed by about 17 more hours turning the argument into Lean, a language in which a computer program checks every logical step. Lean cannot be charmed. A proof either compiles or it doesn't, which makes it the closest thing mathematics has to a game's final score.
The same coverage carries the caveats. The result has not been peer reviewed, mathematicians are still reading it, and parts of the claim are disputed. Two more limits are worth naming, because they apply well beyond mathematics. Lean confirms that a proof establishes the statement it was given; whether that statement is the one the field actually cares about is a question a person still has to check. And a guest post on Terence Tao's blog reminds readers of Yehuda Rav's line that "theorems are the headlines, proofs are the inside story": a verified answer is not the same as an understood one.
Most of the models companies actually ship do not come with a referee. There is a test set someone wrote, sometimes a language model acting as judge, and sometimes a panel of human raters. Each of these is a referee with flaws, and the flaws are statistical:
Sutton's lesson is about generation: more compute buys better search and better learning. It says nothing about how you know the result is good. That part does not scale with compute. Someone has to decide what counts as correct, how many examples are enough, and whether a difference is bigger than chance, and each of those is a human judgment that has to be defended to a customer, an auditor, or a funder.
That is the work we do at DASS: not building the model, but testing whether the referee holds up. If you have an eval score you plan to put in front of someone who will question it, our AI evaluation page describes how we check it, and the free model comparison tool is a quick first look at whether one model really beats another.
Compute makes learning compound. Evidence is what makes the result worth trusting.
One email whenever we publish something new. No spam, unsubscribe anytime.
{{ subscribeErrorMsg }}