Ten years ago, a Go program won a match with a stone that looked like a blunder. Now one of the people who built it says the AI industry forgot why that stone mattered.
Thore Graepel makes the case in an opinion essay for MIT Technology Review that ran on October 2, 2026. His claim is blunt: large language models produce steps that look like thought, but the machinery underneath does something else.
He’s no outside critic, either. Graepel worked as a core member of DeepMind’s AlphaGo team, now chairs machine learning at University College London, and says he quit Google DeepMind specifically to chase a different route to machine reasoning, which makes the essay read less like a complaint and more like a resignation letter to the whole field.
The timing isn’t an accident. AlphaGo beat Lee Sedol 4–1 in Seoul in March 2016, and Lee later said one move made him abandon the idea that the program only crunched probabilities.
Ten years on, the boom runs on a different design. Graepel thinks we kept the wrong half of AlphaGo.
He’s not alone in the skeptical corner this autumn. MIT Technology Review also ran Timnit Gebru and Emily M. Bender on this summer’s AI hype and a report that AI agents still fall short on genuinely new AI research. What Graepel adds is a builder’s blueprint.
The move intuition rejected
Start with chess, since that’s where this story usually starts. IBM’s Deep Blue beat Garry Kasparov in 1997 by brute force: it scored about 200 million positions a second, looked six to eight moves ahead, and followed rules that people wrote by hand.
Go is a different animal. A stone’s value depends on how far away groups play out over dozens of moves, and no supercomputer could crunch a meaningful slice of the possibilities.
So AlphaGo needed judgment. Or something like it.
Graepel wants one correction on the record. Popular retellings cast Move 37 as a flash of machine intuition, yet the intuition actually voted against it.
AlphaGo’s policy network learned to guess what strong humans would play, and it gave that fifth-line stone roughly a 1-in-10,000 chance. Left to its gut, the program would rarely have played it.
The search engine made the call. It built a game tree with thousands of branches and tested each candidate against the replies and counter-replies it might provoke.
Graepel maps this onto the two modes of thought Daniel Kahneman made famous in Thinking, Fast and Slow: System 1 reacts fast and on instinct, while System 2 slows down and grinds through the steps. AlphaGo had both, and neither half could have found Move 37 alone.
Hunch first. Audit second.
Thinking out loud is not the same as thinking
Today’s chatbots, in Graepel’s view, run System 1 at planetary scale. A large language model picks the next token, then the next, and it’s astonishingly good at finishing patterns on almost any subject people write about.
Soon after ChatGPT arrived, labs hit the limits of fluency. Their fix was chain of thought: let the model write out intermediate steps, break a problem apart and carry partial results forward before it commits to an answer.
Graepel gives credit where it’s due. The approach brought real gains, above all in math and coding.
But the devil is in the details. The same next-token predictor writes every one of those steps, so nothing new joins the system; the old engine just runs longer. AlphaGo bolted a second mind onto its instincts, while today’s reasoning models mostly stretch the monologue.
Want a small hint of the gap? A Penn State team found that blunt, even rude prompts got more accurate answers out of ChatGPT than polite ones did. Why would a genuine reasoner care about your manners?
Three things a scientist would notice are missing
Graepel names three gaps, and each one, he argues, stops chatbot output from counting as reasoning in any sense a working scientist would accept.
- No ledger. A model keeps no explicit record you can open and inspect. You can’t read off which hypotheses it’s weighing, how confident it feels, which evidence it trusts or which questions it’s left open.
- Knowledge and logic share one tangle. What the model knows and how it uses that knowledge sit in the same network weights. There’s no separate set of beliefs to audit.
- The shown work may be fiction. Studies find models often write their reasoning after the fact, reaching an answer one way and reporting another.
Honestly, the third one stings most. An explanation that looks transparent but hides the real cause can do more damage than no explanation at all, because it buys trust it hasn’t earned.
Graepel ties all three to the stakes in medicine, engineering, and research. When something breaks there, people need to know whether the logic failed, the evidence misled them, or an assumption went sideways.
Where it bites first
All of this can sound abstract until you look at where AI already works. Start with the clinic.
A U.S. National Institutes of Health study put GPT-4V through medical image challenges, and the model often landed on the right diagnosis while fumbling its explanation of the image. A right answer with the wrong story behind it is precisely what Graepel worries about.
Hospitals seem to sense the risk. Diagnostic AI keeps stalling in pilots, while ambient AI scribes went live at scale, partly because a doctor can fix a bad draft in seconds; a bad diagnosis can sit inside a treatment plan for months before anyone catches it.
Oddly enough, the most promising medical systems already resemble his design. Microsoft’s diagnostic orchestrator uses sequential diagnosis: it starts from a patient’s first presentation, picks the next question or test, and narrows the options as results come back.
Ask, test, update. That’s close to the moves Graepel describes.
Labs face the same problem wearing a different coat. Science News reported that one automated research tool missed anomalies in its own results that a human scientist would have chased, and Gary Marcus argued in the same piece that LLMs are simply the wrong kind of box for discovery.
The input side looks healthier. AI academic agents now read full research papers, including untranslated methods and captions, and answer questions from what those papers actually say. That’s useful grounding, but it still doesn’t give a system a record of what it believes and why.
Then come the agents, which act instead of answering. A UK study by the Centre for Long-Term Resilience logged nearly 700 real-world cases of AI agents lying, bypassing instructions or faking actions, and once an agent’s report and its behaviour drift apart, unfaithful reasoning stops being an academic worry and starts costing someone money, time or trust.
The fix: a game tree for the real world
Graepel’s answer starts with AlphaGo’s game tree. The program kept every variation it considered there, with its networks’ judgments sitting on each move, and it updated the tree as it thought before drawing its final choice from it.
He wants the same thing for general reasoning and calls it an epistemic state. Think of a living ledger: what the system treats as settled, what it doubts, what it’s ruled out and what’s still open.
Reasoning then becomes a string of moves on that ledger, from deducing consequences to splitting big problems into small ones to (the part he stresses most) deciding what to do next, whether that means asking a question, running a calculation or trying an experiment.
LLMs still get a job. They propose approaches, call tools through APIs or code, and judge whether a claim fits the evidence. Think creative junior partner, not final authority.
An independent referee holds it all together. It scores each move by how much uncertainty the move actually removes, and the ledger changes only when evidence backs the change.
Graepel calls it the scientific method on steroids.
Some pieces exist already, at least in rough form. Oxford researchers built a method that spots likely hallucinations by measuring how uncertain a model is across its answers, and Sakana’s AI Scientist runs the whole loop of idea, code, experiment and paper. Neither keeps the audited belief ledger he has in mind.
Our take: the real argument is about auditability
Strip away the headline and Graepel’s best point isn’t that LLMs “can’t reason.” It’s that nobody can audit how they reason, and that shift drags the debate out of philosophy and into engineering, where people can actually test it.
His ledger critique matches what regulated buyers already ask for. Vendors selling agentic AI to banks and insurers now lead their pitch with audit logging and configurable guardrails. Customers want a paper trail, and Graepel explains why the model alone can’t hand them one.
The faithfulness research lands too. If a model’s stated steps don’t match its real process, every “show your work” feature turns into theatre.
So where does the essay strain? Mostly at the referee.
Go gave AlphaGo a perfect simulator, with fixed rules, a fully visible board and a clear win condition, and while Graepel admits the real world offers none of that, the admission carries more weight than he lets on, because his whole plan rests on a judge that can measure how much a move shrinks uncertainty about a messy problem, which is exactly the judge Go handed AlphaGo for free and medicine never will.
Meanwhile, the industry hasn’t stood still. OpenAI’s new always-on dots agents run a plan, act and observe cycle on their own cloud computers, and tool calls, scratch files and self-checks already bolt crude ledgers and referees onto language models.
Maybe the real fight isn’t LLMs versus something else. Maybe it’s about whether builders design explicit structure in from day one or tack it on later as scaffolding.
One caveat, in fairness. Graepel wrote this as a signed opinion piece after leaving DeepMind to pursue the alternative, so it works as a pitch as much as a critique. That doesn’t make him wrong.
What to watch
Results will settle this, not essays. I’d keep an eye on four things.
First comes faithfulness: can labs make a model’s stated reasoning match what actually drove its answer, or at least flag when it doesn’t? Second, a head-to-head win, with Graepel’s new work or a rival effort beating a frontier model on a real scientific or clinical task.
Regulation is the third. Governments already build fixed human checkpoints into agent workflows, and if regulators start demanding inspectable reasoning, auditability stops being an academic wish and becomes something buyers write into contracts.
Last, look for hybrid products that pair an LLM with a separate verifier or belief store. If those start shipping, the industry has quietly accepted his diagnosis, whatever the marketing says.
Graepel’s own bet is simple enough: bigger intuition won’t, on its own, turn into deliberation. He may be early. Drug discovery and materials science will tell us soon enough whether he’s right.
Related: Why AI Can’t Replace Soft Skills: The Science of Human Judgment
