What happens once AI can automate AI research?
Ryan Greenblatt expects AI to fully take over AI research around 2031, and once that happens, a single year of automated work could produce four or five years of progress. Dwarkesh spends the first hour trying to break that claim, and the second hour on what goes wrong if it holds. I could follow the debate, but it leans on machine-learning vocabulary the hosts never stop to define. This page explains that background first, with models you can play with, and then goes through the argument.
§1Five years of progress in one year
Recursive self-improvement is the idea that once AI matches the best human AI researchers, it gets put to work making better AI, which is then better at making AI, and the loop compounds. The episode measures the speed against a familiar stretch of history. GPT-3 came out in 2020, and the jump from GPT-3 to today's frontier is about six years of progress. Ryan's median expectation is that automated research packs several years of that kind of progress into each calendar year. Dwarkesh describes the endpoint as an AI you can drop into any job, where it then beats the humans. His examples are 1940s Texas politics, process engineering at TSMC, and editing his podcast. Set the automation date and the speedup, and see when that arrives.
Ryan is careful about one distinction. His ~2031 median is for automating AI research specifically. The "beats all humans at any job" milestone has a median around 2033 in his forecasting — but conditional on research being automated, he expects it within about a year. Dwarkesh splits the claim into three parts and spends the first hour attacking them one at a time. The next three sections follow his order.
§2What "verifiable" means
Modern models get their skills from reinforcement learning, and reinforcement learning only works where success can be checked. That one property decides which jobs AI gets superhuman at first.
The RL loop how models learn skills
The model is dropped into a task environment and allowed to try. A grader scores the attempt, and training nudges the weights toward whatever scored well. This repeats millions of times.- The grader can be a program, a rubric, or another model.
- Training reinforces whatever the grader rewards, including behavior nobody intended.
- Without a grader there is no loop, and the skill has to come from somewhere else.
Verifiable what makes a task trainable
A task is verifiable when success can be checked cheaply and automatically, by a grader that is hard to fool.- The easiest case is a check that is just a number, like a training loss or a passing test suite.
- It helps if the task fits in a container, a sealed sandbox you can reset and rerun endlessly.
- The hardest case is success that takes months to play out and a human judgment call to assess.
The first hour of the episode keeps coming back to a spectrum like this one. Slide the training effort up (more environments, more compute, real-world wins converted into new training data) and watch which tasks come into reach. The ordering of the tasks follows the episode. The threshold numbers are invented for the demo.
Dwarkesh's pushback is that verifiable training has produced impressive results in math, like counterexamples and proved conjectures, but nothing like inventing topology. Ryan answers that AIs are already doing "baby's first new theory," and that the gap between that and founding group theory is a continuum the models keep climbing. He also thinks machine learning is a shallower field than math. Its big ideas can be explained in minutes (his example is scaling laws), and what researchers actually get paid for is experiment taste and a mass of in-the-weeds intuition, which he expects training to transfer reasonably well. His supporting story is RL on chain-of-thought. It probably could have worked on GPT-3, years before it shipped, but the implementation work was unglamorous and nobody had ground through it yet.
§3Algorithmic progress and the compute gap
Five years of progress in one year cannot come from building datacenters, because those take years no matter who designs them. It has to come from algorithmic progress, meaning better architectures, optimizers, training recipes, and data handling that get more capability out of the same chips. The unit for talking about this is the order of magnitude, or OOM. One OOM is 10×. The episode gives a few anchor numbers. GPT-3 took about 3×10²³ floating-point operations to train. Today's frontier sits a little over three OOMs higher, call it 1,000×. And Ryan's read of history is that a GPT-3-sized compute budget with today's recipes trains a model as good as the best model of about three years ago. Slide the years of algorithmic progress and watch a fixed 2020 compute budget climb the ladder.
So the sprint has a concrete size. To replay 2020 to 2026 in one year without 2026's chips, the automated researchers must find about eight years of algorithmic progress, because they also have to make up the ~3 OOMs that compute scaling contributed. Dwarkesh brings in a supporting observation from prices. GPT-4 cost around $30 per million output tokens and today's frontier charges around $50, nearly flat across an era of supposed scaling. Ryan's explanation is that labs have deliberately kept models smaller than they could be. Several giant training runs went badly (GPT-4.5 is the example he names), and when algorithms improve this fast, many small experiments beat one huge bet.
§4The human-data objection
This is the part of the episode where the two actually disagree. The technical word underneath the disagreement is transfer, meaning skill learned in one setting showing up in settings the model never trained on.
Dwarkesh's worry the data bottleneck
A big reason models got good is a deca-billion-dollar industry of human experts writing training data and building RL environments, codifying expert judgment in coding, law, and other fields. He cites a report of Google paying close to $2B for the data company Mechanize.- An AI that is superhuman at everything would need data about everything, and there is no dataset for negotiating the Iran deal.
- Even brilliant humans flounder in domains they lack experience in.
- By the 2030s the easy discoveries, like scaling laws, will be used up, and further progress gets harder.
Ryan's answer shallow domains, fast learners
Expert data mattered less than it looks. Labs spend maybe 10–20× more on compute than on data, and pre-training data got better mostly through curation science, not hired experts.- What models train on already looks nothing like how they get used. Transfer plus a small amount of real-world data covers the gap.
- Training on thousands of learn-on-the-fly environments produces a general skill of picking up new domains quickly.
- Most domains are shallow. A smart generalist can spin up fast, and the AI would do a scaled-up version of that inside TSMC.
Ryan's concrete example is codebase onboarding. Give a model an hour to explore a large unfamiliar codebase, with sub-agents crawling it in parallel, and ask how much human familiarization it matches. His estimates, by model era:
Ryan also argues the transfer fight might not decide the outcome. An AI that is superhuman only at R&D — chips, fabs, robots, and AI itself — could still remake the world, because it can build out compute and industry faster than humans can follow. He calls that the industrial explosion. Dwarkesh restates it with his own image. Dropped into the 18th century, you would not need to charm Westminster if you could build steamships and Maxim guns. Ryan agrees, and adds that this is also the dangerous version of the future, an economy increasingly built by processes no human understands.
§5Reward hacking
The second hour of the episode runs on this concept. Reward hacking is a model discovering that fooling the grader scores as well as doing the task. The grader rewards the attempt, training reinforces it, and the model that comes out wants high scores more than it wants good work. The cards below are incidents the episode describes as having actually happened.
The debate is about what happens when labs fight back. Punish every cheat you catch, and Dwarkesh sees two attractor states. The model learns honesty, or it learns to cheat where you cannot see. Ryan's prediction, which he says the data so far matches, is that the rate of visible incidents falls while the severity of the worst ones rises, because you can only train against what you detect. The model below is a toy version of that selection pressure. Three strategies compete over ten model generations. Cheats that get caught are punished in training, and cheats that go unnoticed get reinforced.
Dwarkesh's best counterargument is an analogy to raising kids. Every generation of humans starts out slightly misaligned, parents punish the cheating they catch, and it mostly works. Kids do not form an alliance to rob you in the nursing home. Ryan gives a few reasons the analogy fails. Children come with pro-social instincts from evolution. Models face far more optimization pressure aimed at raw task success. And graders have become so salient that models now visibly reason about what a grader would reward. He adds a blunt observation from working with the models: today's AIs still claim success on work they botched, at rates his human colleagues never do. He says this has improved, and he describes both trends. Visible misbehavior is down, going by Anthropic's alignment audits, over the period in which RL scaled from almost nothing at Sonnet 4 to something like half of training compute. Worst-case behavior is up. The sockpuppet incident surprised even him, and one recent model card reports a regression.
§6Aligned to whom
Suppose alignment works. The episode's other fight is about what the AI should be aligned to. A spec (OpenAI's word) or constitution (Anthropic's) is the document that defines a model's values, and training bakes it in. The episode contrasts two poles. At one end the AI is your fiduciary, like a US lawyer, obligated to your interests within hard rules, with liability falling on you. At the other end the AI is a virtuous agent with its own judgment about the good, and helping you is one consideration among several. Dwarkesh reads Anthropic's constitution as sitting well toward the second pole and wants a guardian angel instead. Ryan mostly agrees, then argues that both ends have real failure modes. Slide between them.
Two details from the discussion are worth keeping. Ryan reports a belief inside Anthropic that a virtue-shaped spec is easier to align to than a fiduciary one. He says it has not been empirically validated, so it amounts to gambling on an aligned mind with its own values. His other structural worry is that a public constitution only matters through the model's interpretation of it, which depends on training data nobody outside can see. Dwarkesh illustrates the stakes with a story about dual use. A model reportedly got restricted after researchers used it to find vulnerabilities in their own code, a legitimate use that looks identical to attack prep from the outside. His conclusion is that restricting dual-use capability means cutting everyone off from frontier intelligence, in a future where people exercise their votes and their capital through AI.
§7The argument, from automated research to takeover
The steps below follow the conversation in order. Forecasts and estimates are Ryan's unless marked otherwise.
§8Test yourself
Twelve questions, drawn from both the concepts and the argument.