Frontier labs are making a massive bet on scaling supervision. One part is paying experts (lawyers, poets, doctors, etc.) billions of dollars to generate and evaluate high-quality training data. Another part is verifiable supervision, where success can be checked by an external signal (e.g., whether code passes its tests, a proof checks out, or an agent completes a reproducible workflow in a simulated environment). Generally, the bet is that with enough compute, human judgment data, and verifiable rewards, models will reach superintelligence escape velocity. My guess today is that the models hit a ceiling before they generalize their way out of some of the limitations inherent in that recipe.
Why Both Halves of the Recipe Hit the Same Wall
Start with the human side. Whether experts are writing demonstrations, ranking outputs, or correcting mistakes, that supervision can only encode the knowledge, calibration, and blind spots of the annotator pool. The resulting models tend to mode-collapse toward consensus, damping down the rare, but correct reasoning, and there's no way to certify answers to questions humans can't already adjudicate.
RLVR avoids some of that because the reward can come from the world itself—the proof verifies, the code passes, the task completes. But it has its own constraint: it works best where the task is not only verifiable but cheap to replay, reset, and run at enormous scale. A great deal of consequential judgment satisfies neither condition. In investing, history, science, and business-building, the feedback is slow, ambiguous, or never arrives, which is why the gap hurts most in contested domains like macroeconomics, grand strategy, and historical causation.
2008 Financial Crisis and the Mortgage Market
In 2005, if you'd asked most experts what the odds were that the major banks would melt down because of their exposure to the U.S. housing market, they would have said next to impossible. Ask Michael Burry and you'd have gotten a very different answer. He went through the loan-level detail on subprime mortgage pools, saw what would happen when the teaser rates reset, and started buying credit default swaps against the bonds, a bet almost nobody else wanted to make (Lewis, 2010).
An LLM would find one part of this easy and one part hard. Give it the data and tell it to look, and I think an AI gets to Burry's conclusion pretty quickly. The hard part is deciding that this particular corner of the mortgage market was the thing to obsess over in the first place while every authority was signaling things were fine.
2008, however, is the easy case because the bet resolved. For questions like what really caused the Great Depression, the feedback never arrives at all; ninety years on, monetarists, Keynesians, and gold-standard scholars still weigh the mechanisms differently (Bernanke & James, 1991). There's no verifier to train against. Maybe that matters because different explanations imply different policy lessons.
The Counterexample: Overcoming a Mountain of Wrong Information with Reasoning
Here's a counter to my argument. There's a number in optics called the Abbe number. It's a function of a material's chromatic dispersion, the colored fringing you notice on high-contrast edges when you pixel-peep a photo. Go searching for advice on choosing the clearest sunglass lenses and you'll find a strong, confident consensus that a high Abbe number is critical for clarity. Marketing leans on it hard.
That consensus is mostly wrong. Past a fairly low bar, the lens material isn't what's limiting clarity. People can't tell the difference. The human eye has its own substantial chromatic aberration and the visual system largely works around it (Campbell & Gubisch, 1967; Thibos et al., 1990).
Ask a current model at low reasoning effort how important Abbe number is for sunglass clarity on a scale of one to ten, and in my own little testing it says eight. Turn the reasoning up, let it work through the actual optics, and it comes back with a two. There's a mountain of shallow, repeated, wrong content online, and more inference-time reasoning lets the model climb over its own bad prior and land where the 1980s and 1990s literature already was. This is one small example, but overturning that bad consensus is what I said the models would struggle with. That said, optics has fixed physics and a settled literature to reason toward, which the 2008 mortgage market and an open historical question do not.
So Where Does That Leave It
None of this proves the labs are wrong. Longer context windows, better tools, real-world deployment that generates fresh feedback, self-distillation, some genuinely new form of continual learning—any of those could bridge the gap I'm describing. Maybe the real bet is different from how I've framed it. Maybe scaling current methods is what lets researchers discover architectures that would otherwise have taken decades. So maybe the current recipe doesn't get us to escape velocity, but it builds the engine that does.
My larger point is that there's a real chance AI progress slows or even plateaus for a while. However, that doesn’t mean a plateau in impact. Even frozen at today's level, these systems are powerful enough that their diffusion through the economy will drive a massive restructuring of work.
References
Bernanke, B. S., & James, H. (1991). The gold standard, deflation, and financial crisis in the Great Depression: An international comparison. In R. G. Hubbard (Ed.), Financial markets and financial crises (pp. 33–68). University of Chicago Press.
Campbell, F. W., & Gubisch, R. W. (1967). The effect of chromatic aberration on visual acuity. The Journal of Physiology, 192(2), 345–358.
Lewis, M. (2010). The big short: Inside the doomsday machine. W. W. Norton.
Thibos, L. N., Bradley, A., Still, D. L., Zhang, X., & Howarth, P. A. (1990). Theory and measurement of ocular chromatic aberration. Vision Research, 30(1), 33–49.