YesPress profilePhysics → scaling laws → AnthropicChief Science OfficerNew equations requiredYesPress profilePhysics → scaling laws → AnthropicChief Science OfficerNew equations required

People / Artificial Intelligence / Theoretical Physics

Jared Kaplan Found a Law for Intelligence. Then Came the Fine Print.

A physicist went looking for equations that could keep researchers honest. He found curves that helped remake artificial intelligence - and inherited the harder problem of deciding what their ascent permits.

There is a particular sort of optimism in demanding that every paper contain a new equation. The rule belonged to Jared Kaplan during his years as a theoretical physicist. It was less a boast than a private customs checkpoint: no fresh mathematics, no entry. Clever prose could escort an old idea around the block, but it could not make the idea new. Kaplan wanted the kind of claim that might be wrong in public.

That standard has followed him into artificial intelligence, where being wrong in public is considerably easier and the equations have acquired budgets. Kaplan is now a co-founder and chief science officer of Anthropic, the maker of Claude. He heads technical research and serves as the company's responsible scaling officer, a title that sounds faintly as though someone has been asked to supervise a graph. In Kaplan's case, the description is unusually literal.

His name leads the 2020 paper Scaling Laws for Neural Language Models. Its central result was spare: as researchers increased a language model's size, training data and compute, its performance improved in regular power-law relationships. The paper tracked some trends across more than seven orders of magnitude. A field accustomed to temperamental experiments had been handed a ruler.

7+orders of magnitude tracked in scaling trends
2020scaling-laws and GPT-3 papers published
2021Anthropic co-founded

A tent between worlds

Kaplan did not begin with talking machines. A Chicago native who attended the Illinois Mathematics and Science Academy, he was on the United States Physics Team in 2001. He studied physics and mathematics at Stanford, then wrote a Harvard doctorate called Aspects of Holography under Nima Arkani-Hamed. After research posts at SLAC and Stanford, he joined Johns Hopkins in 2012.

In graduate school, he initially wanted practical work connected to the Large Hadron Collider. The attraction drifted toward stranger territory: quantum gravity, conformal field theory and the holographic principle, the proposal that a world with gravity may be described by information living in fewer dimensions. Kaplan once said he had pitched a tent between string theory and phenomenology. It is a good image for his career - temporary shelter on the border between grand ideas and things one can calculate.

“I prefer to work on problems that are precise enough that it's clear whether we've made progress, but vague enough that we currently lack both the tools and the vision to anticipate the solution.”Jared Kaplan

His colleagues noticed the combination. A former graduate student, Nikhil Anand, described Kaplan as someone who could extract the governing physical principle from a complicated problem and push it as far as it would go. Kaplan put the preference more mischievously: he wanted to be confused on some basic level, but confused precisely enough to persuade another scientist that progress had occurred.

There was another laboratory, padded and less forgiving. Kaplan practiced Brazilian jiu-jitsu with the Johns Hopkins club and had earned a blue belt by 2017. Earlier he had taken up Muay Thai. Jiu-jitsu worked because it occupied him completely; for a couple of hours, research disappeared. The metaphor is tempting - leverage, position, patient pressure - but the more useful fact is simpler. Even a mind devoted to the universe occasionally needed someone to try to put it in a chokehold.

Jared Kaplan speaking onstage at a technology event
Kaplan onstage at the Bloomberg Technology Summit in London. The physicist's second arena has fewer mats and rather more microphones. Photo: Chris J. Ratcliffe/Bloomberg via Getty Images.

The curve that made a wager calculable

Kaplan joined OpenAI in 2019. The move from quantum gravity to machine learning looks abrupt only if one ignores his taste in questions. Neural networks were complicated systems producing empirical regularities that theory did not yet explain. Researchers could train a larger model and watch it improve, but how much improvement should they expect? Which scarce ingredient should receive the next dollar?

The scaling-laws work, written with Sam McCandlish, Tom Henighan and other colleagues, gave quantitative answers. Loss declined as a power law with model size, dataset size and compute. Within the observed ranges, the trends were smooth and remarkably indifferent to architectural details. Bigger had already been a hunch. The paper turned it into a forecast and offered guidance on how to allocate a training budget.

The same year, Kaplan co-authored the paper introducing GPT-3. The model's ability to perform tasks from instructions and examples, without conventional task-specific retraining, made the scaling argument visible to people who would never inspect a loss curve. He also contributed to Codex. In two years, the physicist who had studied whether spacetime was emergent helped demonstrate that linguistic competence could appear from a sufficiently enlarged prediction machine.

The curve was influential because it removed one kind of uncertainty. It did not remove all of them. Later work would refine the compute-optimal balance between model size and training data. Capabilities that matter to users can appear unevenly even when aggregate loss moves smoothly. And prediction does not answer the older political question hiding beneath every technical forecast: who gets to act on it?

When the graph acquires a gate

Kaplan co-founded Anthropic in 2021 with six former OpenAI colleagues, including Dario and Daniela Amodei, Jack Clark, Tom Brown, Sam McCandlish and Chris Olah. The company would build frontier models while studying how to make them reliable, interpretable and steerable. Kaplan's research instincts suited the bargain: work on the system itself, measure its changing behavior and make safety an empirical discipline rather than an ornamental paragraph.

Anthropic's Responsible Scaling Policy attempted to connect capability thresholds to stronger safeguards. In 2024, Kaplan became responsible scaling officer. In 2025, internal evaluations led him to place Claude 4 under AI Safety Level 3 protections, which included stronger security and resistance to attempts to bypass safeguards. TIME put him on its TIME100 AI list that year and quoted his hope that such commitments could create a “race to the top.”

1 · Measure

Evaluate models for capabilities that could create severe risks.

2 · Trigger

Use capability thresholds to activate stronger safeguards.

3 · Decide

Translate evidence into training, release and security choices.

By February 2026, the policy had changed. Anthropic removed its earlier categorical promise not to train systems beyond a threshold unless protections were already in place. The revised approach emphasized disclosure, comparison with competitors and the possibility of delay under narrower conditions. Kaplan argued that stopping unilaterally would not help if rivals continued and Anthropic lost its ability to understand the frontier. He rejected the description of a U-turn.

This is the fine print of the scaling law. A power law describes what may happen when inputs grow. It says nothing about market competition, international coordination or how much voluntary restraint survives contact with a rival's data center. The responsible scaling officer can watch the gauges; he does not control the weather, the traffic or the other vehicles.

The curve offered foresight. It did not grant command.

Potato planets and the human enterprise

Kaplan's public persona is quieter than the consequences attached to his work. In 2020, he made three short science films with a producer for Scientific American. One imagined what different strengths of gravity might do to planets. Weak gravity could leave them potato-shaped; stronger gravity might permit something closer to the tiny world in The Little Prince. Another explained cosmic structure by comparing matter in the universe to a herd of animals.

He wanted the films to be “fun and somewhat whimsical,” and wanted viewers to understand science as a human enterprise. The phrase deserves to travel with him. Scaling laws can make machine learning look like destiny plotted against compute. But research is conducted by people, inside organizations, amid incentives and incomplete information. The humanity does not vanish because the axes are logarithmic.

There is a useful modesty in the potato planet. Change one constant and the familiar world becomes comic, then strange, then perhaps impossible. The exercise does not predict another universe; it sharpens one's sense of what this universe depends upon. Kaplan's AI work proceeds through a related maneuver. Vary model size, data or compute, then observe which regularities survive. The difference is that alternative training runs are not imaginary worlds. Their descendants become products, colleagues, competitors and tools. A physicist can admire the regularity while the executive must answer the telephone.

▶Scaling and the Road to Human-Level AIWatch Kaplan's 2025 talk at Y Combinator's AI Startup School

Kaplan has said he would prefer artificial general intelligence to arrive in 2032 rather than 2028. Preference is not prediction, and neither is control. It is, however, a glimpse of the person standing beside the famous curve: excited by what intelligence might become, impatient with imprecision, and aware that additional time has value.

His career has expanded the old publication rule. A new equation remains evidence that an idea has earned attention. Now comes the harder test: whether the institutions built around that idea can produce rules equal to its consequences. Physics allowed Kaplan to ask what the world is. Artificial intelligence has added an impolite follow-up - what, exactly, are we going to do about it?