There is a moment in a career when competence becomes suspiciously comfortable. Tom Brown reached it in 2015. He had graduated from MIT, spent eight years building web startups, won Y Combinator backing and grown a company to thirty people. By Silicon Valley arithmetic, he had put in his ten thousand hours. Then he looked toward artificial intelligence and discovered that his accumulated fluency bought him very little. He would have to begin again.
The prospect was not flattering. Brown remembered a B-minus in linear algebra, though he has since joked it may have been a C-plus. He had no doctorate and no neat research pedigree. His friends did not disguise their doubts. AI safety sounded remote to them, and some were not convinced that Brown would be good at the work anyway. Their candor was bracing. It was also useful.
He gave himself three months. At South Park Commons in San Francisco, he took an online machine-learning course, tried Kaggle projects, worked through Linear Algebra Done Right and joined a paper-reading group. The place was full of people changing direction. Expertise was welcome; attachment to looking expert was not. Brown later wrote that it let him put aside his “armor of expertise.” Six months of focused work later, he had a job at OpenAI.
“It gave me a community where I felt comfortable putting aside my armor of expertise.”Tom Brown, on starting over at South Park Commons
Before the machines, there were drinks
Brown's route to that study table began at twenty-one. In the summer of 2009, two friends started Linked Language and he became the first employee. The small company supplied a lesson more durable than any particular line of code: nobody was coming to assign the next important task. Brown has described the difference with a slightly unruly animal metaphor. School had trained him to wait like a dog for a bowl. A startup required a wolf that could hunt.
He did not romanticize his abilities. At MoPub, where he was an early engineer on the mobile advertising company's core server, he recalls struggling as a programmer. The scale forced improvement. Then came a Winter 2012 Y Combinator company, first called Solid Stage, conceived before Docker made container tooling legible. The idea failed its own interview rather memorably: a frowning face appeared on the board beside the question, “What are you actually going to build?”
The answer became Grouper, a social service that arranged a meeting between two groups of three, often over drinks. Its origin was personal rather than algorithmically majestic. Brown had been, in his telling, an extraordinarily awkward kid. Going out with friends made meeting people feel safer. Grouper turned that social crutch into a product and produced what Brown cheerfully calls reliable shenanigans.
One enthusiastic Grouper customer was Greg Brockman, then at Stripe. That friendship would become Brown's bridge to OpenAI. When OpenAI was announced, Brown messaged Brockman and asked whether a web engineer with an undistinguished linear-algebra grade might help. Brockman saw a scarcity of people who understood both machine learning and distributed systems. Brown kept studying and checking in. Eventually, there was work for him.
The line that kept going
At OpenAI, Brown worked for a year, spent another at Google Brain, then returned. From 2018 into 2019, the work gathered around a blunt proposition: with the right recipe, larger models trained with more computing power became more capable in a remarkably predictable way. Brown co-authored the neural scaling-laws paper and turned his attention toward scaling. The graphs ran straight across ranges so vast that he found them almost comic - twelve orders of magnitude, an absurd distance for an empirical relationship to behave itself.
The theory acquired a name and then a famous demonstration. Brown was lead author of the 2020 GPT-3 paper, which described a language model with 175 billion parameters and an unusual ability to perform tasks from instructions and a few examples without task-specific retraining. The paper became one of the hinge documents of the present AI era. Its opening name belonged to an engineer who had wondered, five years earlier, whether he was qualified to enter the field.
Brown's contribution was not merely literary order on a long author list. The work depended on reliable distributed systems and on moving the stack toward GPUs with software that allowed researchers to iterate quickly. His old life had returned in disguise. Serving mobile advertisements and keeping startup infrastructure alive were hardly glamorous rehearsals for a frontier model, but rehearsals they were.
The career pivot looked like discarded experience. GPT-3 revealed it as stored energy.
Seven people, no certain product
The group around OpenAI's safety and scaling work had become close. They communicated in public Slack channels, took the scaling evidence seriously and worried about whether increasingly capable systems would remain aligned with human intentions. In late 2020, members of that group left. In 2021, Brown became one of seven co-founders of Anthropic.
He is notably unsentimental about the founding tableau. OpenAI had money, recognition and a formidable roster. Anthropic had seven co-founders working remotely, a mission and no certainty that it would produce a product. The first recruits could have chosen better-known companies and, in many cases, better-understood jobs. Brown credits that early concentration of mission-driven people with preserving a culture willing to object when the institution drifts.
Claude existed as a Slack bot months before ChatGPT appeared, but Anthropic hesitated. The team had not settled whether releasing it was wise and had not built the serving infrastructure a rush of users would require. After ChatGPT's arrival, Anthropic launched its API and then Claude.ai. Brown has called the delay a lesson. Conviction about consequences is necessary; so is preparing the machinery for whichever decision you make.
The company's product footing became clearer with Claude 3.5 Sonnet and its aptitude for programming. Claude Code began as an internal tool assembled to help Anthropic's own engineers. Brown offers an endearingly strange account of its advantage: the team treated Claude itself as a user. Give the model good tools. Give it the right context. Think about what makes its work easier. It is product empathy extended to a stakeholder made of matrices.
The physical consequence of a graph
Brown is now Anthropic's Chief Compute Officer. The official sentence is tidy: he leads the technical organization responsible for securing, scaling and effectively using its computing resources. The reality includes clouds, accelerators, networking, performance software, electricity and construction. A line on a scaling-law chart eventually becomes a building with a formidable power bill.
Anthropic uses Nvidia GPUs, Google's TPUs and Amazon's Trainium. Supporting three platforms divides performance-engineering effort and multiplies the software work. Brown accepts the nuisance for two reasons: the company can reach a larger combined pool of scarce capacity, and it can match the right chips to training or inference jobs. Optionality is expensive. Dependence can be dearer.
The strategy is visible in Project Rainier, the vast AWS cluster built for Anthropic on hundreds of thousands of Trainium 2 chips. Brown manages Anthropic's relationship with Amazon and compares the partnership to Netflix's early role as a demanding AWS customer. A frontier lab hacks through the jungle first; the cloud provider improves the trail; later customers get a path. His metaphors still prefer wolves and machetes to management diagrams.
Compute has also carried Brown beyond software. He expects power to become the central bottleneck in the United States and favors more generation of every kind, including nuclear and renewables, alongside easier data-center construction. He has described the current expansion as humanity's largest infrastructure buildout. The phrase sounds extravagant until one remembers that his day job is to turn extravagant demand into an operating system.
An idealized version, minus the armor
Brown's public story has a useful lack of varnish. He says he was bad at programming. He mentions the bad grade. He admits Anthropic did not know what product it would make and was surprised by Claude's success in coding. This is not false modesty so much as an engineer's running error log. A system improves when its failures remain observable.
Asked what he would tell a younger version of himself, Brown did not prescribe a credential or a specialty. Take more risks, he said. Work on something that would excite your friends, or make an idealized version of yourself proud if you succeeded. The advice is intrinsically motivated and pleasantly inconvenient. It offers no guaranteed promotion and no respectable timetable.
His own timetable took him from arranging drinks among strangers to arranging chips across continents. The apparent leap hides the continuity: build systems that allow unlikely interactions, learn where they break and refuse to wait for somebody else to place the next task in the bowl. The armor came off. The work got larger.