FIELD NOTES
Greg Kamradt ✦ Needle in a Haystack ✦ ARC Prize Foundation ✦ The test keeps getting harder

Profile / Artificial intelligence

Greg Kamradt and the Art of Asking AI a Better Question

Greg Kamradt turned a sandwich hidden in a mountain of text into a revealing AI test. Now, as president of ARC Prize Foundation, he is asking a harder question: how quickly can a machine learn something new?

Early in 2024, Mike Knoop had a proposal for Greg Kamradt. Knoop, a co-founder of Zapier, wanted to make a little-known AI benchmark impossible to ignore. Kamradt had first encountered him through a talk about how Zapier was bringing AI into its own operations, then invited him into a recorded conversation. They kept in touch. Now Knoop was starting a competition with François Chollet, the researcher behind the benchmark, and needed someone to help run it. Kamradt's answer, as he later recounted it, was short: “Yes! I'm in.”

The invitation led to ARC Prize, a $1 million competition organized around a deceptively modest question. Can an AI system work out a new rule from a handful of examples? The first year brought together more than 1,400 teams and 17,000 entries. By January 2025, Kamradt, who had co-led the competition with Bryan Landers, was president and a board member of the new ARC Prize Foundation. It was a striking turn for someone whose public work had largely been about helping people use AI applications.

Yet the turn makes sense when you look at the questions he had been asking. Kamradt had already made one AI test travel far beyond his own notebook. It involved a long document, one sentence about a sandwich, and a colorful chart showing when a machine lost track of lunch.

Greg Kamradt, Mike Knoop, François Chollet and Bryan Landers together at an ARC Prize gathering
Four people, one open competition. Greg Kamradt, Mike Knoop, François Chollet and Bryan Landers at an ARC Prize gathering in New York. Photograph published by Kamradt.

First, find the sandwich

A language model may accept an enormous amount of text. That number, the context window, is easy to print on a product page. It tells you how much material the system can take in, but it leaves a practical question hanging: can the system retrieve a fact buried inside that material? Kamradt devised Needle in a Haystack to make the answer visible. He inserted a known sentence into long stretches of unrelated prose, changed its position and the total length, then asked the model to find it.

The sentence was wonderfully ordinary: the best thing to do in San Francisco is eat a sandwich and sit in Dolores Park on a sunny day. An answer about sandwiches was easy to recognize. A missed answer was just as clear. As the test varied the needle's location and the haystack's size, the results could be plotted as a grid. One glance showed where retrieval held up and where it failed. The image reached people who would never have read a table of raw model outputs.

The test had limits. Recovering one planted sentence does not prove that a model understands a long document, reasons across it, or can use it well. Its strength was its narrowness. It translated an advertised capacity into a question that anyone could understand and other researchers could repeat. Kamradt published code, and the method became a familiar way to probe long-context systems. In a field fond of numbers with a great many zeroes, a sandwich proved useful precisely because it required no interpretive acrobatics.

“Benchmarks also serve as a compass for AI research. They tell us what directions to explore.”Greg Kamradt, writing about ARC Prize

A practical education

Kamradt's route into this work was not an academic procession from one laboratory to the next. His earlier jobs included finance at NetApp, strategy and predictive analytics at Salesforce, and operations and product at Digits. His public writing from those years has the tone of someone who prefers an actionable list to a grand theory. In a 2018 post on effectiveness, he urged readers to write down the next step, ask for specific feedback, and tell the story of the impact their work produced.

That instinct carried into his AI education. Through Data Independent, his YouTube channel, he published tutorials and interviews about building applications with language models. His material walked developers through tools such as LangChain with an emphasis on practical use. He also founded Leverage, through which he helps people build with AI. The point was to make unfamiliar machinery approachable enough that someone could try it, test it and put it to work.

The side projects tell the same story in miniature. ChunkViz shows how different ways of splitting text create different pieces for a language-model workflow. Agent Traffic Control offers a radar-like view of agents doing work in parallel. ReadSide lets someone discuss a particular passage while reading an article. There is even Terra Mano, a project in which Kamradt managed 3D printing, bronze casting and hand finishing of maps of American mountains. A maker's curiosity is allowed to wander. It usually returns with a prototype.

$1MPrize pool at the 2024 launch
1.4K+Teams in the 2024 Kaggle contest
17KEntries submitted that year

The call that changed the question

When Knoop contacted him, Kamradt was already used to translating AI ideas for builders. ARC Prize offered a different job: help create a public reason for researchers to work on a particular limitation. Chollet's Abstraction and Reasoning Corpus presents small visual puzzles. A solver sees a few examples of a transformation and has to work out the rule for a new case. The puzzles are built to reward generalization to something unfamiliar, rather than the rehearsal of a known task.

Kamradt's account of the first competition is refreshingly unromantic about the work involved. The team needed a website, a technical guide, clean data, testing procedures, a public leaderboard, university visits and an active community. Landers handled much of the design and front end. Kamradt and Landers led the day-to-day push, with Knoop and Chollet as hosts and collaborators. Behind every tidy score sat the untidy labor of making a contest people could actually enter and trust.

The team also wanted researchers to show their methods. The prize encouraged open solutions, so a strong result could become more than a headline on a scoreboard. Kamradt had come from the world of APIs, customer value and applications; he later wrote that ARC Prize gave him a welcome change of pace toward AI's more fundamental questions. He also wrote that returning to a small team after much solo work reminded him how much he liked collaboration. “Mission Driven is fun,” he put it. The line has the cheerful plainness of someone surprised to find a new motive as compelling as a new market.

The question kept moving
01 / RETRIEVECan it find a planted fact?
→
02 / REASONCan it infer a new rule?
→
03 / LEARNCan it explore an unknown world?
A sketch of the progression from Needle in a Haystack to the interactive ARC-AGI-3 benchmark.

Take away the instructions

ARC Prize Foundation's next turn made the test feel less like an exam sheet and more like a game whose manual has gone missing. ARC-AGI-3, launched in March 2026, places AI agents in original interactive environments. They are given no rules and no stated goal. To succeed, an agent must try actions, notice effects, form a theory about the environment, revise it when necessary and carry useful lessons into later levels. The measure includes how efficiently it acts compared with people.

That last detail matters. A system that eventually stumbles into success may have taken far more steps than a human player. Kamradt has described efficiency as central to the foundation's view of intelligence. An interactive test can reveal whether an agent learned from its experience or simply spent enough actions to arrive at the finish. It is also easier to see the character of a failure. Did the agent miss a clue? Did it cling to a bad theory? Did it discover a rule in one level and forget it in the next?

The foundation began with human players. In an April 2026 account, Kamradt described a controlled study of 458 participants used to establish human performance on the interactive benchmark. The team adjusted scoring details after observing how people actually played. That is a reminder that a benchmark is built, calibrated and revised by people. A score may look like an immutable fact, but its meaning depends on the care taken with the test beneath it.

At launch, the foundation reported a large gap between its human reference point and frontier AI on ARC-AGI-3. The precise percentages will change as systems improve and scoring matures. The design question lasts longer: what can a system learn when the familiar clues are removed? Kamradt's role is to help make that question legible to researchers, competitors and a public accustomed to polished demonstrations.

A cabinet of small experiments

Even while leading the foundation, Kamradt keeps shipping smaller projects. His 2026 list includes agnts.sh, which presents links in forms useful to agents; InterAuth, which offers scoped read access to a Google Doc; ReadSide; and RunReq, a monitor of hiring activity at frontier AI organizations. Their range resists a neat job title. So does the SnakeBench project, which set language models against one another at Snake. The tests, tools and occasional joke all belong to the same workshop.

There is a personal continuity here, too. In 2018, Kamradt advised readers to document their work and explain its impact. Years later, his own account of ARC Prize itemized the labor behind a competition instead of skipping to the triumphant photograph. That choice makes the later title, president, feel less like a sudden promotion than the result of a familiar habit: make a problem concrete, gather people around it and keep track of whether the work changes anything.

The sandwich still makes a good opening scene because it is so small. It asked one machine to find one sentence. The ARC Prize work asks whether a machine can infer rules, explore, and learn with something like human economy. Kamradt has moved from teaching people to use AI to helping define what AI must demonstrate. Each new test raises the difficulty, and each one begins with the same editorial discipline: write a question whose answer can be seen.