Parth Patel works in the part of artificial intelligence where confidence meets a locked door. On one side sits a model, fluent and impatient, producing code that looks finished. On the other side sit hidden tests, clean baselines, grading rules and reviewers asking the impolite question: does any of this actually work? Patel is a founding engineer at AfterQuery, a San Francisco applied research lab built around the notion that expert judgment can be captured, structured and made useful to machines. It is a practical occupation with a philosophical shadow. Answers are plentiful. Judgment remains expensive.
Patel's route to this work runs through UC San Diego's Jacobs School of Engineering, where his public profile places him from 2021 to 2025. Soon after came AfterQuery, an early-stage company formed to turn real professional work into training material for AI. The title “founding engineer” is concise, but the territory is broad. AfterQuery describes four kinds of material: supervised examples, reinforcement-learning prompts and rubrics, software environments for agents, and recordings of people completing computer tasks. All four try to preserve something ordinary datasets tend to flatten: the sequence of decisions between a problem and its answer.
That sequence matters because professional work is rarely a clean exchange of prompt and response. A lawyer weighs exceptions. An analyst decides which comparison is honest. A programmer notices that a fix passes the obvious test but creates a less obvious defect. The finished artifact hides these forks in the road. AfterQuery's premise is that a useful model needs to encounter the forks, not merely admire the destination.
“The scarce ingredient is not another answer. It is the judgment that makes an answer defensible.”
The engineer at the examination desk
A public credit makes Patel's corner of that premise unusually concrete. ReactBench, a benchmark for coding agents working on React software, names him as a data adviser. The benchmark asks agents to perform two kinds of work. In “Writing React,” an agent implements a real task drawn from a merged open-source pull request. In “Fixing React,” it improves a component with known problems without being told what those problems are. The setup resembles a demanding take-home exam written by people who have watched too many clever students exploit the wording.
The reference solution is withheld. Hidden behavior tests are injected later. A separate verifier checks whether the submitted code introduced React problems involving correctness, state, performance, accessibility or security. The agent cannot inspect that verifier during its run. An adversarial agent is then invited to try to cheat the system. If the trick works, the task is repaired or removed.
It is tempting to call this quality assurance, which is accurate in the way that calling a kitchen a room with heat is accurate. Evaluation design determines what counts as progress. If the test rewards a shortcut, the model learns the shortcut. If it misses accessibility or security, a polished score may conceal an ugly product. A benchmark must test the agent, test its own tests and retain enough skepticism to know the difference.
The numbers expose the editing. ReactBench began with 23,100 candidate pull requests. Only 3,852 showed the right React signal. Filters narrowed that pool to 2,172. Manual review left 129. Thirty-nine tasks survived. The retained share was about 0.17 percent of the original pool. Software culture celebrates scale, but useful evaluation data often arrives through subtraction. The benchmark becomes credible because nearly everything is refused.
A company built around the missing middle
AfterQuery calls itself an applied research lab. Its website opens with a sentence that doubles as a manifesto: “We teach machines how experts think.” The claim leads directly to the missing middle between instructions and finished work. Models can learn from prompt-response pairs, but difficult jobs also contain tools, context, revisions and standards. A good answer in finance is not good because it resembles finance. A good program is not good because the syntax is tidy. It must survive contact with the world it is meant to serve.
The company therefore builds more than static datasets. Its agent environments reproduce APIs, tools and services so models can be trained inside workflows. Its rubrics turn expert judgment into reward signals. Its computer-use trajectories record how people navigate software from beginning to end. The work asks engineers to build both the stage and the critic: a place where an agent can act, plus a reliable way to decide whether the performance deserves applause.
Patel's advisory credit sits inside this discipline. ReactBench does not use a general-purpose model as the sole judge. It pins the scanner version, keeps the grading container separate and restores protected configuration before checking a submission. Reviewers inspect trajectories and final patches before turning failures into conclusions. A zero can mean the model failed, but it can also mean the instruction was vague, the test was brittle or the infrastructure broke. The grown-up work begins when the scoreboard stops pretending those are the same event.
There is a pleasing irony here. AI systems are marketed as a way to remove labor, while reliable AI evaluation demands a great deal of careful human labor. Someone must select the task, clarify the requirement, write the hidden test, challenge the verifier and inspect the strange run that almost passed. Automation does not abolish the editor. It gives the editor a larger stack of drafts.
From student to early operator
Patel's public professional timeline is compact. It shows four years at UC San Diego, then the early engineering team at AfterQuery. His X profile describes him simply as “founding engineer @AfterQuery.” His GitHub includes solutions to Advent of Code, the annual series of programming puzzles. In 2026, he was named as one of two founding-engineer speakers for an AfterQuery leadership panel co-hosted by Scope and the Trojan Investing Society at the University of Southern California.
The panel placed engineering beside research, growth and strategic projects. That grouping is revealing. Expert-data work is not a sealed laboratory exercise. It requires recruiting people who know a domain, translating their work into repeatable tasks, building the software that delivers those tasks and checking that the result measures what its designers claim. The founding engineer is not merely maintaining a product after the thesis has been settled. The software helps settle the thesis.
This is an operator's version of engineering: close to the uncertain edge, where categories overlap and clean job descriptions arrive late. The visible outputs may be datasets and benchmark scores. Underneath are decisions about whose expertise counts, how a task should be represented and which failure matters. Those choices are technical, editorial and commercial at once.
A benchmark is an argument in executable form. It tells the machine what matters, then waits to see if the machine finds a loophole.
The test becomes the product
For years, AI benchmarks were treated like scoreboards hung outside the factory. Build the model, run the test, publish the number. Coding agents complicate that picture because the agent's output can change a living system. A patch may pass the stated requirement while slowing the page, breaking keyboard access or planting a security problem. Evaluation is no longer a ceremonial lap after development. It becomes infrastructure used to shape training, select models and decide which work can be trusted.
That shift makes Patel's field less peripheral than it first appears. The valuable artifact is not only a smarter model. It is also a more discerning environment around the model: tests that remain hidden, rubrics that reward the right behavior, tasks drawn from real work and reviewers who can tell a model error from a measurement error. The more fluent AI becomes, the more valuable disciplined doubt becomes.
Patel is early in his career and early in this layer of the industry. The available record already ties him to a clear problem: translating human judgment into machinery without pretending judgment is simple. That work will not always make for a dramatic demo. It may look like a filter rejecting 22,971 pull requests, a second container checking the first, or a reviewer staring at a failed run and refusing the easy explanation.
The machine produces an answer. Patel's world begins one beat later. It asks for evidence, checks the seams and keeps the door locked until the work can open it. Fluency may knock first; verification still keeps the key.
Keep reading