Founding engineer at AfterQueryCo-author of UI-Bench and IDE-BenchFrom Fortuna to real-world AI evaluationFounding engineer at AfterQueryCo-author of UI-Bench and IDE-BenchFrom Fortuna to real-world AI evaluation

People / Engineering / Applied AI

Agustin Garcinuno Took the Long Route to Useful Data

From a redwood town to resale software, Wall Street and AI evaluation, his path keeps returning to the same practical question: does the thing actually work?

The shortest description Agustin Garcinuno gives of himself is also the most revealing. He is a developer in San Francisco, he likes building things with code, and when he is not doing that he likes being outside. Lately, he says, he spends most of his time trying to make useful data at AfterQuery. It is an almost aggressively modest summary of a public geography that reaches from the cold southern tip of Chile through a Northern California company town, a university in Philadelphia, a resale startup and two investment firms, then arrives at the intricate plumbing of artificial intelligence.

The phrase to notice is “useful data.” There is data that fills a dashboard, data that decorates a pitch deck, and data that settles an argument. Garcinuno’s public work belongs mostly to the last category. His recent projects ask AI systems to design interfaces and edit real software repositories, then surround those tasks with rules sturdy enough to tell whether the machines succeeded. In a field that produces confident demonstrations by the hour, he has been helping build the measuring tape.

A route in five stops
Punta Arenas
Scotia
Fortuna
Philadelphia
San Francisco

Before the benchmark, a listing nobody wanted to write

Garcinuno’s public biography points first to Punta Arenas, Chile, then to Scotia, California. Scotia sits among Humboldt County redwoods and is known for its long association with the timber business. He graduated from nearby Fortuna High School in 2019, then went east to the University of Pennsylvania, where he studied Environmental Management and Computer Science. He received the Mark E. Reed Scholarship three times, in 2019, 2021 and 2022.

At Penn, the two halves of that degree made an unusual pair: one concerned with natural systems and their stewardship, the other with systems people make. His working life initially tilted toward finance. He spent the summer of 2021 at Soros Fund Management and, after university, worked as an analyst at Goldman Sachs for the first seven months of 2025. But between school and Goldman sat a project closer to the habits of an engineer.

Flipkit began with secondhand clothes and an exquisitely ordinary complaint. Listing is boring. Garcinuno and co-founders Haley Kang and Tavi Kim had experience reselling clothing and knew the repetitive work of putting the same inventory across marketplaces. They spoke with hundreds of other resellers and built software to automate omnichannel sales. The product served secondhand apparel businesses, joined Penn Venture Lab’s VIP-X alumni ranks and became a semifinalist in Penn’s 2024 Startup Challenge.

Flipkit did not become a permanent stop. Garcinuno lists the company as a one-year chapter, from January through December 2024. Still, it supplied a pattern that would return in a more technical setting: begin with an actual workflow, locate the part consuming human attention, and make the machine carry more of it.

Giving AI design a scoreboard

In August 2025, Garcinuno joined AfterQuery as a founding engineer. That same month, he, Sam Jung and Spencer Mateega released UI-Bench, a benchmark for a new class of AI products that turn text prompts into apps and websites. These tools were promising polished results in minutes. The problem was not a lack of impressive examples. It was a lack of public, rigorous comparison.

10text-to-app tools compared
300generated websites reviewed
4,000+expert pairwise judgments

UI-Bench used 30 prompts to produce 300 sites across 10 tools. More than 4,000 expert judgments compared the results in pairs. A model derived from TrueSkill, the rating system developed for competitive matchmaking, turned those comparisons into rankings with confidence intervals. The team released the prompt set, evaluation framework and a public leaderboard. Visual taste, famously slippery in conversation, had been given a repeatable contest.

That scale was part of the argument. A single beautiful page might say more about a lucky prompt than a capable system. Thirty prompts make luck work harder. Expert reviewers add another constraint: the judgment comes from people expected to notice hierarchy, composition and polish, not from an automatic proxy that merely resembles taste. The benchmark’s confidence intervals also leave room for a useful admission. A ranking is evidence with uncertainty attached, not a marble tablet delivered from the laboratory.

The method matters. Asking whether a generated interface is “good” invites a fog of personal preference. Asking a qualified reviewer which of two interfaces is better creates a decision. Repeating that decision thousands of times creates evidence. UI-Bench did not eliminate judgment; it organized judgment so the results could be examined, rerun and challenged.

UI-Bench
  • Visual design under review
  • 30 prompts
  • 300 generated sites
  • Expert pairwise comparison
IDE-Bench
  • Software work under review
  • 80 engineering tasks
  • 8 unseen repositories
  • IDE-native tools and tests

Then the test entered the codebase

Five months later came IDE-Bench, co-authored by Garcinuno with Spencer Mateega, Jeff Yang, Tiana Costello, Shaurya Jadhav and Nicole Tian. If UI-Bench asked whether AI could produce a convincing surface, IDE-Bench asked whether an autonomous coding agent could survive the basement: find the right files, understand unfamiliar logic, make a change and prove it worked.

The benchmark contains 80 tasks across eight repositories spanning C, C++, Java, Python, TypeScript and a full MERN stack. The repositories had not previously been published, a deliberate defense against testing a model on code it might have encountered during training. The tasks resemble daily engineering work: fixing bugs, adding features, refactoring systems and improving performance.

The repositories are deliberately varied: a network traffic analyzer, a memory profiler, an event callback system, a game-engine service and a cross-lingual translation application among them. That variety prevents “coding ability” from collapsing into one favored language or one familiar project shape. An agent must repeatedly orient itself, form a plan and leave the codebase in better condition than it found it. The test cares about the modification, not the eloquence of the explanation preceding it.

Crucially, the agents do not receive only a terminal. They work through a structured collection of tools resembling an AI-native development environment: code search, file inspection, editing, tests, API calls, database queries, browser interaction and WebSocket checks. This makes the benchmark less like an exam question and more like a workday with unusually attentive supervision.

There is a quiet consistency between Flipkit and IDE-Bench. Both begin by respecting the workflow. A reseller moves inventory through listings and marketplaces. A developer moves through search, edits, execution and verification. Automation becomes credible only when it enters that sequence without pretending the difficult parts have vanished.

The week hidden inside a cluster

Benchmarks depend on environments, and environments have a talent for becoming someone’s entire job. At AfterQuery, Garcinuno and research operations lead Sam Jung initially built a Google Kubernetes Engine cluster to generate sandboxes for AI-agent simulations. Each sandbox needed its own resource limits, file paths, API keys and saved state. As tasks multiplied and agents ran for hours, the manual versioning and maintenance began crowding out the research the infrastructure existed to support.

AfterQuery wordmark on a dark technical grid
THE WORK BEHIND THE WORDMARK - AfterQuery’s simulations require isolated, versioned environments that can keep running when agents take the scenic route.

The team moved the runtime layer to Daytona. The replacement took one day. Preconfigured snapshots could start warm, hundreds of isolated sandboxes could run in parallel, and developers could preview servers, inspect logs or intervene through a terminal. Sandbox creation time fell by 86 percent. More than 40 engineering hours a week were recovered.

That last number deserves translation. Forty hours is not an optimization hidden after a decimal point. It is a person-week. It is the difference between tending the apparatus and improving the experiment. Garcinuno described versioning as central to the process, and said moving that work into snapshot infrastructure had been a major benefit. The next ambitions are similarly practical: resize CPU and storage during long trials, and iterate faster on simulated iOS workflows.

1 dayto replace the in-house setup
86%less sandbox creation time
40+engineering hours saved weekly

A career arranged around proof

It would be easy to describe Garcinuno’s route as a collection of pivots: Chile to California, environmental study to software, resale to finance, finance to AI. That reading gives the nouns too much authority. The verbs are steadier. Observe a process. Find the drag. Build the system. Decide how success will be measured.

His public biography remains concise, and the work is still early. There are no grand declarations attached to it, only a preference for useful data and a growing set of artifacts that define usefulness with uncommon specificity. A good interface must win comparisons. A coding agent must change a real project without breaking it. A research environment must preserve state, isolate risk and return time to the people running the experiment.

Outside, in Garcinuno’s formulation, is where he goes when he is not building with code. It is also a fitting corrective to the abstractions of AI. The world outside the demo has stubborn edges. Software has dependencies. Design has viewers. Infrastructure has limits. Claims must eventually meet all three. His recent work is about arranging that meeting, then keeping score.