A poker player and a language model share an awkward problem: both must act before certainty arrives. The difference is that the poker player knows the chips are real. In Ivey, a reinforcement-learning engine Aaron Jin built with Ryan Cheng at Stanford, the machine learned pre-flop decisions through Q-learning. With careful state design, the pair reported that ordinary reinforcement-learning methods could discover the familiar logic of position and hand strength. The project was named, with appropriate ambition, after Phil Ivey. It is a small clue to Jin’s larger habit. He likes problems in which judgment can be represented, put under pressure and made to answer for itself.
That habit now operates at industrial scale. Jin is a founding engineer at AfterQuery, the San Francisco company that turns expert work into data for training and post-training large language models. His own description is admirably free of incense: he builds “high-throughput data pipelines.” The phrase evokes loading docks more readily than artificial intelligence, which is part of its charm. Models may speak in smooth paragraphs; somewhere behind them, an engineer still has to decide which examples move where, how quickly they arrive and whether the lesson survives the trip.
AfterQuery’s thesis is that expert knowledge lives not only in final outputs but in choices, tradeoffs and context. Its products include supervised fine-tuning data, reinforcement-learning tasks and rubrics, agent environments and demonstrations of people operating software. This is an operationally fussy business disguised as an epistemological one. If you want a model to learn how a lawyer synthesizes a case, how an analyst examines a transaction or how a developer fixes a repository, the polished answer is only the last square of a much larger board.
First he built the feature. Then he built the test.
Jin arrived at this layer by moving steadily upstream. At Apple, he says, he integrated Apple Intelligence into Safari Suggestions and built predictive models of battery performance. One piece concerned what a product might recommend; the other concerned how the device itself would behave. Consumer software permits no airy distinction between the model and the machine. A useful prediction that drains the battery is a failure wearing a clever hat.
At LinkedIn, Jin moved from intelligent features toward the machinery that judges intelligence. He built a distributed benchmarking platform to orchestrate model evaluations. “Evaluation” can sound like the quiet hour after the real work. In machine learning it is closer to the constitution. An eval decides what counts as better, what regressions are visible and which improvements deserve to ship. Distributing those tests across infrastructure turns a research question into an engineering system: repeatable, comparable and difficult to flatter.
Product intelligence and performance prediction
Distributed model evaluation
Training and post-training data pipelines
The progression is neat without being preordained. First, make machine intelligence useful inside a product. Next, make its quality measurable across a platform. Then work on the examples and feedback that shape the model before the user ever sees it. Feature, test, lesson. Engineers rarely receive careers with such tidy chapter headings, but Jin’s work has supplied them anyway.
The career-fair photograph provides a pleasant reversal. Jin, a Stanford BS/MS computer science graduate, is back on campus on the employer’s side of the table. AfterQuery’s banner stands behind the group. Tote bags, QR codes and neat rows of cards cover the table. Recruiting fairs are where ambitious abstractions become ordinary conversation: What do you build? Who needs it? Are you hiring? The company was recruiting across engineering, research, operations, growth and sales - the human departments required to manufacture better machine reasoning.
A workbench with unusually wide shelves
Jin’s public projects make more sense as a workbench than a portfolio. There is a GPT-2-style decoder-only transformer, complete with sentiment analysis, paraphrase detection and sonnet generation. There is PolymarketFM, a foundation model for zero-shot prediction-market forecasting. There is Tungsten, a compiler built from scratch without libraries. Flow automates knowledge work; Roam turns inspiration into trips. Slop generates short videos, Jaike turns lectures into “brainrot,” and Jonin is a pixel ninja platformer. Solemnity would not survive this repository list, and that is to its credit.
The list also resists a fashionable fiction: that serious engineers are defined by a single stack or problem category. Poker, compilers, 3D vision, travel planning and video generation do not share a market. They share the discipline of implementation. Each forces an abstract claim to survive contact with software. Does the agent learn? Does the parser work? Can the model represent a room without becoming too large for the phone? Does the app remain amusing after the demo?
The room, reconstructed
PocketNeRF, a 2025 Stanford computer-vision project by Jin, Lucas Brennan and Ryan Suh, offers the sharpest technical example. Neural radiance fields can reconstruct a three-dimensional scene from images, but their appetites for time and memory make mobile use difficult. The team built a pipeline for indoor spaces using sparse smartphone images. Jin handled the integration of structural priors: the useful assumption that rooms contain planes, walls tend to be vertical and indoor geometry is not an anarchist collective.
The project combined those constraints with content-aware quantization, allowing different parts of the model to use different numerical precision. The team reported faster training and roughly an order-of-magnitude reduction in model size without sacrificing visual quality. The underlying maneuver is elegant: give the system a little more knowledge about the world, then ask it to carry fewer bits. Intelligence, in this case, looks like knowing which details deserve expensive representation.
This is the thread connecting a student project to AfterQuery. Training systems improve when the examples contain the right structure. In PocketNeRF, geometry narrows the problem. In poker, the state representation determines what the agent can learn. In post-training an LLM, a rubric tells the system which response is better and why. None of these structures is simply “data” in the warehouse sense. Each is a considered account of the world.
Building inside a fast-growing machine
The context around Jin’s current work has grown quickly. In April 2026, AfterQuery announced a $30 million Series A at a $300 million valuation and said it had passed a $100 million revenue run rate. Private-market records later tracked a Series B in August. Those are company figures, not personal achievements, but they alter the conditions of the engineering job. A pipeline that serves an experiment and a pipeline that serves a rapidly scaling supplier of training data are relatives, not twins.
Scale does not dissolve the original question. It makes the question expensive. How do you preserve expert nuance when throughput rises? How do you make an evaluation broad enough to matter and precise enough to diagnose failure? How do you distinguish a model that has learned a pattern from one that has learned to please the grader? The sensible answer is not a slogan. It is a set of data models, queues, tests, interfaces and judgment calls, revised until they behave.
Jin’s title is “founding engineer,” a designation that can cover a small republic of responsibilities. What is public is narrower and more useful: the pipelines are high-throughput, and their purpose is training and post-training LLMs. There is no need to decorate that description. Infrastructure is where ambition is forced to acquire error handling.
His career so far has moved from predictions inside a device, to tests across a network, to lessons for models that will operate across many kinds of work. The public trail is technical rather than confessional. It says little about grand aspirations and quite a lot about what he chooses to build. That may be the more reliable document. A compiler, after all, is an unforgiving diary.
Return to the poker table. The agent cannot see the opponent’s cards. It still needs a representation of the situation, a reward and enough trials to tell a sound strategy from a lucky hand. The same essentials survive in more elaborate forms throughout Jin’s work. Structure the problem. Move the examples. Measure the result. Repeat before certainty arrives.
Artificial intelligence is presently rich in people selling answers. Jin’s engineering has followed the less theatrical work that surrounds them: the product constraints, the evaluations, the training data and the systems that keep all three moving. The interesting moment is not when the model finishes its sentence. It is the moment after, when someone must decide whether the sentence was any good - and then build a machine capable of learning from the verdict.