Long before Eddie Zhou was explaining the architecture of AI agents, he was standing in front of a fire engine with two Princeton classmates and an iPad. The year was 2013. Their project, FireStop, was built around a problem that sounded ordinary until the sirens began: firefighters needed building plans, hydrant locations, hazard warnings and inspection data, and they needed all of it before a bad situation acquired worse manners.
The team wanted to shave 90 seconds from the opening of a response. It built a cloud-connected system that could keep recent information available when a vehicle lost its connection on the way to a call. There was no fashionable talk of agents or context windows. There was simply context, a user under pressure, and a machine expected to be useful.
That early project is a neat key to Zhou's later career. At Google Brain, he helped put deep-learning models into products used by vast audiences. At Glean, where he became a founding engineer and AI team lead, he worked through the company's progression from enterprise search to an assistant and then to agents. The surface changed. The stubborn problem underneath did not: how do you get the right knowledge into the right hands, while respecting who is allowed to see it and how much confidence the system deserves?
A young engineer learns to cut
Zhou studied Operations Research and Financial Engineering at Princeton, adding a certificate in Statistics and Machine Learning. His extracurricular list had the cheerful excess of an undergraduate determined to try the whole buffet: club basketball, Princeton Data Science, the Entrepreneurship Club, Tau Beta Pi, Sigma Xi and Phi Delta Theta. FireStop, co-founded with Charlie Jacobson during their sophomore year, was the practical item on the plate. It moved from campus accelerator to pilot programs with New Jersey fire departments.
His senior thesis went in the opposite direction, deep into models. It contained three projects: temporal statistical models for vision-based autonomous driving, nonparametric long short-term memory networks, and an R package for Bayesian-optimized deep learning. A Princeton faculty member described it as three independent projects in one. The university's Center for Statistics and Machine Learning gave Zhou its first Independent Work Award in 2016.
The thesis was 34 pages. In its acknowledgements, Zhou credited his adviser, Han Liu, with teaching him the value of brevity alongside statistical parsimony. It is a revealing detail. Machine learning encourages abundance: more parameters, more data, more compute, more pages explaining the first three. Zhou learned early that subtraction can be an engineering act.
After Princeton came Google's Applied Brain team. Zhou worked where research meets the rather less forgiving world of production. The models reached Web Search ranking, Assistant language understanding, YouTube and advertising recommendation systems. This was deep learning after the conference applause, when latency, reliability and product constraints arrive carrying clipboards.
The small search box and the large mess behind it
Zhou joined Glean's early team in 2019. The company's first public problem was enterprise search, a category that had accumulated more disappointment than glamour. Web search could make the world's information feel orderly. Inside a company, information was scattered across documents, messages, code repositories and business applications. Each organization had its own acronyms, habits and tribal map. Permissions made every answer personal. The same query from two employees could properly produce two different sets of results.
In a 2021 essay, Zhou and colleague Mrinal Mohit laid out why the gap persisted. Company content is unique, often specific and legible mainly to insiders. There are fewer repeated pages that can answer the same question, and less public feedback from which a ranking system can learn. The search box may be small; the epistemology behind it is warehouse-sized.
Then large language models became fluent. In February 2023, amid the first feverish months of ChatGPT, Zhou and Glean chief executive Arvind Jain published a distinction that has aged well: coherence is not knowledge. A model can arrange words convincingly while lacking current company facts. An enterprise system also has to understand permissions, explain where an answer came from and adapt to an organization's private vocabulary. Eloquence, in this setting, is a charming colleague who has misplaced the filing cabinet.
Zhou's public enthusiasm was real. He wrote that he wanted to bring ChatGPT to the workplace. Yet his excitement came with conditions. The model needed grounding in enterprise knowledge. It needed to know that a mention of personally identifiable information might live in a document about log data. It needed to surface the information an employee could actually access. Trust was architecture, not garnish.
Search widens into action
By 2025, Zhou described Glean's product evolution as a widening slice of a knowledge worker's journey. Search helped after a person had already decided what information to find. An assistant could synthesize what it found. An agent could move earlier and later in the task: help decide what was needed, make a plan, call tools, retrieve more context and take an action.
His definition of an agent is deliberately provisional. One useful framing is a system that can form a plan from a set of tools and then execute it. But the interesting questions begin immediately. Can it revise the plan? Are its tools fixed functions or other agents? How many attempts should it get? When does flexibility become a politely automated form of wandering around?
Zhou favors practical bounds. A fixed limit on steps can keep an agent from looping indefinitely. High-stakes work may deserve a frozen workflow. Dynamic systems can be valuable where discovery matters, then their successful paths can be turned into repeatable routines. This is engineering by adjustable dial, not manifesto.
He also sees a role for delegation. A central agent can call specialized subagents, allowing separate teams to improve separate capabilities. The idea borrows from familiar software design: modular systems are easier to develop, diagnose and scale than one magnificent knot. Zhou joked that experienced software engineers were probably punching the air as they watched agent systems ignore lessons that distributed computing had already paid to learn.
That line captures his public temperament. Zhou does not pretend every component must be invented inside Glean. Open-source frameworks are candidates when they meet the requirements. The team builds its own when efficiency, security or deployment constraints demand it. Reuse is leverage, not surrender. Reinventing the wheel is only romantic if one has never maintained a wheel.
The discipline after the demo
The most difficult part of an enterprise agent may be reflection: deciding what to do after a search returns incomplete results. A model cannot reliably notice a missing company document if it does not know the document exists. It does not know what it does not know, and a confident instruction to “never lie” cannot repair absent context.
Zhou's answer is measurement. Glean evaluates systems offline in batches, creates known-correct and adversarial examples, traces failures through compound systems and asks which stage produced the bad result. Online guardrails can catch narrower problems, such as whether an answer is grounded in the material it received. They cannot certify every output as true. Users care about the whole answer; engineers still need to isolate the broken part.
This measured caution has accompanied increasingly ambitious releases. Glean reported 94 percent completeness for its second agentic engine on a suite of complex enterprise tasks in 2025. In 2026, Zhou co-authored the launch of Waldo, a specialized model that plans search before a frontier model writes the final answer. In Glean's production harness, the company reported roughly half the latency and one-quarter fewer tokens for search, without a measured quality loss.
Indexed to the prior harness at 100. Figures reflect Glean's reported production results, not an independent benchmark.
Waldo is an apt endpoint for Zhou's arc so far. The model is specialized, assigned the job of finding and assembling evidence while a frontier model handles broader reasoning. It plans the search, decides which tools to use and stops when it has enough. Beneath the contemporary terminology is the old FireStop requirement: gather what matters, preserve the necessary access and deliver it before the person has lost valuable time.
A life outside the context window
Public technical profiles have a habit of turning people into lists of systems. Zhou's official biography supplies a useful correction. Away from models, he plays basketball, enjoys California sunshine and watches his cat do weird stuff. His LinkedIn profile lists English and Chinese at native or bilingual proficiency and French at professional working proficiency. The club basketball line from Princeton survived into adulthood as a hobby; the machine-learning line became the job.
His stated ambition for Glean has remained broad and plain: make life easier for knowledge workers. Search, in his telling, was the first delivery mechanism, not a discarded identity. Assistants and agents extend help across more of the task. The continuity matters because it resists the industry's fondness for declaring a new epoch every quarter.
Zhou's career is less a sequence of pivots than a repeated encounter with the same unruly fact: information has a location, an owner, a permission boundary and a moment when it becomes useful. The model may be clever. The interface may be charming. The system earns its keep only when it knows which plan to retrieve, which tool to call, which document a person may read and when to admit that it needs another pass.
The fire engine in that old photograph is bright red and gloriously unsubtle. Zhou and his classmates stand in front of it smiling, an iPad held between them. Enterprise AI is harder to photograph. Its work happens in indexes, graphs, model calls and access checks. Yet the promise fits in the same frame: when someone needs to act, make sure the knowledge gets there first.