A quarterly business review is not supposed to be dramatic. It is a deck: charts, account notes, a diagnosis of the quarter, a proposal for the next one. Yet this modest artifact contains exactly the kind of trouble that exposes an ambitious piece of software. The data lives in different systems. The customer has said one thing on a call and done another in the product. A recommendation can be plausible, polished, and embarrassingly wrong. At Handshake, Darshdeep Hora and his colleagues decided to give the whole job to an AI agent.
Their first question was not how fluent the agent sounded. It was whether an account manager could carry the finished deck into a customer meeting without fixing it. The distance between those standards is Hora's territory. A machine that performs most of a workflow is an amusing colleague. A machine entrusted with the whole workflow needs judgment, evidence, and the good manners not to pitch a customer something it already owns.
That distinction is easy to lose in a demonstration. A slide can be handsome while its premise is false. A number can be accurate while the story built around it is nonsense. Quarterly reviews make such mistakes unusually public: the customer is sitting in the room, often with the records open and an opinion about what happened. The useful unit of progress was therefore not a completed slide or even a completed deck. It was a deck that could withstand informed disagreement. That set a sterner target than fluency and a more interesting one than speed.
The agent they built reduced preparation time from 45 minutes per account to around 10. Across a 70-account book, the team calculated that each seller recovered roughly 40 hours every quarter. The arithmetic is attractive. The route to it is more revealing: interviews, failed assumptions, 54 weighted criteria, and a repeated willingness to discover that the software was confidently misunderstanding the assignment.
A standard before a system
Hora, John Kottlowski, and Datta Kaligotla began with the people who already knew the work. Handshake's account managers described the recurring labor of assembling a QBR: collect usage data, reconcile metrics, review call logs, understand the customer's goals, form a narrative, and build the deck. Some of it was button-clicking. Some was interpretation. All of it had to arrive in one coherent document.
The first evaluation plan sounded sensible. Gather strong historical decks and compare the agent's output with those examples. It soon collapsed under the variety of human practice. Different account managers emphasized different things, and hundreds of old decks would be needed to produce a useful signal. History, in this case, was an archive of preferences rather than a clean answer key.
“And to make an agent good, we first have to define ‘good.’”Darshdeep Hora, John Kottlowski and Datta Kaligotla
So the team made judgment explicit. Could the customer verify every claim from its own records? Did the recommendations follow from stated goals? Did the title agree with the metrics? Each question became a criterion. Criteria acquired weights. The rubric eventually covered growth, risk, positive value, and mixed signals. This was not a philosophical seminar about quality. It was a scorecard with consequences.
One failure earned a special place in the account. Early versions recommended modules the customer had already bought but had not yet deployed. The sentence looked like an expansion pitch. The actual conversation should have been about adoption. Another flourish of prompting would not solve it. The agent needed better context about the customer's existing products before it tried to write the narrative. The error was linguistic on the surface and architectural underneath.
There is a tidy moral here, which makes it suspicious. The more useful conclusion is messier. Building the agent changed the work because it forced a group of experts to say what they meant by good. The account managers supplied the standard. Engineers turned it into tests. Failures improved both the software and the definition. The system did not remove people from the process. It concentrated their contribution where it mattered.
The career in the feedback loop
Hora's public biography is compact, but its turns are instructive. He grew up in Lucknow, India, attended The Doon School, then crossed an ocean to Wesleyan University in Connecticut. From 2011 to 2015 he studied computer science and physics. He also spent all four undergraduate years on the men's crew team, an activity in which synchronized effort is less a metaphor than a rule of staying upright.
In his senior spring, a late scratch changed Wesleyan's first varsity lineup before the season opener. Hora stepped into his first varsity race without practice in that combination. The boat won. A teammate praised him for stepping up. It is a small campus anecdote, worth keeping small, but it fits the later record: enter an unfamiliar system, find the rhythm quickly, and contribute without making a ceremony of the adjustment.
College also supplied a stranger rehearsal for verification. In a summer forensics class, Hora photographed a staged crime scene and analyzed hair and fiber samples. Elsewhere on campus, he worked as a research intern in physicist Francis Starr's group. The future engineer was moving among evidence, models, and the awkward fact that an elegant theory still has to survive what is lying on the laboratory table.
Computer science, physics, laboratory research, and four seasons of men's crew at Wesleyan.
Engineering roles at Microsoft and Coinbase.
Founding engineer at Blackbird Labs, working close to the beginning of a restaurant technology company.
Public work on agent monitoring at Sierra, followed by the QBR-agent case study at Handshake.
After Wesleyan came engineering roles at Microsoft and Coinbase, then a founding-engineer post at Blackbird Labs. Those names describe very different operating conditions: a mature platform, a fast-moving financial network, and a restaurant technology company at its beginning. The dates and private details are not needed to see the range. Hora learned inside both large systems and unfinished ones.
An earlier side project shows the same breadth in miniature. Playlift was an iOS app for managing collaborative playlists. Hora's description of the work runs from a Django and PostgreSQL backend to Swift device code, Spotify and Apple Music integrations, and an AWS deployment pipeline. It is the kind of project where the interface appears simple only because several disagreeable services have been persuaded to cooperate behind it.
Former colleagues described him as tenacious, quick to learn unfamiliar technologies, proactive about engineering bottlenecks, and optimistic under difficult deadlines. One recalled a stable game battle delivered shortly after Hora joined the team, along with work on an event pipeline that needed careful failure and retry logic. Reliability again, although then it arrived wearing the less fashionable clothes of retries and deadlines.
Who monitors the monitors?
At Sierra, Hora co-authored a 2026 explanation of an always-on evaluation layer for customer-service agents. These monitors use a language model to review conversations for qualities such as looping or rising frustration. The immediate puzzle is deliciously circular: if a model judges the agent, who judges the model that judges the agent?
The answer was another grounded loop. Teams labeled conversations. Multiple models evaluated them. Disagreements exposed definitions that were too broad, too narrow, or missing context. Those cases returned to the training and evaluation sets until the models agreed more consistently and could explain each flag. A polite word might signal patience in one exchange and sarcasm in another. The monitor had to understand the difference well enough for a human reviewer to trust its rationale.
This work and the Handshake project share a temperament. Neither treats human judgment as mystical decoration. Both begin by gathering difficult examples, describing the desired behavior precisely, and checking the machine against a standard people can inspect. The ambition is large, but the method is almost domestic: write down the rule, test it, notice the mess, clean up, repeat.
The future arrives as a workflow before it arrives as a revolution.
Hora's latest work argues for an unfashionable form of confidence. It is not confidence that the model will somehow manage. It is confidence earned through visible criteria, inconvenient counterexamples, and numbers that survive contact with the office. The 78 percent reduction matters because it describes time that people received back. The 54 criteria matter because they describe the care required before that exchange became responsible.
The final 22 percent is where a system learns the customer's actual situation, where a generic recommendation becomes an informed one, and where a draft becomes something a person will sign. It is also where Hora's story is most legible. The engineer does not appear at the center of a grand prophecy. He appears beside a feedback loop, asking whether the output is ready yet. Then, inconveniently and usefully, asking again.
Continue reading