Field NotesNew York●AI agents meet network infrastructure●Experiments, checks and one $250 lesson●

People / Artificial intelligence / Infrastructure

The Engineer Who Keeps AI on a Short Leash

Anton Bubna-Litic spent years turning mathematical models into working systems. At NetBox Labs, he is applying the same practical instinct to AI agents - testing the fashionable future until it becomes useful infrastructure.

There is a particular kind of optimist who carries a checklist. Anton Bubna-Litic appears to be one. He works at the loud edge of software, where AI agents are promised as tireless colleagues and every week brings a fresh reason to rebuild the toolchain. Yet his public record reads less like a manifesto than a lab notebook: test the holdout, reuse the code, watch the outputs, write down the mistake. The future is welcome. It will simply have to pass inspection.

That combination now has a title. Bubna-Litic is the Founding AI Engineer at NetBox Labs, the company built around NetBox, an open-source system for modeling networks and infrastructure. He joined in August 2025 after roughly a decade at the data and analytics company Quantium. The move put him in New York and inside a problem that is suddenly urgent: if AI agents are going to act on real infrastructure, what must they know, where may they act, and how can anyone tell whether they did the right thing?

His answer begins a long way from server racks. It begins with eyesight.

01 / The first patternBefore the agent, there was the texture

At the Australian National University, Bubna-Litic studied mathematics in the Bachelor of Philosophy program, an honours degree designed around research. In 2009, he undertook a computational-vision project with visual scientist Ted Maddess. The subject was texture: the subtle statistical structures that make one field of dots or marks look different from another, even when their simpler properties appear alike.

The work matured into conference proceedings and papers about texture discrimination and perceptual salience. One paper investigated how many distinct mechanisms might be required to tell apart patterns defined by higher-order spatial correlations. In ordinary language, it asked what machinery a visual system needs before a complicated difference becomes visible.

It is tempting to turn every early research topic into prophecy. Careers are usually untidier than that. Still, the continuity is striking. Bubna-Litic began by studying the hidden structure that lets a human recognize a pattern. He now builds around the hidden structure that lets software act on a network. The nouns changed. The underlying question kept its desk.

2009Computational vision research at ANU
10 yearsApproximate tenure at Quantium
$250A runaway agent test worth documenting

There were useful side roads. He studied architecture at ETH Zurich, another discipline preoccupied with structure, constraints and the awkward behavior of reality. He co-founded Escape Room SA. In 2014, he reached the final of a GeoNext and GoGet hackathon with a data-visualization applet and presented it at the conference. A mathematician making a visual tool from transport data is not a dramatic reinvention. It is evidence of an appetite for problems that eventually have to leave the notebook.

02 / ProductionThe decade of ugly edges

Quantium supplied ten years of those departures. Bubna-Litic moved through analyst, senior data scientist and lead data scientist roles before becoming a vice president of data science. His work there included retail forecasting and an LLM-powered tool for generating SQL. The corporate progression matters, but his comments about the job are more revealing than the promotions.

In 2020, when “production AI” usually meant machine learning rather than a chatbot wearing a necktie, he explained how his team deployed models. Docker and Kubernetes helped keep training and serving environments stable. Models were tested on out-of-sample holdouts. Features and errors were checked against development behavior. Outputs were watched through both metrics and visualizations. Training and serving reused as much code as possible.

“We choose simple models and features over more complex ones.”Anton Bubna-Litic on production machine learning

This is a useful sentence because complexity has excellent manners in a demonstration. It arrives dressed as sophistication. Then the upstream data changes, a package version shifts, or a pipeline produces a plausible but wrong number at two in the morning. A simple model may be less enchanting, but it is easier to diagnose and correct. Bubna-Litic’s warning was sharper: machine-learning pipelines can fail silently.

He also emphasized the relationship between data scientists and engineers. Productionizing a statistically difficult model, in his telling, was often more engineering than statistics. The work improved when the two groups stayed close, learned from each other and enjoyed being a team. Reliability was technical, but it was also social. Someone had to feel comfortable saying that the elegant answer had become a maintenance problem.

A practical operating loop
01

Prefer the legible

Choose models and features the team can inspect when reality disagrees.

02

Test the seam

Keep training and serving close, use holdouts and check upstream behavior.

03

Watch what ships

Monitor outputs with numbers and pictures, because quiet failures still count.

03 / The guildA $250 receipt from the future

At NetBox Labs, the models gained initiative. Bubna-Litic describes an AI team constantly experimenting with agentic development: software that can plan, use tools and perform sequences of work. He runs a weekly augmented engineering guild where colleagues share ideas, try new tools and report what happened. The ritual is important. An experiment held privately is a curiosity. An experiment made legible to colleagues can become an operating advantage.

Two results from that guild are wonderfully specific. Directory-specific prompts performed ten times better than one large configuration file. A runaway test-driven development session burned $250 in tokens. Both entered the shared playbook. One was a performance gain. The other was a small invoice from chaos. Together they reveal more about an engineering culture than a dozen declarations about innovation.

“Failures are first-class findings.”Anton Bubna-Litic on the NetBox Labs experiment log

Failure is easy to praise in the abstract and embarrassing in the expense report. Treating it as a first-class finding means preserving the conditions, cost and lesson well enough that the next person does not have to purchase the same knowledge. It also prevents a team from remembering only its victories, which is how folklore replaces engineering.

The approach is not timid. Bubna-Litic has said his team is pushing at what is possible, and the cadence he describes is active: new ideas nearly every day, tools tested as they appear, an internal agent platform used across the company. The restraint lies in what follows enthusiasm. Results are compared. Errors are recorded. A tool does not become useful merely because it can finish a demo while everyone is watching.

NetBox Labs diagram linking the network, observability, automation and operations in a continuous lifecycle
The loop behind the work: network data moves through observability, automation and operations. An agent needs context at every turn, not merely confidence.

04 / ContextGiving the agent a map

NetBox is unusually suited to the next part of the story. It models devices, addresses, racks, circuits and the intended state of infrastructure. That structured context is valuable to people managing complex networks. It is also the sort of context an agent needs before touching them. A language model can produce a polished answer without knowing whether a cable exists. Infrastructure is less forgiving of eloquence.

In June 2026, NetBox Labs launched a broader infrastructure intelligence platform that included a platform MCP server and a library of agent skills. MCP is a standard way for AI applications to connect with tools and data. Bubna-Litic singled out the server with particular pride. It allows agents to interact with the platform through a codemode sandbox, an environment in which an agent can write and execute code within boundaries. His team’s evaluations, he said, found that agents using this approach outperformed agents limited to direct tools.

The detail matters because it turns a broad claim about intelligent infrastructure into an engineering choice. A sandbox can give an agent room to combine steps, transform data and inspect results without handing it an unrestricted kingdom. NetBox supplies the modeled world. The interface supplies controlled access. Evaluation supplies the reason to prefer one design over another.

Mathematics and research at the Australian National University.

Analyst to vice president of data science at Quantium.

Joined NetBox Labs as Founding AI Engineer.

Helped bring agent access and a codemode sandbox to the infrastructure intelligence platform.

05 / The useful futureOptimism with observability

Bubna-Litic’s career is not a conversion story in which a sober data scientist suddenly discovers AI. It is a story about the widening radius of a familiar practice. In computational vision, the question was whether a distinction could be perceived. In retail analytics, it was whether a prediction could survive production. In agentic infrastructure, it is whether software can act usefully inside a complicated, changing system.

Each stage rewards the same habits: formal curiosity, a taste for visible evidence and respect for the inconvenient edge case. Even the escape room fits, if only as a joke with good structural integrity. A puzzle is satisfying when the clues are sufficient, the constraints are real and the solution works outside the solver’s imagination.

The aspiration visible in his recent work is larger than developer productivity. NetBox Labs’ internal agent experiments are feeding future product work for customers, while the platform’s MCP server opens structured infrastructure to agents. If that direction succeeds, network teams will gain collaborators that can navigate operational context, not just discuss it. The promise is leverage across a company. The obligation is to make that leverage observable.

This is where the checklist earns its place beside the optimism. An agent may be new, but silent errors are old. A sandbox may feel futuristic, but boundaries are ancient wisdom. A tenfold prompt improvement is valuable because it was measured. A $250 mistake is useful because it was remembered. Bubna-Litic’s contribution is not to drain the wonder from AI. It is to give wonder a test environment, a monitoring panel and colleagues who meet next week to compare notes.

The future often arrives as theatre. His version comes with logs.