OnDeck AI Wants You to Just Ask the Footage
A Vancouver startup replaced months of computer-vision labeling with a sentence. Its Perception-0 model finds anything in video - no training data. It began by counting fish and now watches autonomous naval fleets.
The pitch is almost suspiciously short. Point a camera at anything - a harbor, a factory floor, the deck of a boat pitching in swell - and instead of building a model to recognize what you care about, you type a sentence. "Find every time someone crosses the yellow line." "Flag the vessel that stops and waits." OnDeck AI's model reads the footage and answers. No labeled dataset, no six-month engineering project, no data scientist on retainer.
That is the whole idea behind OnDeck AI, a Vancouver company in Y Combinator's Summer 2025 batch. Its flagship model is called Perception-0, and the tagline the founders use is blunt: analyze any footage, without training a model. For anyone who has actually tried to ship a computer-vision system, that sentence lands somewhere between relief and disbelief.
01 / THE PROBLEMWhy watching video is still mostly a human job
Cameras are everywhere. Understanding what they record is not. Traditional computer vision works by memorization: you collect thousands of labeled examples of the thing you want to detect, train a narrow model, and hope the lighting, the angle, and the weather never change too much. Change the camera and you often start over. Ask it to notice a behavior - not "a person" but "a person loitering" - and the whole approach strains.
So most footage goes unwatched, or gets watched by a tired human scrubbing through hours of it for the ten seconds that matter. OnDeck's bet is that vision language models - the same family of AI that can describe an image in words - can reason about video the way a person does, from a handful of examples instead of a mountain of them.
The distinction matters more than it sounds. A memorizing model can tell you a red truck is a red truck because it has seen ten thousand red trucks. It cannot easily tell you that the red truck is idling in a place it should not be, or that it arrived, waited, and left twice in an hour. Those are questions about context and sequence, and they are the questions most operators actually care about. Dungate frames the shift as models that reason from minimal images, closer to human cognition than to a lookup table.
02 / THE ORIGINIt started with fish
OnDeck did not begin in defense. Its legal name is still OnDeck Fisheries AI Inc., a fossil of where it started. The founders - Alexander Dungate and Sepand Dyanatkar, friends since grade 6 - first pointed the technology at fishing boats. Working with First Nations longline fishers around Vancouver and lake fisheries in northern Manitoba, they used AI to do the tedious, regulated work of electronic monitoring: counting halibut, blackcod, and rockfish, and identifying species from deck cameras.
Fisheries turned out to be an excellent teacher. The cameras were bad, the lighting worse, the events rare, and labeled data almost nonexistent - exactly the conditions that break conventional computer vision. Solving the ugly version of the problem first gave the team a product that generalized. The same model that separates a rockfish from a blackcod can separate one behavior from another on the deck of an autonomous vessel.
The fisheries work also grounded the company in something most AI startups skip: consequences. Electronic monitoring exists because regulators want to know what a boat actually caught and threw back, and getting that count wrong has legal and ecological weight. Dungate has been candid about how far that work still has to go - regulatory limits, financial constraints, and technical hurdles like the context window that forces long footage to be split into pieces. He has estimated full deployment on fishing vessels could take three to five years with enough investment, and that early funding covered only a fraction of what the fisheries mission needs. That honesty about limits is unusual, and it reads as a company that has met the hard parts rather than skipped past them.
03 / THE PRODUCTFour verbs: search, understand, report, customize
Perception-0 is sold as infrastructure rather than a dashboard. The workflow reduces to four moves, and most customers live in the first two.
Search
Type what you're looking for in plain language and scan across large footage datasets.
Understand
Grasp qualitative behaviors and multi-step event sequences, not just single objects.
Report
Turn what it finds into structured reports for the people who need to act.
Customize
Adapt agentic workflows to a customer's specific cameras, footage, and rules.
The design choice worth noticing is what OnDeck chose not to build. There is no army of labelers, no bespoke model per camera, no configuration marathon. The interface is the question. That is a harder thing to engineer than it looks, and it is the part competitors using off-the-shelf models struggle to match on hours-long footage.
Perception-0 is also pitched as a self-adapting, agentic system rather than a fixed classifier. In practice that means the same model can be pointed at a port one day and an offshore rig the next without a retraining cycle in between. For a buyer, the appeal is less about any single accuracy number and more about time: the months of computer-vision development that OnDeck says it removes are months a customer would otherwise spend before seeing a first useful result.
04 / THE CUSTOMERSFrom First Nations fishers to a navy
The customer list reads like two different companies stapled together, which is roughly the point. On one end: NGOs and fisheries programs. On the other: international defense and security agencies, robotics research teams, autonomous-asset operators, port monitors, and offshore oil and gas platforms.
The headline account is Singapore's Ministry of Defence, which ran a pilot using OnDeck to find targets and behaviors in footage streaming from a fleet of autonomous vessels. That a defense ministry handed a months-old startup a paid pilot says something specific: the team proved it on real footage, not slides.
What ties the two ends of the customer list together is the shape of the problem, not the industry. A fisheries regulator and a naval command both have more footage than eyes, both need to catch rare events buried in long stretches of nothing, and both operate cameras that were never standardized for AI. OnDeck sells to whoever has that shape, which is why a company that started with longline boats can credibly quote robotics researchers, security firms, and offshore energy operators on the same page. The common thread is the question the customer wants to ask, not the vertical they sit in.
05 / THE DIFFERENCEReasoning instead of memorizing
Set against the field, OnDeck's position is narrow and deliberate. Below is the rough shape of the alternatives a buyer weighs.
| Approach | What it needs | Where it strains |
|---|---|---|
| Traditional computer vision | Thousands of labels, a model per task | New cameras, behaviors, rare events |
| General-purpose VLMs | Prompting, but no video tooling | Long clips; forgets earlier context |
| OnDeck Perception-0 | A sentence describing the target | Purpose-built for long-video search |
The differentiator is not that OnDeck invented vision language models - it did not. It is that the team built the plumbing around them for long, messy, real-world video and pointed it at customers whose footage is genuinely hard. The published work backs the claim: the company has a NeurIPS workshop paper arguing its VLM methods beat traditional computer vision on niche tasks.
06 / THE TEAMA biologist and a rocket engineer walk into YC
The founding pair is an unusual mix. Alexander Dungate, the CEO, holds degrees in both computer science and biology and is a National Geographic Explorer - the fieldwork background that explains the fisheries roots. Sepand Dyanatkar, the CTO, brings a Cambridge machine-learning master's and stints at the European Space Agency, Amazon, and swarm-robotics labs. One founder knows what it's like to stand on a boat; the other knows how to make many machines coordinate. The product sits at that intersection: video treated less as pixels to classify and more as a scene to understand.
The company is small - six to seven people - and backed by a notable roster: Y Combinator, National Geographic, NVIDIA, and Eric Schmidt's family foundation. It is SOC 2 compliant, the kind of unglamorous checkbox that separates a demo from a vendor a government will actually pay.
There is a cultural throughline in how the founders talk about the work. Dungate's stated motto - stay stoked, never lose the passion and the energy - sounds like surf talk, and given the ocean roots, maybe it is. But it also describes a specific bet on retention through mission: a six-person team selling to defense ministries and fisheries regulators at the same time needs a reason to hold together, and the reason here is the belief that making footage legible is worth doing across whichever industry needs it next.
07 / THE MARKETWhere it fits
OnDeck sells into a widening gap. The world is generating far more video than anyone can watch, and the buyers who feel that most acutely - defense, security, robotics, industrial monitoring - are the ones with the budget and the urgency. The business model is straightforward B2B SaaS: land a pilot on a hard problem, prove it on the customer's own footage, and grow into a recurring contract. From $0 to six-figure ARR since May 2025 suggests the motion works.
Whether OnDeck stays a specialist in long-video understanding or gets pulled into the broader defense-AI arms race is the open question. For now the company is doing the thing that tends to compound: solving one hard, specific problem for customers who cannot solve it themselves.