There is a specific request that shows up in Shofo's pitch, and it sticks with you: one hundred thousand hours of cooking videos where someone is holding a pan, with reasoning annotations on top. It is oddly precise. It is also exactly the kind of thing an AI lab now needs, cannot easily get, and will pay real money to have delivered by Friday.
Shofo, a startup in Y Combinator's Winter 2026 batch, exists to fill that request. The four-person company, based in San Francisco, describes itself in five words: "The World's Largest Video Library." The longer version is that Shofo crawls billions of short-form videos from across the internet, cleans and labels them, and sells the resulting datasets to the labs building the next generation of AI models. Its internal shorthand is blunter, and it is the phrase the company keeps coming back to: Common Crawl for video.
That reference is worth unpacking, because it explains the whole bet. Common Crawl is the open repository of web pages that quietly fed a decade of language models. Nobody outside the field talks about it, and almost everyone inside the field has used it. It is infrastructure. Shofo is wagering that video needs the same thing, that no one has built it yet, and that the endless scroll of TikTok is actually the largest untapped training set on the planet.
01 / The problemThe six-month bottleneck
To understand why Shofo has customers, you have to understand what its customers were doing before. A lab that wants to train a model to understand video - to recognize objects, track actions, reason about what is happening on screen - needs enormous volumes of footage that has been labeled for exactly those tasks. Historically that meant one of two things: license a pile of data at a price measured in the millions, or assign a team to collect and annotate it in-house, a process that runs for months.
Neither option is fun, and both are slow. The data is the constraint, not the ambition. Shofo's read on the market is that generating or structuring the right data has quietly become more valuable than building the model that consumes it. Sell the shovels, in other words, to the people digging for gold.
The customers here are narrow and specific: research labs and companies training or fine-tuning multimodal models, the kind of teams for whom a missing dataset is the difference between shipping this quarter and shipping next year. These are not casual buyers. They know precisely what they need, they can tell a good label from a lazy one, and they are the ones who have been complaining loudest about how long video takes to assemble. When your buyer is that sophisticated, a mediocre product gets found out fast. That cuts both ways - it is a hard audience to fool, and a valuable one to satisfy.
02 / The productCrawl, sanitize, label, deliver
What Shofo actually operates is a pipeline. A lab arrives with a request in plain language - the cooking-videos-with-a-pan example is the company's own - and Shofo's agents search the index, pull the matching subset, run it through labeling, and hand back a dataset in standard formats. The parts break down roughly like this.
Collection
A fleet of crawlers pulls short-form video from TikTok, Instagram Reels, YouTube Shorts and the wider web, supplemented by private data partnerships.
Sanitization
NSFW screening, quality filtering, and perceptual hashing to remove duplicates - the unglamorous cleanup that makes the rest usable.
Labeling
Fast object detection paired with vision-language reasoning annotations, layering activity recognition and segmentation onto each clip.
Delivery
Packaged in formats labs already use - WebDataset, Hugging Face, or raw archives - ready to feed straight into training.
The individual tools here are not secret. Object detectors and vision-language models are available to anyone. What is hard, and what Shofo is really selling, is the combination running at scale, on video, without falling over - and the index that already exists on the other end of the request. A lab asking a competitor for a niche dataset starts a collection project. A lab asking Shofo is running a query against something that is already there. That difference - a search versus a scavenger hunt - is the entire pitch, and it is why the company keeps stressing the word "index" over the word "labeling."
The business model follows directly from that. Shofo sells datasets, not seats or subscriptions. Price scales with how much work each video requires: a set with plain object labels costs a fraction of a set carrying full reasoning annotations, activity recognition and segmentation stacked on top. A lab placing a large recurring order - hundreds of thousands or millions of labeled clips a month - is where the economics get interesting, because the marginal cost of pulling more from an index you already run is low, while the value to the buyer stays high. It is a volume business dressed as a bespoke one.
03 / The originThe pivot that read its own metrics
Shofo did not start as Shofo. The founders - a group connected largely through UC Santa Barbara - previously built Correkt, an AI search engine for multimodal content that reached more than 40,000 users. Along the way they built crawling infrastructure to feed it. At some point they noticed the plumbing was worth more than the product it was plumbing for. They kept the pipes and rebuilt the company around them.
The current team is small and technical: Braiden Dishman as chief executive, who did a stint at AWS before this; Alexzendor Misra as chief technology officer, who ran Correkt; and Andre Braga leading AI, with a statistics and data-science background. Four people, one focused bet. It is the kind of team you assemble by accident and only recognize in hindsight - people who met building something else and discovered, mid-build, what they were actually good at.
The expertise that matters here is not glamorous. It is knowing how to keep a crawler alive against a platform that would prefer it dead, how to deduplicate at a scale where near-copies of the same clip number in the millions, and how to attach labels that a picky lab will actually trust. This is systems work and data-plumbing work, closer to logistics than to research. The founders spent two years learning it on someone else's problem before turning it into their own.
04 / The moatWhy this is hard to copy
A skeptic's first question is fair: if the tools are off the shelf, what stops a well-funded rival from cloning this? Shofo's answer is not a clever algorithm. It is time and access. Crawling major platforms at scale means surviving anti-scraping defenses that are aggressive and constantly changing. That is years of unglamorous infrastructure work, plus private data partnerships that give the company inventory a newcomer simply cannot reach. The head start is the product.
05 / The go-to-marketGive it away, then sell it
Here is a move that looks backwards until it doesn't. Before selling a single enterprise contract, Shofo published a sample dataset of roughly 58,000 videos on Hugging Face and let the research community download it - more than 25,000 times. For a company whose product is data, giving data away sounds like a mistake. It is closer to a handshake. You do not win a lab's trust by being a stranger. You win it by showing your labels are good, in public, for free.
06 / The marketStanding next to the labeling giants
Shofo enters a field that already has incumbents - Scale AI, Labelbox, Surge and others that built businesses labeling data customers already hold. Shofo's angle is different in two ways: it owns a large pre-built index of public video before any customer shows up, and it is pointed specifically at video rather than at everything. Whether that focus becomes a durable advantage or an incumbent's next feature is the open question. The company is early, the funding is modest, and the thesis is unproven at scale. But the bottleneck it is attacking is real, and labs are the ones complaining about it.
Will this matter in ten years? That depends on whether video becomes as central to AI as text has been. If it does, someone will own the layer underneath it. Shofo is a small, specific attempt to be that someone.
Find Shofo
- shofo.ai - official website
- Y Combinator company page
- LinkedIn - Shofo
- X / Twitter - @shofoai
- Hugging Face - Shofo datasets
- Braiden Dishman - Co-Founder & CEO
- Fondo - "Shofo Launches: Common Crawl for Videos"
Note: video interviews and a product demo were not publicly available at the time of writing. Check shofo.ai and @shofoai for the latest.