Eighty million legal documents are an awkward thing to ask a machine to read. Activeloop says one of its earliest customers had exactly that problem. The company helped cut a model-training job from two months to a week. The striking part is not that a model got clever. It is that the documents became usable. Before an algorithm can find an answer, someone must collect the files, represent them consistently, move them to compute, and keep track of which version was used.
- Activeloop sells infrastructure for storing, searching and training on mixed AI data: text, images, video, audio and embeddings.
- Deep Lake grew from an open-source dataset format into a multimodal database; newer products address agent state and shared memory.
- Its best public evidence comes from specific workflows at Bayer Radiology, Matterport, Flagship Pioneering and IntelinAir, rather than a claimed universal speedup.
Founder Davit Buniatyan had seen the difficulty at Princeton, where neuroscience research produced data at a scale that made storage and streaming part of the scientific problem. The company began in Y Combinator's summer 2018 batch as Snark AI. It built custom systems for legal documents and agricultural imagery. Buniatyan later described the lesson from those projects and hundreds of customer conversations: excellent databases existed for rows and analytical queries, but machine learning teams still assembled their own contraptions for images, labels and other unruly material.
The first answer was a chunked dataset format called Hub, released as open source in 2020. As version control, querying, visualization and data loaders accumulated, Hub became Deep Lake. The company even admitted a less noble reason for the rename: everyone seemed to have a “hub.” The joke is useful. A format had become a system, and the old name no longer told the truth.
When the dataset is the experiment
A conventional application database is comfortable with records. A machine learning dataset might contain a room scan, a label attached to a particular frame, an embedding generated from that frame, and the earlier version of the label that proved wrong. Deep Lake's pitch is to put those related objects in one versioned place, then query or stream them without a hand-built conveyor belt for each model. The open-source library gives developers a way in; managed and enterprise offerings sell the operational layer around it.

Matterport offers a neat demonstration. Its machine learning manager Alan Dolhasz said Deep Lake let the team train on a different dataset by changing one line of code, a switch that previously took at least a day. The point is not the line. It is the ability to vary training data while leaving the rest of the experiment intact. That makes the comparison between models less about who rebuilt the pipeline most recently.
“With Deep Lake, it's literally changing one line and we can train on a completely different dataset.”Alan Dolhasz / Matterport machine learning development
Earthshot Labs found another use for the same version history. During a tree-segmentation project, a researcher saw inexplicable false positives after retraining. She inspected dataset versions and found faulty masks, then rolled back to an intact commit while investigating. A database had become a record of why a model changed. It is a mundane achievement until the alternative is redoing field work or blaming the model for bad labels.
The bill arrives before the breakthrough
IntelinAir's crop imagery made the scale plain. Its Activeloop case study describes roughly 1.5 petabytes of raw aerial data in 2020, assembled from repeated flights and several kinds of sensor readings. The published result claims 50% lower compute and storage costs, 30% less storage and 12% higher accuracy against a baseline. Those are customer-case figures, not a promise that every dataset will behave the same way. They do show what the buyer was paying for: less time and money spent preparing the next training run.
Bayer Radiology's problem was different in kind but similar in shape. Medical images, records and laboratory data did not naturally arrive as one tidy AI dataset. Activeloop says Deep Lake was integrated into Bayer's cloud in two weeks. The case study describes less preparation work and a way to query biomedical material, including scans, in natural language. The deployment in Bayer's own cloud matters as much as the interface: sensitive data cannot simply be thrown into the first convenient service.
Flagship Pioneering tested a more search-centered feature, Deep Memory, on biological research questions. Its case study reports better retrieval than baseline vector or lexical search, with a chart showing recall at ten results rising from 65.3% to 79.7% in one comparison. It is a useful reminder that “find the right document” is a measurable engineering task. A persuasive chatbot with the wrong paper underneath it is still wrong.
A lake with a Postgres door
Activeloop's $11 million Series A in March 2024, from Streamlined Ventures, Y Combinator, Samsung Next, Alumni Ventures and Dispersion Capital, put its disclosed funding at about $20 million. The company framed the money around selling Deep Lake to large enterprises. The later product sequence shows a second idea taking shape: once a system can store and retrieve the raw material of AI, it can also hold what AI agents do while they work.
In 2025 Activeloop introduced L0, a service that reasons over mixed documents and returns answers with evidence. In 2026 it described Deeplake as a serverless, PostgreSQL-compatible database for short-lived agents. The wording needs care. Its engineers say PostgreSQL supplies the familiar connection protocol, parsing and authentication, while Deep Lake stores the data and DuckDB executes queries. Familiar clients can connect, but some PostgreSQL behavior differs. The company explicitly names features such as SELECT FOR UPDATE, advisory locks and LISTEN/NOTIFY as outside or different from ordinary Postgres semantics.
The parts have different jobs. Calling the whole arrangement “Postgres” would hide the interesting engineering.
For an agent that spins up, searches a pile of PDFs, writes interim state and then goes quiet, the architecture promises quick provisioning and no always-on idle database. But the same design is a poor assumption for a workload that relies on PostgreSQL's full transaction behavior. Activeloop's practical market is therefore specific: teams working with bursty agents and multimodal data, especially when they also need the originals for training or audit. Standard Postgres and specialized vector stores remain sensible alternatives when those extra jobs are absent.
What the next agent inherits
Hivemind extends the thought from data to experience. It captures coding-agent sessions and turns their traces into skills a team can reuse. A June 2026 example with ScrapeGraphAI showed a short skill enriched with live web research while keeping each addition tied to a source. On its main site, Activeloop places Hivemind between Deeplake, the data layer, and Refinery, a system for testing software improvements. The proposed loop is easy to understand: observe a run, remember its useful result, try an improvement, verify it, then carry that learning forward.
That is also a description of what a careful human team already does. The difference is whether the work is captured in a form that survives the next session. The company has moved from helping a scientist feed a model to helping an agent inherit another agent's work. It has also widened its competition: cloud object stores, vector databases, managed Postgres and agent-memory tools can each do pieces of the job. Activeloop must make the combined workflow simpler enough to justify adopting a specialized stack.
The transferable lesson is less grand than “build a database for AI.” Start with the waste. Measure how long it takes to prepare a dataset, switch a version, recover a bad label, or reconstruct what an agent learned. Keep the original material next to its derived index. Make every experiment identifiable. And when a customer says a task went from a day to a line of code, ask what disappeared in between. That vanished work is often the product.