A data lake sounds restful until somebody asks what happened to yesterday’s numbers. Then the scene changes. Files must arrive, updates must land in the right order, small files must be packed, catalog entries must agree, and a query that worked on Monday must still work on Friday. Onehouse has built a company around that less photogenic side of the lakehouse: the work required to make open data useful every day.
The short version
- Onehouse runs managed lakehouse infrastructure in a customer’s cloud account, with data stored in open table formats.
- Its core jobs are ingestion, incremental processing, table maintenance, catalog synchronization, and, more recently, query serving.
- Its founder, Vinoth Chandar, created Apache Hudi at Uber; the company has raised $68 million through its 2024 Series B.
- For a buyer, the test is practical: fresher data and lower operating effort on the buyer’s own workload, after all service costs.
The trouble began at Uber
Chandar built Apache Hudi at Uber to solve a problem that traditional data lakes handled awkwardly: data changes. A ride, transaction, or account record can be corrected after it first appears. Dumping an immutable batch once a day leaves an analytics system behind the business it is meant to describe. Hudi brought transactional updates and incremental processing to lake storage. It was a technical answer to a distinctly human complaint: “Why is the dashboard still waiting?”
Onehouse, founded in 2021 and publicly introduced in 2022, is the commercial sequel. Its premise is that a clever open-source format does not, by itself, operate a production data platform. Engineers still have to connect sources, tune tables, monitor freshness, and make the data legible to different query systems. Onehouse sells a managed service for those chores. The data lives in the customer’s cloud environment; the company says its platform works with Apache Hudi, Apache Iceberg, and Delta Lake rather than requiring one proprietary storage format.

The expensive part is the boring part
The first version of Onehouse focused on the managed foundation: ingest data, transform it incrementally, and keep tables fit for querying. Today its catalogue is wider. OneFlow handles ingestion. Table Optimizer automates tasks such as clustering and compaction. LakeView offers Hudi table observability. OneSync publishes table metadata to multiple catalogs, including Snowflake and Databricks Unity Catalog. Quanton runs SQL and Spark ETL jobs on a managed compute runtime. Lakegres adds a Postgres-compatible query endpoint for Hudi and Iceberg tables, aimed at fast lookups as well as analytical work.
That sequence explains the business model. Onehouse sells modular managed software and expert support to enterprise data teams. A free observability product draws Hudi users into the orbit; paid services take on more of the operating burden. The costs a buyer can measure are more useful anyway: cloud storage, compute, software fees, engineering hours, and the price of a stale or broken pipeline.
Consider the table itself. Incoming records do not arrive in the neat blocks that query engines prefer. Small files accumulate; changes demand reconciliation; a scan that once finished quickly begins to drag. Table maintenance is scarcely the sort of work people boast about at conferences, which is precisely why automating it can matter. Onehouse’s product claim is that a team can preserve the economics and portability of open storage without assembling a small guild to keep it healthy.

A customer’s first broken thing
At Apna, the Indian jobs platform, the original stack centered on BigQuery and batch updates that ranged from hourly to daily. That cadence was a poor fit for a business using data to improve job matching. Apna moved toward streaming with Confluent Cloud and a Hudi lakehouse managed by Onehouse. Its engineers used append-only bronze tables so downstream data could be rebuilt after errors, while Onehouse handled updates to the lakehouse tables. The customer account describes fresher data, fewer failure points, and a way to keep BigQuery SQL while adding other engines.
This is a useful distinction in a market fond of declaring revolutions. Apna did not have to forbid its analysts from using familiar SQL. The architecture changed underneath them. In the published account, the old arrangement failed first on freshness, cost, and maintenance; the decision to stream data followed from those bottlenecks. A reader contemplating the same move can copy the diagnostic, even if the vendors differ: time one critical record from source to query, count the handoffs, and measure the cost of each copy.
NOW Insurance offers another angle. In a Onehouse customer webinar, it reported saving more than 40 engineering hours a month on pipeline building and maintenance, as well as cutting Airflow operating expense by half. Those are customer-reported figures, not a universal discount coupon. Olameter’s problem was stranger: years of meter-reading XML had piled up because an earlier custom .NET approach was too slow to process it economically. Onehouse says its team built a custom XML ingestion path in two weeks, making that data queryable for predictive maintenance. In both cases, the obstacle was not a shortage of ambition. It was the stubborn shape and movement of data.
“We generally prefer open source vendors to avoid lock-in.”Subham Todi, lead data engineer at Apna, in a published customer account
Open, with a bill attached
Onehouse occupies an awkward and interesting market position. Databricks and Snowflake offer broad managed platforms. AWS EMR and similar services let teams run processing engines. A sufficiently staffed company can assemble Hudi or Iceberg, Spark, a catalog, ingestion tools, and monitoring itself. Onehouse’s pitch is narrower than owning every query and broader than a single ETL tool: operate the shared lakehouse layer so those other systems can work from open data.
The strongest expression of that pitch is interoperability. Apache XTable, which Onehouse developed with partners including Microsoft and later donated to the Apache Software Foundation, addresses translation among Hudi, Iceberg, and Delta table metadata. OneSync tackles a second problem: catalogs. An open table is less liberating if every engine expects a different directory of tables and permissions. Onehouse’s answer is to keep those directories in step. It is a practical idea, though a proof of concept should include schema changes, deletes, access rules, and failures, not just a clean demo table.
In 2025 the company widened its remit from managed storage toward compute. Quanton promises better ETL price performance by doing less redundant work in recurring jobs; the numbers on Onehouse’s product pages are benchmarks and guarantees framed by particular workloads. In 2026, Lakegres targeted another weakness: quick, selective questions against data kept in the lakehouse. Its Postgres-compatible endpoint is designed to let existing clients query Hudi and Iceberg tables without first copying them into a separate serving store. The company’s newest AI training pipelines prepare and version datasets, then hand training jobs to Baseten, Fireworks, or Together AI.
That expansion tells us what changed Onehouse’s mind. As Chandar described in a 2025 review, open formats had won acceptance, but teams still struggled with unpredictable costs and operational complexity. Onehouse chose to take responsibility for more of the system: processing, engines, and serving, not merely the file layer. The advantage is fewer seams for a customer to stitch together. The risk is familiar to any platform buyer: each additional service must earn its place on a real workload, rather than on the elegance of an architecture diagram.
The test worth copying
For a data team, the Onehouse question is not whether “open” sounds attractive. Choose one expensive, troublesome pipeline. Record its freshness, failure rate, cloud bill, and engineer time. Rebuild it with data in an open table format; connect the query engines people actually use; then run schema changes, late records, and a recovery drill. If the numbers improve after the Onehouse fee and migration work are counted, the service has done something concrete. If a small, stable workload already runs cheaply, or an organization depends on tightly coupled features of one warehouse, the case may be weaker.
A lakehouse can be a magnificent piece of architecture. It can also be a place where somebody must take out the rubbish. Onehouse’s wager is that the latter job deserves a product, a budget, and perhaps a little more respect.