Imagine changing a column name in a data model. The edit takes a moment. Finding out what it breaks can take rather longer. Somewhere downstream, another query expects the old name. You send work to a warehouse, wait for a response, and learn what a compiler might have told you before you left your desk. The cloud has become an unusually elaborate spellchecker.
- SDF moved SQL validation and dependency analysis into local development.
- Classifiers added business meaning beyond basic SQL data types.
- dbt acquired the team in January 2025; its technology helped power Fusion.
- The transferable lesson: check the code before paying to run it.
SDF Labs built a business around shortening that conversation. Its multi-dialect SQL compiler could understand queries locally, inspect their dependencies, and flag errors before execution. It also offered a transformation framework and an analytical execution engine. The initials stood for Semantic Data Fabric; the practical proposition was simpler: give data developers earlier feedback. The original toolchain brought compilation and execution into the same development workflow.
For an analytics engineer, that could mean catching a missing column while editing, tracing where a field came from, or checking what a proposed change would affect. For the person paying the warehouse bill, it meant an opportunity to stop using remote compute for questions that could be answered locally. The distinction matters: making a development loop cheaper does not make every production query cheaper.
A family business with a compiler habit
The company began in 2022 with four co-founders: Lukas Schulte, Wolfram Schulte, Michael Levin, and Elias DeFaria. Lukas was CEO; his father Wolfram was CTO. Levin had worked at Microsoft and Meta. DeFaria and Lukas had worked at PiñataFarms, a consumer video startup. It was an unusual mixture of database research and software built for people impatient to see something happen. Their June 2024 public debut came after roughly two years of development.
Wolfram supplied a particularly relevant apprenticeship. At Meta, he had worked on tracking personally identifiable information through pipelines spanning more than a million tables, according to dbt founder Tristan Handy's account of the acquisition. At that scale, a diagram of which table feeds which other table is only the beginning. A privacy question can concern one field, copied, renamed, or transformed across many models.
That background helps explain why SDF was interested in meaning as well as punctuation. SQL can tell you that two columns contain integers. It cannot, by that fact alone, tell you that one integer is a customer identifier and the other is a marketing event identifier. Human beings know these are different things. A conventional type check may be perfectly satisfied when they are joined.

Two integers walk into a flower shop
SDF's Mom’s Flower Shop tutorial makes the problem wonderfully ordinary. A model joins in-app events to marketing events using an event identifier. The two identifiers describe different kinds of events. The SQL can look respectable while the business logic is wrong. SDF's example annotates the source fields with different classifiers, propagates those labels, and adds a check that catches the mixture.
Classifiers were an extra layer of metadata on columns and tables. They could describe sensitive information, financial concepts, identifiers, or retention requirements. The classification system allowed developers to define how labels should survive a transformation. A copied phone number and a count of distinct phone numbers need different treatment. A label that follows everything indiscriminately soon becomes another thing to ignore.
Column-level lineage supplied the map. SDF's lineage documentation distinguishes a direct copy, a modification, and a filter dependency. Those distinctions help explain how an output field depends on its inputs, rather than merely drawing an arrow between two tables. In its model deprecation tutorial, the developer checks downstream references before retiring an old model. The apparently small maintenance task becomes a question the system can help answer.
A useful detail sits in the checks tutorial: a privacy check initially passes because the relevant classifier has not been added. Once the label is attached, the check fails. There is a lesson here for anyone buying governance software. A passing check only means something if the underlying description is complete. Metadata deserves the same suspicion as a spreadsheet with suspiciously tidy totals.
EVENT.inappEVENT.marketingThe rival became the engine
SDF occupied the transformation and developer-tooling layer of the data stack. It was aimed at people preparing warehouse data for analysis, rather than selling a dashboard to everyone in the business. Its approach combined three jobs: parsing SQL syntax, compiling it with schema knowledge, and executing locally through Apache DataFusion. dbt's technical explanation described all three levels. Support for multiple dialects mattered because SQL changes its accent from warehouse to warehouse.
There were other routes to SQL understanding. SQLMesh uses SQLGlot for semantic analysis and column-level lineage. Warehouse-native validation was another alternative. SDF's distinction lay in its particular combination of Rust implementation, static analysis, classification, and local execution - not in being the only company that noticed SQL needed to be understood.
Then came the distribution decision. On January 14, 2025, dbt Labs acquired SDF Labs, bringing the team into dbt. A startup building tools around an established workflow was now helping rebuild that workflow's engine. The acquisition turned SDF's technical work into a foundation for dbt's developer experience.
Integration exposed a practical obstacle: SDF was written in Rust, while the existing dbt engine was written in Python. dbt described a five-month rewrite that incorporated SDF's technology. Fusion entered public beta on May 28, 2025, alongside dbt's VS Code extension. This was more involved than putting a new badge on an existing compiler.
“SDF provides a software toolset for faster developer feedback as you’re writing data models.”Lukas Schulte, June 2024
Nine million dollars, and a different answer to “free”
SDF announced $9 million in seed funding in June 2024. RTP Global says it led the seed round; other reported investors included Two Sigma Ventures and Founders’ Co-op. The money backed development of the company. Funding is not revenue, and it was not the sale price: financial terms of the acquisition were not disclosed.
The commercial evidence was more modest and more useful: at launch, GeekWire reported paying customers including ClassDojo, Obie, and Linqto. SDF's public-beta announcement also thanked engineers at Patreon, Deel, and Cybersyn, among others. Those acknowledgments establish collaboration; they do not make every named organization a paying customer. The business sold developer software to data teams, with public libraries and documentation surrounding its core technology.
After the acquisition, the business model belonged to dbt's wider mix of free local tools and commercial platform features. In May 2025, Handy explained a change of mind: SDF's compiler source was proprietary when acquired, and dbt initially intended to keep that strategy. It instead chose a mixture of proprietary, open-source, and Elastic License 2.0 components. The source-available license allowed local use while restricting a competing hosted service.
The arrangement changed again. On June 1, 2026, dbt announced the first alpha of dbt Core v2.0, sharing Fusion's Rust foundation under Apache 2.0. Fusion continued to add capabilities above that foundation. By July, dbt explicitly distinguished Fusion's SQL comprehension from the open-source Core baseline. For users, “free,” “source-available,” and “open source” remained three different questions, even when the engines shared ancestry.

Borrow the sequence, test the assumptions
The most useful thing to copy is the order of work. First describe the source schemas accurately. Then validate code, inspect dependencies, and apply checks before running transformations against expensive infrastructure. SDF's local compilation guide requires column names and data types for source tables. A compiler reasoning from stale schemas can give an impressively prompt answer to the wrong question.
Next, distinguish a property of the code from a property of the data. SDF's data tests concern values in tables; static checks concern what can be inferred from models and metadata. A valid query can still process duplicate records or encode the wrong definition of revenue. Local validation reduces one class of surprises. Production testing and business review still have work to do.
Finally, keep the workflow connected. The documented Dagster integration exposed SDF models as assets with compile-time and runtime metadata, tying authoring to orchestration. The useful adoption question is whether the engine understands your actual dialect, functions, macros, and source definitions. A tidy demonstration is a beginning; a representative project is the examination.
SDF's durable idea is that a developer should learn about a mistake while it is still easy to fix. That sounds almost too polite to build a company around. Yet moving knowledge earlier changes who must wait, what must run, and how much a mistake has already cost. Sometimes the consequential innovation in a data warehouse is the query you never had to send.
Follow the code, meet the people
The original SDF site now leads into dbt. These doors lead to the technology and the people who built it.