The file
2016 ● Founded in San Francisco2018 ● Capital One named as customer2020 ● $29.6m announced funding2023 ● Acquired by Databricks2016 ● Founded in San Francisco2018 ● Capital One named as customer2020 ● $29.6m announced funding2023 ● Acquired by Databricks

01 / Company Data governance

The Company That Tried to Put One Lock on Every Data Door

Okera sold a deceptively simple idea: write the rule for sensitive data once, then make it follow the data everywhere. Databricks bought the company in 2023, just as AI made the old permission problem larger and stranger.

Imagine an analyst at a bank asking a reasonable question about customers. The answer is somewhere in the company's data lake. So are account numbers, private addresses, and records the analyst has no business opening. A simple “yes” grants too much. A simple “no” kills the analysis. The interesting answer lies in between: this person may see these rows, with this field hidden, for this purpose. Okera built a company around making that answer possible at scale.

In brief
/ 04 points
  • Okera made software for discovering sensitive data, writing access policies, enforcing them, and auditing use.
  • Capital One was a named early customer; large enterprises were its market.
  • It raised $29.6 million in announced funding and sold to Databricks in 2023 for an undisclosed price.
  • The useful lesson: classify data and attach the rule before another dashboard or model starts using it.

The company began in 2016 with two founders who knew big data from the inside. Amandeep Khurana had seen the difficulty of granting appropriate access across large organizations. Nong Li had helped create Apache Parquet, a file format that made analytic data easier to store and move. Portability, it turned out, had an awkward sequel. Once data could travel freely, who would ensure its permissions kept up?

At its 2018 public launch, Okera announced a $12 million Series A, general availability of its Active Data Access Platform, and Capital One as a customer. The bank was also represented among the investors through Capital One Growth Ventures. That combination matters. A large financial institution was not merely applauding the idea from the sidelines; it had a direct reason to make data useful without turning access into a free-for-all.

“Okera helps enable us to provide teams with secure and appropriate access to the right data.”Devon Cavanagh, Capital One, at Okera's 2018 launch
The product / 01

Four parts, one difficult question

Okera's original product brochure divided the platform into a data access service and three catalog services: a schema registry, a policy engine, and an audit engine. In plain English, it kept an inventory of data, recorded what the data meant, decided who could see it, and logged what happened. The target was an enterprise using several storage systems and analysis tools at once, often spread across clouds and older infrastructure.

A simplified Okera access path
1 / ClassifyFind data and mark sensitive fields.
2 / DecideMatch a user and purpose to policy.
3 / EnforceFilter, mask or allow at query time; record the event.
One rule can be reused across tools only when the systems actually pass through the enforcement path.

Consider a postal code. A marketing analyst may need the first three digits to map regional demand but not the full address. A fraud investigator may need more. Okera's later Snowflake integration described the ability to filter, redact, or transform values during a query. That is a more useful vocabulary than the blunt permit-or-deny menus that often greet a data steward. It also requires careful mapping of identities, tags, and workflows; a wrong classification can make a beautifully written policy useless.

Okera sold this to the people caught between competing incentives. Analysts and data scientists wanted speed. Data owners wanted to know where their datasets went. Privacy and security teams wanted a clear rule and a record of access. The platform's pitch was that each group could work from the same policy layer, even when their favorite tools differed. In that sense, Okera's most important product was agreement: a rule everyone could point to when somebody asked why a field was visible.

The turn / 02

The lake had another problem: nobody knew what was in it

The first bottleneck was permission. Then came the inventory. In 2019, Okera previewed Spotlight for AWS, a dashboard that combined audit information from Amazon S3 with classification and usage analysis. It could show who was touching tagged data, which tools were busiest, and which enormous datasets were collecting dust. Nong Li described a customer with several petabytes they suspected were unused. The question was wonderfully unromantic: why pay to keep a mountain of data if nobody climbs it?

Okera Spotlight analytics dashboard showing user activity, applications and daily trends
381 active users, 34 touching tagged assets. Spotlight made the invisible traffic inside a data lake look rather like rush hour.

The dashboard also revealed a shift in the company's thinking. Access control alone tells you whether a rule fired. Usage intelligence tells you whether the rule makes sense in the real world. If an approved dataset is never opened, perhaps no one can find it. If a sensitive field is queried in an unexpected pattern, perhaps a policy needs attention. Okera began with the gate; Spotlight looked at the crowd moving through it.

$12mSeries A announced
May 2018
$15mSeries B announced
April 2020
2023Databricks acquisition
announced May 3

In April 2020, Okera announced a $15 million Series B led by ClearSky Security, bringing its stated total to $29.6 million. Nick Halsey joined as CEO; Li returned to the CTO role before later becoming CEO again. The announced use of proceeds was familiar startup arithmetic: engineering, sales, and marketing. What the money actually cost in dilution is not public, nor is the price Databricks eventually paid. The figures say how much capital was raised, not whether investors made a profit.

The market / 03

The rule followed the data

A single-purpose data lake product would have met an increasingly scattered customer. Okera followed the work into AWS, Snowflake, and existing catalogs. Its Collibra integration used classification attributes defined in Collibra to drive Okera access rules at query time. It joined Snowflake's Snowpark program and announced a Snowflake-specific SaaS edition in 2021. It also announced an Amazon EMR edition for data stored on S3. Each move addressed the same irritation: a privacy rule written in one place can become a second, slightly different rule somewhere else.

Okera was neither the only vendor asking this question nor the only way to answer it. Immuta and Privacera were competing in data access governance, while cloud platforms offered native controls. Okera's distinction was its insistence on policy across a mixed estate: multiple clouds, multiple stores, multiple tools. That is valuable when a company truly has such a mixture. When one platform owns the whole workflow, native permissions may be simpler. And no shared layer can protect a route it never sees. The dull work of identifying data, assigning owners, and connecting each access path remains essential.

By 2023, the problem had acquired a new customer: AI. Models and their builders wanted broad access to training and retrieval data. Databricks announced it was buying Okera that May and planned to fold its classification and policy technology into Unity Catalog. Public accounts described Okera's isolation technology, then in private preview, as another attraction. The acquisition terms were not disclosed. It is safer to describe the deal as a plan to deepen Databricks' governance capabilities than to claim every Okera feature arrived intact.

A permission is only useful if it survives the journey from the catalog to the actual query.The operational lesson behind Okera's design

There is a practical idea to copy here, even without buying a platform. Pick one sensitive dataset. Mark its risky fields. Write down who may see which version and why. Test that rule in each tool that can reach the data. Then inspect the audit trail to see what actually happened. The exercise often discovers a surprise before it discovers a breach: an analyst using an old copy, a service account with extravagant rights, or a supposedly popular dataset nobody touches.

Okera's story is an unusually tidy illustration of a messy principle. The data industry's favorite verbs are collect, move, and analyze. A serious enterprise eventually adds another: allow. Okera built its business in that little word. The company disappeared into Databricks, but the question it posed has only become harder to avoid: when data turns up in a new place, does the rule arrive with it?

Follow the trail

See the original product description, the Spotlight dashboard, the acquisition account, and a short video interview with Nong Li.