Imagine a photograph of a pedestrian. One engineer fixes its label. Another removes it from a training set. A third gets a better model score and cannot quite explain why. This is an illustrative scene, but the difficulty it describes is the territory Graviti occupies: the awkward distance between possessing data and knowing exactly what you have done with it.
- Graviti organizes images, video, audio, annotations, and metadata for AI teams.
- Git-like branches and commits give datasets a traceable history.
- Its commercial platform bundles curation, visualization, permissions, and automated workflows.
The company’s proposition is wonderfully unglamorous. Before a machine can learn from a dataset, people must find it, inspect it, correct it, share it, and preserve it. Graviti sells software for that work. A model may receive the applause; someone still has to keep the rehearsal notes.
The engineer who kept meeting the same problem
Edward Cui came to this problem through self-driving cars. After switching from mechanical engineering to computer science at the University of Pennsylvania, he joined Uber’s Advanced Technologies Group in 2015. Cameras and lidar produced a daily accumulation of unstructured sensor data. In a 2022 interview, Cui described shortages of compute and data-management infrastructure early in that work.
The order matters. The practical difficulty appeared while engineers were trying to make AI work with real-world inputs. Cui founded Graviti in 2019 after seeing an opportunity in managing those inputs at scale. The company’s origin is a reminder that an engineering bottleneck can be a business idea, provided enough other people keep reaching the same narrow passage.

By July 2021, Graviti’s product team was introducing TensorBay to developers as a Git-like management tool. In January 2022, the company publicly launched the broader Graviti Data Platform. Its intended users include enterprise AI teams and researchers; a published logistics testimonial describes better data preparation and training automation, without naming the company.
A dataset should remember who touched it
TensorBay’s documentation describes a sequence familiar to software engineers: create a dataset, open a draft, upload data, and publish a commit with a message and tag. Once released, that version cannot be modified. The point is to preserve a stable object that a later experiment can refer to, rather than a folder whose contents quietly change beneath its name.
Graviti adds branches, version comparison, and visual histories. Teams can work on datasets simultaneously and inspect differences between versions. Splitting or merging existing datasets is advertised as requiring no additional storage for the copied data. That makes the dataset a working structure with relationships and history, rather than a pile of files duplicated whenever somebody wants to try something.
The platform also previews data and annotations online. Supported material includes images, video, and lidar formats; labels include boxes, polygons, and classifications. Distribution views help engineers inspect what a dataset contains. In practical terms, a team can look for poorly represented examples or suspicious labels before those problems become a debate about the model.

Collaboration has its own plumbing. Graviti describes dataset-level permissions for viewing, editing, using, and managing data, alongside activity records. Action, its workflow feature, connects operations such as filtering and pre-annotation. Configurable triggers and node-level logs make the repetitive work easier to run and inspect. The attraction is the combination: fewer handoffs between otherwise separate jobs.
The meter runs on data
Graviti’s published pricing makes its business model unusually legible. Starter is free and lists 100GB of storage, 50,000 rows, and two hours of XS compute. Standard starts at $200 monthly; Premium at $800. Both advertise unlimited seats. Storage, records, compute, and transfers carry their own rates. These are published plan terms, not a quote for a particular workload.
That arrangement invites collaboration while putting the bill against workload. It also means a team should count more than colleagues: retained data, processing hours, and transfers all matter. On-premise plans are offered by arrangement. The buying question is whether managing the work together saves enough engineering effort to justify the subscription and consumption charges.
There are alternatives. DVC brings data and pipeline versioning into Git-based workflows. Voxel51’s FiftyOne concentrates on inspecting and curating visual datasets. A team can assemble tools around existing storage, or consider Graviti’s bundled approach. This is a comparison of documented capabilities; the useful distinction is how much integration work a buyer wants to own.
Sharing requires more than a download button
Graviti also helped start Project OpenBytes, announced by the Linux Foundation in November 2021. Its focus was open-data standards, formats, and the conditions that make sharing easier. Motional supported the initiative with its freely available nuScenes and nuPlan datasets. The obstacle here includes licensing uncertainty and inconsistent formats, both of which can make an available dataset surprisingly troublesome to use.
“Acquiring higher quality data is paramount if AI development is to progress.”Edward Cui, November 2021
Graviti’s Open Datasets listings range from autonomous-driving data to MNIST’s handwritten digits. The company did not create those datasets; it provides a route to discovering and using them. GroundTruth adds documented annotation tools, while Sextant supports evaluation against selected dataset versions. Together they show how widely the same question travels: what, precisely, are we feeding the machine?
A habit worth stealing
The practice a reader can copy is simple: record the dataset version beside every experiment. Keep labels with their inputs. Inspect distributions before collecting more examples. Set permissions deliberately. Automate a repeatable operation only after deciding what a successful run looks like. Those habits remain useful whether the software comes from Graviti or somewhere else.
The fit depends on the problem. A small, stable dataset may need little infrastructure. A team with an established versioning and inspection stack may gain less from another platform. And orderly storage cannot supply missing examples or settle whether a label is correct. Graviti’s proposition becomes most interesting when data changes frequently, several people touch it, and yesterday’s result must still be explainable tomorrow.