The virtual lab that makes computational research reproducible, traceable and shareable - so an analysis runs the same way today or a decade from now.
Every computational scientist has lived the same small tragedy. A result looks clean, the paper ships, and then someone - a reviewer, a colleague, the researcher's own future self - tries to run the code again and it collapses. A library version has moved on. A data file is missing. The laptop it lived on has been wiped. The finding was real, but the machinery that produced it is gone. This is the reproducibility crisis, and it is less about fraud than about entropy: software environments rot.
Code Ocean, a New York company founded in 2016 out of a postdoc at Cornell Tech, was built to stop that rot. Its central idea is a Compute Capsule - a single, version-controlled bundle that holds the code, the data, and the exact software environment together, tied to the results it produces. Open a capsule and it runs. Share it and it runs on someone else's screen too. Publish it and it gets a DOI, the same kind of citable identifier a journal article carries. A running piece of analysis becomes a durable, checkable object rather than a disposable byproduct.
"A consistent, automated virtual lab is absolutely critical to today's digital age of science."
- Simon Adar, Co-founder & CEO, Code OceanThat framing - a virtual lab rather than a file format - is the whole company. Code Ocean does not just store notebooks. It gives researchers workstations they recognize (JupyterLab, RStudio, MATLAB, VS Code, Shiny), auto-generates the Dockerfile most scientists would rather not write by hand, provisions GPUs in a few clicks, and records where every result came from. The pitch to a research organization is blunt: the science you paid for should still work after the person who did it has left.
The Allen Institute for Neural Dynamics is the clearest public example of what the platform looks like when it is fully loaded. According to an AWS case study, the deployment supports hundreds of researchers running more than 450,000 computations across two-plus petabytes of reproducible data - managed by an IT team of roughly five people. That ratio is the argument for platform thinking in one line: you do not scale science by hiring more system administrators, you scale it by making reproducibility and compute self-serve.
Code Ocean sells to organizations where a computational result has to hold up - pharmaceutical and biotech companies, academic and non-profit research institutes, and the publishers who referee their work. Named users span biotech firms such as Lantern Pharma and CytoReason through to large research institutions. Panna Sharma of Lantern Pharma has described the platform as "an ideal environment for our scientists and engineers to design and scale our proprietary A.I. for drug discovery."
Three, mainly. First, reproducibility: capsules freeze the environment so a result can be reproduced on demand. Second, traceability: an immutable lineage graph records how every output was derived, which matters enormously in regulated life sciences where an auditor may ask exactly how a number was produced. Third, collaboration and continuity: work no longer lives on one person's machine, so a team can pick up an analysis - and a new hire can inherit one - without archaeology.
Plenty of tools touch pieces of this. Docker captures an environment; a notebook captures code; a data lake stores files; a workflow engine runs pipelines. Code Ocean's distinction is combining code, data, environment, version control and result lineage in one traceable unit, and running it securely inside the customer's own cloud. It installs as a VPC in the customer's environment with no data egress, carrying ISO 27001, HIPAA, GxP and GDPR compliance - a design that matters more to a pharma buyer than any single feature.
"Code Ocean's software-as-a-service platform is now the one place where research comes together in a trusted, virtual laboratory space."
- Scott Tobin, General Partner, Battery VenturesThe core unit: code, data and environment version-controlled together, auto-generated Dockerfiles, fully exportable, publishable with a citable DOI.
Visual pipeline builder or one-click nf-core import, running on AWS Batch with automatic scaling - Docker, git and Nextflow handled for you.
Turn capsules and pipelines into parameterized apps a bench scientist can run without ever opening a terminal.
FAIR-compliant data management with custom metadata plus an immutable provenance graph capturing how each result was produced.
GPU environments in a few clicks and model tracking from development to production, with reproducibility and lineage built in.
A secure AI assistant running Anthropic Claude via Amazon Bedrock inside the customer's cloud, with every action logged for 21 CFR Part 11 audit trails.
The business model is enterprise SaaS. Code Ocean deploys inside a customer's own cloud on custom pricing, aimed at pharma, biotech and research institutions, with a separate academic tier that leans into open science. Its integrations with publishers act as a distribution channel as much as a feature: when Nature Portfolio journals and IEEE let reviewers run a paper's code in the browser, every submitting author meets the platform. Springer Nature expanded that partnership across the full Nature Portfolio in 2024.
In the wider market, Code Ocean sits among reproducible-compute and bioinformatics platforms - the nearest alternatives include Seqera (Nextflow Tower), DNAnexus, Latch Bio, Terra, Rescale and Databricks. Where several of those lead with pipeline execution or raw scale, Code Ocean's center of gravity is the reproducibility and traceability of the whole analysis, packaged for regulated environments. It is less "run this workflow fast" and more "prove this result, forever."
"Computing has given us vastly expanded speed and simultaneously dramatically increased complexity."
- Simon Adar, Co-founder & CEOCode Ocean was co-founded by Simon Adar, whose doctorate was in hyperspectral image processing and who worked with a national space agency before turning to research reproducibility, and Ram Dayan, a veteran engineer who led the build of the platform. The pivot from remote-sensing imagery to scientific plumbing was not a leap so much as a recognition: Adar kept watching good research get stranded inside code no one else could run. That lived experience - not a spotted market gap - is the origin story most founders in this space share.
"Code Ocean is uniquely positioned to solve fundamental issues in computational research reproducibility, collaboration, security and productivity."
- Irad Dor, Partner, M12 (Microsoft's Venture Fund)Simon Adar and Ram Dayan launch Code Ocean out of a postdoc, with an early seed round.
Journals begin letting authors and reviewers run code tied to articles through the platform.
Battery Ventures leads a round, bringing total funding to about $21M to accelerate in biotech.
Battery Ventures and Microsoft's M12 co-lead a round to expand the digital lab platform.
The platform scales to a major neuroscience institute with petabyte-scale reproducible data.
Springer Nature extends code deposition and peer review across the full Nature Portfolio.
A secure AI agent powered by Claude via Amazon Bedrock launches inside customer clouds.
AI can execute real analyses with every action versioned, traceable and compliant.
Adar pivoted from hyperspectral imaging and space-agency work to research reproducibility.
Published Compute Capsules get a DOI, making a running analysis a citable scholarly object.
Nature and IEEE reviewers can execute a paper's code with no install at all.
The Aqua assistant runs Claude inside the customer's own cloud, so regulated data never leaves.
Idle compute instances auto-terminate after roughly two hours - cost control baked in.
It provides a cloud platform that makes computational research reproducible and traceable by packaging code, data and the exact software environment into shareable Compute Capsules that anyone can re-run.
Computational scientists, bioinformaticians and ML researchers at pharma and biotech companies, universities and research institutes such as the Allen Institute, plus journals like Nature and IEEE for code sharing and peer review.
It combines code, data, environment, version control and result lineage in one traceable, shareable unit - auto-generating Dockerfiles, giving published capsules a DOI, and running securely inside a customer's own cloud with compliance built in.
Over $30M disclosed, including a $15M Series A in 2021 and a $16.5M Series B in 2022 co-led by Battery Ventures and Microsoft's M12.
Yes. Its Aqua assistant and Trusted Agents run Anthropic's Claude via Amazon Bedrock inside the customer's private cloud, so AI can create and run analyses while every action stays versioned, traceable and audit-ready.
Watch: product feature demos and talks are published on Code Ocean's YouTube presence and blog - search "Code Ocean product demo" on YouTube for the latest walkthroughs, and see the founder interview on Zenodo.
Photograph / caption, Vincent Musi style: The wordmark sits still while the work moves - hundreds of thousands of computations, petabytes of data, a five-person team keeping it all reproducible. Code Ocean is the kind of company you only notice when the science holds up.