Breaking
YC P26 Astraea launches AI operating system for clinical trial biometrics Raw study data → FDA-ready outputs in days, not months Founders: Joshua Wang (CEO) & Sanmay Sarada (CTO), ex-Stanford Runs on-prem inside the sponsor's own environment Target: Phase II/III oncology & rare-disease studies Claimed timeline compression: 30-50% YC P26 Astraea launches AI operating system for clinical trial biometrics Raw study data → FDA-ready outputs in days, not months Founders: Joshua Wang (CEO) & Sanmay Sarada (CTO), ex-Stanford Runs on-prem inside the sponsor's own environment Target: Phase II/III oncology & rare-disease studies Claimed timeline compression: 30-50%
Company Health · AI Y Combinator P26

Astraea Wants to Turn the Slowest Part of Drug Development Into a Weekend Job

Every new medicine waits on months of manual data wrangling before it reaches the FDA. A two-person San Francisco startup thinks that math is broken - and is rebuilding clinical trial biometrics as an AI pipeline.

After a clinical trial finishes collecting data, most people assume the hard part is over. It isn't. What follows is one of the least visible bottlenecks in medicine: months of manual work to convert raw study data into the exact datasets, tables, and documents the FDA will accept. Teams of five to ten specialists spend the better part of a year on it, one spreadsheet and one SAS program at a time.

Astraea, a San Francisco startup in Y Combinator's Spring 2026 batch, was built around a single observation: the slowest step in getting a drug approved is also the most manual. Its founders, Joshua Wang and Sanmay Sarada, are trying to replace that step with software - an AI-native platform that ingests a protocol and raw data at one end and produces submission-ready regulatory outputs at the other.

~9 mo
Typical biometrics cycle
5-10
People per study
Days
Astraea's target
30-50%
Claimed timeline cut

The problemThe paperwork between a finished trial and a filing

Clinical trial analysis has its own vocabulary, and none of it is friendly. Raw data has to be mapped into SDTM, a standardized structure regulators expect. That gets transformed again into ADaM, the analysis-ready version. From there, statisticians generate TFLs - tables, figures, and listings - plus the statistical programs behind them, and run rounds of quality control before anything is packaged for submission.

In most organizations this happens across a patchwork of Word documents, Excel specifications, and SAS code, with repeated QC cycles catching errors by hand. It is slow, expensive, and easy to get wrong. It is also unavoidable: skip a standard and the submission bounces.

The unsexiest bottleneck in drug development is manual statistical programming. Astraea decided that was the point. The wedge

How it worksA protocol goes in, a submission comes out

Astraea describes itself as an operating system for clinical development. The pipeline starts by ingesting the raw clinical data, the protocol, and study metadata. From the protocol and study design it generates TFL shells, the statistical analysis plan, and eCRF structures. It then builds CDISC-compliant SDTM datasets, derives ADaM datasets, and produces the tables, figures, listings, and traceable statistical programs on top - running QC checks, including Pinnacle 21 validation, along the way.

1
Ingest
Protocol + raw data + metadata
2
Map
CDISC-compliant SDTM datasets
3
Derive
ADaM + TFLs + stat programs
4
Submit
QC'd, FDA-ready outputs
The Astraea pipeline, compressed. What normally moves through a relay of specialists over months is run as one auditable system.

The word the company keeps returning to is auditable. Automating the labor is one thing; convincing a regulator to trust the output is another. Astraea's answer is to keep humans in the loop for review, validation, and every regulatory decision, while the machine does the wrangling. The platform is built to be 21 CFR Part 11 compliant with audit-ready traceability, and it runs inside the sponsor's own environment so raw patient data never leaves the building.

Manual today
~9 months
With Astraea
days
Astraea's central claim, drawn to scale. The gray bar is the status quo; the orange bar is the pitch. Figures are the company's own.

Under the hoodFour engines instead of one big model

Rather than route everything through a single model, Astraea splits the work across four specialized engines - Standards, Compliance, Evidence, and Stats. In a domain where every output has to be explainable to an auditor, that separation is a design choice as much as a technical one: narrow, traceable components are easier to defend than one black box.

StandardsCDISC-native mapping to SDTM and ADaM.
Compliance21 CFR Part 11, Pinnacle 21 validation.
EvidenceTraceability from raw data to output.
StatsTFLs and the programs behind them.

The foundersA math olympian and a fetal-medicine data engineer

The pairing behind Astraea is the kind investors like for deep-domain problems: one founder who speaks agents and one who speaks the regulated clinical world. Joshua Wang, the CEO, studied CS and math at Stanford before leaving to build - a USAMO top-250 qualifier and USACO Platinum competitor who was the first intern at Salient. Sanmay Sarada, the CTO, is a Stanford CS graduate who worked as a data engineer at Baylor College of Medicine and co-authored four published papers in fetal medicine.

For deep-domain startups, the co-founder fit is the moat. Astraea needs someone who can build multi-agent systems and someone who has lived inside regulated clinical data. Why the team matters

The marketStarting where it hurts the most

Astraea is aimed at pharmaceutical and biotech sponsors running Phase II and Phase III trials, and it starts in oncology and rare disease - the studies where analysis complexity is highest and every saved week matters most. That is a deliberate inversion of the usual "start with the easy case" playbook. The hardest trials carry the most manual labor to remove and the customers who feel the pain most acutely.

The competitive field is not other chatbots. It is contract research organizations and legacy tooling - the human services firms and SAS-based workflows that do SDTM and ADaM programming by hand today. Astraea's bet is that automating the whole pipeline as an agentic system, run privately inside the customer's walls, beats renting people or buying point tools. On business model, it is straightforward enterprise software: sold to sponsors, deployed in their environment, priced against the months and headcount it removes.

Where it standsEarly, but pointed at a real number

Astraea is still small - a founding team of a few people, a $130K seed tied to its YC batch, and a public launch in May 2026. Whether it can earn a regulator's trust at scale is the open question, and the company would be the first to say a human signs off on the parts that count. But the thing it is aiming at is not vague. It is a nine-month cycle, measured in specialists and QC rounds, sitting between finished science and an approved drug. Shrink that, and everything downstream moves sooner.

#clinical-trials#biometrics#cdisc#sdtm#adam#fda-submission#ai-agents#healthcare-ai#yc-p26#life-sciences