Albi Celaj had an inconvenient problem: the models were good. At Deep Genomics, researchers were building systems that beat published benchmarks for individual pieces of RNA biology. Yet the collection became difficult to maintain and harder to turn into drug programs. Success at the small task was obscuring trouble with the larger one.
- Deep Genomics models RNA biology to help design genetic medicines.
- Its platform combines broad AI models, specialist tools and laboratory feedback.
- BioMarin is a disclosed pharmaceutical collaborator.
- The latest results concern drug design, with clinical benefit still a separate test.
A very successful wrong turn
Celaj’s October 2025 account supplies an unusually candid origin for BigRNA. Splicing, the process that edits RNA transcripts, had offered a tractable early problem. Moving into drugs that increase gene expression meant confronting many interacting mechanisms. Which one mattered for a particular gene? A small team could accumulate models faster than it could organize their use.
“It was exciting, but it didn’t scale or give us any drugs.”
Albi Celaj · October 2025
In late 2021, Celaj tried a broader model over a weekend, using an older Google TPU through Colab. The session required restarting every 24 hours; he checked progress on his phone. It was an unglamorous beginning for BigRNA, which would learn RNA regulation across tasks instead of demanding a separate tool for every mechanism. The organizational frustration supplied the research question.
RNA has more than one instruction
Deep Genomics grew out of University of Toronto research and launched in 2015. Founders Brendan Frey, Andrew Delong, Hannes Bretschneider and Hui Yuan Xiong joined expertise in machine learning with genome biology. Their central concern was the gap between observing a genetic change and understanding what it does inside a cell.
Consider the company’s published Wilson disease research. A variant called ATP7B Met645Arg was associated with a missing segment in the RNA transcript: exon 6 was being skipped. The company announced a candidate, DG12P1, designed to correct that mechanism. The useful insight was the explanation connecting a particular variant to defective RNA processing. A mutation’s name alone would not have supplied a treatment strategy.

This puts Deep Genomics in a specific part of the AI drug-discovery market: the regulation of RNA and the design of genetic medicines. Envisagenics’s SpliceCore is an adjacent approach focused on RNA-splicing discovery. Deep Genomics has broadened its platform across expression, stability and editing, with experimental capabilities alongside its computational tools.
The pipette gets a vote
The company calls this system its BioFM Platform. BigRNA supplies a broad model trained on more than a trillion genomic signals. REPRESS predicts microRNA binding and messenger-RNA degradation. DeepADAR supports guide design for ADAR-mediated RNA editing. These are different questions about what happens to RNA, gathered into a workflow for molecular design and target biology.
- 01Design dataChoose the biological question
- 02PredictRank mechanisms and molecules
- 03TestMeasure actual cell behavior
- 04Learn againUpdate the next decision
Conceptual workflow, based on the company’s platform description.
Its two experimental facilities, in Toronto and Cambridge, together contain more than 10,000 square feet of laboratory space. Scientists generate datasets for particular tasks, check predictions with other assays and return the results to model training. A laboratory is therefore both a judge of the model and a supplier of its next lesson.

An expensive idea needs a customer
The financing grew with the ambition. The university reported a US$3.7 million seed round in 2015. A US$13 million Series A followed in 2017, then US$40 million in January 2020. SoftBank Vision Fund 2 led the US$180 million Series C in July 2021, alongside investors including CPP Investments and Fidelity. These are investments in the company, rather than published costs of producing a drug.
Round sizes shown on one scale. Investment is not R&D expenditure.
The clearest disclosed customer relationship is the BioMarin collaboration announced in November 2020. It covered four rare-disease indications. Deep Genomics would identify and validate mechanisms and lead candidates; BioMarin would advance selected programs into development. The agreement included an upfront payment, potential development milestones and exclusive options on the programs.
That arrangement explains the business better than a software subscription analogy. A pharmaceutical collaborator needs a scientific result that fits its development decisions. Patients are the intended eventual beneficiaries; the immediate users are research and development teams. Brian O’Callaghan became CEO in September 2023, while Frey moved into the chief innovation officer role.

The drug you would rather reject
On October 6, 2026, Deep Genomics introduced DeepRNAi, a model for predicting gene silencing caused by siRNA treatment. The central difficulty is familiar: a molecule may suppress its intended transcript while also suppressing others. A candidate needs potency and restraint. Simple sequence matching, the company argues, cannot fully predict the behavior of the silencing machinery.
Its reported tests included cell-line data and rat experiments. In a retrospective exercise, DeepRNAi scored 3,599 sequences complementary to PCSK9, the target of the cholesterol drug inclisiran. The company reports that inclisiran ranked as the most potent candidate with low predicted off-target burden. The exercise tested whether the model could recognize a useful existing molecule; it did not produce a new approved treatment.
That distinction matters. Retrospective ranking supports a design hypothesis. It cannot, by itself, establish safety in people. Celaj has also described difficulties predicting subtle differences between individuals and transferring models trained on healthy data into disease settings. Tissue, disease state and experimental context remain part of the question.
Borrow the loop
For non-commercial researchers, REPRESS offers a practical entry through its public GitHub repository. Pharmaceutical teams can explore collaboration around target biology and molecule design. BigRNA remains proprietary. Access therefore depends on the capability being sought and the terms attached to it.
The wider lesson is available without a license: organize research around the next decision. Generate data that can change a ranking. Keep experimental scientists and model builders close enough to argue productively. Deep Genomics’s most instructive move was recognizing that a pile of impressive tools could still leave the important work undone.