On the wire
01 OPEN-SORA 2.0: REPORTED $200K TRAINING COST02 COLOSSAL-AI: 41K+ GITHUB STARS03 GPU CLOUD, FINE-TUNING AND MODEL APIS01 OPEN-SORA 2.0: REPORTED $200K TRAINING COST02 COLOSSAL-AI: 41K+ GITHUB STARS03 GPU CLOUD, FINE-TUNING AND MODEL APIS

Company profile / AI infrastructure

The $200,000 Answer to a Billion-Dollar Question

HPC-AI Tech sells a way to make expensive AI work less expensively. Its open-source video model supplies the most interesting receipt: an 11-billion-parameter experiment built for a reported $200,000.

Here is an awkward question for anyone selling GPU time: what if the customer could use less of it? HPC-AI Tech, a Singapore-based AI infrastructure company, has spent years trying to answer exactly that. Its Colossal-AI software helps large models spread work across chips. Its cloud rents the chips. Its Open-Sora project, an open video model, offers a public test of the entire arrangement. The two businesses seem to pull in opposite directions until you notice what customers actually buy: a finished training run, an inference service, a working product. A faster route to that result can be more valuable than another hour on the meter.

The short version

  • Founded in 2021 by computer scientist Yang You, HPC-AI Tech combines open-source model software with paid AI infrastructure.
  • It offers Colossal-AI, Open-Sora, on-demand GPU machines, reserved clusters, fine-tuning tools and hosted model APIs.
  • The team says Open-Sora 2.0, an 11-billion-parameter video model, cost about $200,000 to train. That figure covers a reported training effort, not every expense of building a research program.

The problem starts before the GPU arrives

AI models have become large enough that an ordinary development habit can turn punitive. A model may not fit in one GPU's memory. Splitting it across several machines introduces communication overhead. Data has to arrive when the processors are ready, and the processors have to agree on what happens next. The hardware bill is only the visible part of the problem. An idle accelerator still clocks up an invoice.

Yang You came to this problem from high-performance computing. He earned a computer science PhD at UC Berkeley and became a professor at the National University of Singapore. He founded HPC-AI Tech in 2021. Its stated mission is to increase AI productivity, a phrase that sounds abstract until translated into the engineer's daily work: more useful model progress per unit of compute and effort. Colossal-AI was the company's first public proof. The open-source system provides parallel training strategies, memory management and tools intended to let developers scale work that begins on a smaller machine.

Yang You, founder of HPC-AI Tech, in a formal portrait
Yang You knows the price of a slow machine. His company now sells both the fix and access to the machine.

The origin matters because this is a market full of people promising access to scarce chips. HPC-AI Tech's different claim begins in software. Its public GitHub projects let developers inspect the methods, run them elsewhere and argue with the results. By September 2026, Colossal-AI had roughly 41,000 stars and Open-Sora roughly 30,000. Stars do not buy capacity, but they do show that the code has drawn an audience beyond a sales deck.

“We aim to support you to write your distributed deep learning models just like how you write your model on your laptop.”Colossal-AI project description

A video model becomes the invoice

Open-Sora makes the company's efficiency argument easier to see. The project began with short clips and published training code, weights and technical reports as it evolved. Its 1.2 report disclosed a training run of 35,000 H100 GPU hours for a 1.1-billion-parameter model. It also listed limits: some 720p, 16-second settings were outside the training distribution, and certain inference settings required more than one 80 GB GPU. The disclosure is useful precisely because it does not pretend that an impressive demo erases a hardware constraint.

Earlier versions had more basic trouble. In the 1.1 report, the team described failed generations on complex or long scenes, noisy motion and weak consistency over time. It suspected a temporal attention problem and said its frame-dropping approach cost fluency. For 1.2, it introduced a video compression network that compressed time as well as space, among other changes. This is the closest thing in the public record to a change of mind: a concrete defect, a stated technical diagnosis, and a revised design. The work did not turn every long, high-resolution prompt into a reliable video, but the team left enough detail for another engineer to learn from the miss.

In March 2025 the team released Open-Sora 2.0, an 11-billion-parameter model, with code and weights. It reported a training cost of about $200,000. The figure is eye-catching, but it needs a careful reading. It is a reported model-training cost, not a complete ledger for data acquisition, earlier experiments, salaries, infrastructure development or serving every future video. A team copying the result still needs the technical skill to prepare data, run the pipeline and measure quality. What the number does show is that engineering choices can move the price of a serious model substantially.

$200KReported Open-Sora 2.0 training cost
11BOpen-Sora 2.0 parameters
35KH100 GPU hours in Open-Sora 1.2 report

The engineering work was not a single trick. The public reports describe video compression, parallelism, data preparation and architecture changes across releases. Open-Sora 2.0's repository also publishes an inference table: at 768 by 768 resolution, adding GPUs reduced a test run from 1,656 seconds on one GPU to 276 seconds on eight. That is the kind of result a buyer should ask for, along with the conditions under which it was measured. Speed claims detached from model size, resolution, precision and hardware are advertisements in a lab coat.

A published inference example

1 GPU
1,656 sec
8 GPUs
276 sec
Open-Sora 2.0 repository figures for 768 × 768 text-to-video generation. More hardware shortens elapsed time here; it does not make the extra hardware free.

The open workshop, the paid machine

The commercial path is unusually legible. Developers can discover Colossal-AI or Open-Sora on GitHub. They can try the model and read the documentation. If they need a ready machine, HPC-AI.COM offers on-demand GPU rentals and reserved clusters, including H200 and B200 systems. The cloud advertises per-second billing, network and storage aimed at multi-GPU workloads, and full eight-card machine rentals for its on-demand offering. Its quoted starting rates vary by hardware and time, so any comparison needs a current quote for the exact configuration.

The company now sells higher layers of the stack too. Its fine-tuning SDK lets developers submit training work to managed infrastructure. Model APIs provide hosted inference through an OpenAI-compatible interface and charge by token usage. Video Ocean, its creator-facing video service, shows another way the same technical base can become a product. This is why calling HPC-AI Tech merely a GPU rental company misses the point. It is trying to capture value at every step where a team moves from model idea to deployed application.

One workload, three doors

01 / SOFTWAREBuild and optimize with Colossal-AI or Open-Sora.
02 / COMPUTERent GPU capacity or reserve a cluster for the run.
03 / SERVICEFine-tune a model or call a hosted model API.
The doors are related, but customers can enter at different points. Open code is also usable outside the company's cloud.

The paying audience includes AI startups that cannot buy a cluster, researchers whose workloads briefly exceed institutional capacity, and enterprises that want large models without hiring an infrastructure team for every experiment. In its 2024 financing announcement, HPC-AI Tech said it had 50 key clients, including four Fortune 500 companies. It named organizations such as AWS, IBM and Intel among users. Those names should not be mistaken for identical commercial contracts or formal partnerships. The stronger evidence of market fit is the widening product line: different buyers need code, compute and inference in different combinations.

Where the bargain stops

A bargain in GPU computing is conditional. A small job that needs only one GPU may not suit an eight-card rental. A model with unusual memory needs may make a cheaper chip the expensive choice if training stalls. Data transfer, storage, engineering time and downtime can swallow a low hourly rate. A hosted API may be simpler until token volume makes direct deployment cheaper. HPC-AI Tech competes with hyperscalers such as AWS and Google Cloud, specialist GPU clouds such as CoreWeave, Lambda and Runpod, and open toolchains that let skilled teams manage their own clusters. There is no universal cheapest option; there is a cost for a particular workload, at a particular scale, with a particular tolerance for operational work.

The useful lesson to copy is the company's way of making an infrastructure claim testable. Publish the code. State the hardware. Show the measured run. Separate the price of the final training pass from the cost of discovering how to do it. HPC-AI Tech did not solve the entire economics of AI with a $200,000 figure. It did something more useful: it gave developers enough detail to ask whether their own GPU bill is a fact of nature or a design decision.