COMPANY SIGNAL
2023 $9M seed for AI accelerator research2025 $28M Series A for software-first infrastructure2026 Tachyon compiler and runtime take center stage2027 public beta planned for summer2023 $9M seed for AI accelerator research2025 $28M Series A for software-first infrastructure2026 Tachyon compiler and runtime take center stage2027 public beta planned for summer

Company profile / AI infrastructure

The Chip Company That Followed the Bottleneck

Lemurian Labs set out to build a better AI chip. Two pivots later, its bigger wager is that the costly part is teaching software to use the chips we already have.

The robot behaved beautifully in simulation. Then it met the real world. Jay Dawani, later the co-founder of Lemurian Labs, says a collaborative robotics project he led around 2019 saw its agents fail on physical machines after working in a virtual environment. Making the simulation more detailed made matters worse: the agents learned its quirks. Lower fidelity, plus more randomness, helped them generalize, though he puts reliability at only about 40 to 50 percent. That is a bruising result for anyone trying to build autonomous machines. It is also a useful beginning for a company now selling an argument about computer infrastructure: the apparent problem is often one layer above the real one.

The short version
  • Lemurian Labs builds Tachyon, a compiler and runtime intended to move AI workloads across CPUs, GPUs and other accelerators.
  • It began with robotics hardware, moved to a data center chip, then made portable software its main product.
  • The company has announced a $9 million seed round and a $28 million Series A; Tachyon beta testing is planned for summer 2027.

Dawani founded Lemurian Labs in 2022 with Dr. Vassil Dimitrov. The original ambition was a processor for autonomous robots: a machine that could perceive, plan and act without burning through a battery. The company called its proposed architecture the Spatial Processing Unit, or SPU. It also developed PAL, a parallel adaptive logarithm number format meant to make AI arithmetic more efficient. These were serious ideas, and expensive ones. A new chip needs more than a clever design; it needs a software system that developers can use, production capacity, and buyers willing to rewrite the habits of their teams.

The expensive lesson inside a small failure

The robotics experience supplied the first warning. A model can look excellent within the rules of its training environment and still be useless when the environment changes. That lesson applies to processors too. A faster accelerator on paper has little commercial value if a developer must spend months rebuilding software around it. Dawani later told interviewers that Lemurian moved from robotics to data center chip design, then to the software layer for heterogeneous systems. The company had followed the obstacle from the robot to the processor and finally to the act of programming processors.

Members of the Lemurian Labs team gathered together
The people hired to change the plumbing. AI may be glamorous upstairs; someone still has to make the basement work.Image: Lemurian Labs newsroom

The second warning came from scale. In a 2023 interview, Dawani recalled abandoning an effort to train a model with roughly 2.1 billion parameters because the cost could not be justified. That number now looks modest beside the largest models, but the decision captures the question Lemurian keeps asking: when a workload grows, which part of the system becomes unbearable first? In 2023, the company’s public answer was specialized silicon. It raised $9 million in seed funding led by Oval Park Capital and discussed a platform built around PAL and the SPU. Its claims of up to 20 times the throughput of legacy GPUs referred to simulated or projected performance, not a published production comparison.

That shift is more interesting than a simple pivot story. The team had to keep asking what a customer would actually buy. A new AI accelerator asks a customer to bet on the device and its supporting tools. Portable software offers a different bargain: keep the model, try another machine, and judge it on cost and performance. In Dawani’s account, prospective customers repeatedly identified software maturity as the problem. He said the company abandoned its hardware focus and put its engineering effort into Tachyon.

“For decades, faster chips delivered ‘free gains,’ but now the real bottleneck is software.”Jay Dawani, December 2025

A compiler wants to see the whole room

The technical target is familiar to anyone who has tried to move a machine learning model between processors. The high level code may be portable; the fast path often is not. Hand-tuned kernels, vendor libraries, memory layouts and scheduling choices bind a workload to a particular architecture. A switch of chip or cloud can become a fresh performance project. Lemurian’s claim is that Tachyon can ingest PyTorch models and optimize execution across CPUs, GPUs and specialized accelerators without developers hand-writing a new collection of kernels for each target.

Its pitch is not merely translation. The company says Tachyon looks at the system as a whole, allowing it to overlap communication with computation and schedule work across processors. A kernel optimized in isolation cannot see all the work moving through a cluster. If Lemurian can make that broader view fast enough, a developer could test a different accelerator or cloud instance before committing to a migration. That would make hardware comparison an engineering experiment rather than an organizational ordeal.

$9M2023 seed round backing chip and arithmetic research
$28M2025 Series A backing the software-first direction; includes earlier convertibles

There is an important tense in that description. Tachyon is the flagship product Lemurian describes today, and the company says it has built the underlying compiler and runtime. Its website schedules beta testing for summer 2027. Independent production benchmarks, a price list and named Tachyon customers are not public. The right comparison is therefore between a stated technical ambition and the measurable results a beta will need to produce. The older 20-times chip figure should not be carried forward as a Tachyon result.

Jay Dawani speaking at a PyTorch event
Jay Dawani at a PyTorch event. The audience knows that “runs anywhere” and “runs well anywhere” are two different promises.Image: Lemurian Labs newsroom

Who pays for freedom of movement?

The natural buyers are teams whose AI models already run at a meaningful scale: enterprises, cloud providers, AI developers and chip companies trying to make new accelerators useful. A small team with one model on one well-supported GPU may have little reason to change its workflow. A team managing different GPU generations, rented capacity across clouds, and edge deployments has more to gain from making the workload less dependent on any one vendor. The costs to compare are concrete: engineering time, utilization, power, migration delay and the price of the next available machine.

Lemurian has not published a Tachyon price or a detailed commercial model, so its eventual economics remain open. The company has described interest from cloud providers, semiconductor companies and enterprise AI users, and Dawani has discussed an unnamed supply chain planning project using some of its earlier multi-agent research. That is evidence of technical engagement, not a public roster of Tachyon customers. Pebblebed Ventures and Hexagon co-led the December 2025 Series A, which Lemurian said included capital previously raised through convertible securities. The sum bought time to hire engineers, build the product and work with ecosystem partners; it did not settle the performance argument.

Its team shows where the bet has moved. Christopher Vick, who worked on the HotSpot Java optimizing compiler at Sun Microsystems, became vice president of engineering in 2024. In 2025 Lemurian added infrastructure and product leaders, including Adam Robertson and Abhay Chitral. In March 2026 it appointed Java pioneer Kim Polese to its board and MIT compiler researcher Saman Amarasinghe as a technical advisor. The company’s own culture page emphasizes testing quickly, candor and changing course when evidence demands it. Those are credible rules for a team that has already changed course twice.

The test that matters now

For a buyer, the useful lesson is less romantic than “follow the bottleneck.” Write down the workload first. Measure the present cost per training run or inference request, including the people who tune it. Then move the same model to a second processor or cloud and count every rewrite, failed optimization and lost week. A portability tool earns its place only if that bill falls without a painful loss of speed, accuracy or operational control. It may disappoint on highly specialized workloads that already have excellent vendor kernels, or where a team will never leave its chosen hardware. No abstraction gets to declare victory by hiding the invoice.

Lemurian Labs has made its position unusually plain: the celebrated chip is only useful if software can put it to work. The irony is that a company formed to design one more accelerator now hopes to make it easier for customers to choose among all of them. By summer 2027, if its beta arrives as planned, the debate can move from speeches about kernels to a more satisfying question: how much work does it actually save?