The first thing to understand about artificial intelligence is that the exciting part ends early. A model identifies the prescription, draws the frame or chooses the next model in an agent’s chain. Then comes the plumbing: memory, queues, GPUs, routing, autoscaling, audit logs and an invoice with the emotional range of a parking ticket. Simplismart works in this second act. It was founded in 2022 by Amritanshu Jain, an Oracle Cloud engineer, and Devansh Ghatak, who had worked on search at Google. Their origin story does not begin with a grand theory of intelligence. It begins with Jain stuck in a loop, writing repetitive production code late at night.
- Simplismart deploys, tunes, scales and monitors open-source and custom AI models.
- Customers can use shared APIs, dedicated GPUs, their own cloud account or on-premises hardware.
- Its public cases span Tata 1mg, Physics Wallah, Sanas, Invideo, Ema and Dashverse.
- The business charges by token or audio minute, by GPU hour, and through negotiated enterprise contracts.
- The useful lesson: optimize for the actual workload, not the most impressive model or GPU.
The failure happened after the demo
Jain and Ghatak had seen the same small tragedy from different desks. A machine-learning experiment worked; moving it into production required specialized engineers to rebuild the surrounding system. Models had to be compiled, versioned and observed. Traffic did not arrive politely. A generic endpoint might be convenient but expensive, inflexible or impossible to use with sensitive data. The founders say two hackathon wins and even a rejected acquisition offer came before the conclusion that changed their minds: the model lifecycle was not the side problem. It was the product.
Early Simplismart was pitched as a no-code MLOps system, with a charming metaphor about assembling a Subway sandwich. By 2024, generative AI had clarified the sharper opportunity. The company moved the center of gravity to inference - the moment a trained model does useful work - and raised a $7 million Series A led by Accel. A reported $9 million Series B followed in August 2026.
“We wanted to build exactly what we wish we had.”Amritanshu Jain, co-founder and CEO
Tailor-made is a technical claim
Simplismart’s platform takes models from sources such as Hugging Face or a customer’s repository, compiles and serves them, then manages autoscaling, traffic, versions and monitoring. It supports language, speech, image, video and multimodal systems. More than 150 open-source models are available through shared endpoints; custom weights can run on dedicated infrastructure. The same control plane can sit over Simplismart’s cloud, AWS, Azure, Google Cloud, a private VPC or an on-premises Kubernetes cluster.
The distinction from a generic model API is not merely where the server lives. Simplismart tunes the combination of model, runtime, chip and traffic pattern. A voicebot cares about pauses measured in milliseconds. A university lecture can be processed asynchronously but may arrive in a tidal wave at the end of class. A healthcare workflow cares about data residency and auditability. Calling all three “inference” is like calling a bicycle, a freight train and an ambulance “transport.” Accurate, and not terribly helpful.
a model
and benchmark
and observe
The numbers are the plot
Ema, an enterprise agent company, had two routing models that decided which foundation model should handle a request. Their sequential journey took more than ten seconds. Simplismart introduced continuous batching, overlapped CPU and GPU work, adjusted kernels and removed scaling bottlenecks. The reported median fell to 0.9 seconds. It is hard to sell an “instant” assistant when the receptionist spends ten seconds deciding whom to call.
At Sanas, a real-time voicebot ran on expensive H100 GPUs. Custom INT4 quantization moved the Gemma 3 model to A10Gs while retaining a reported 99.99% of task accuracy and cutting cost by 53%. Invideo’s image and video pipeline used fused kernels, caching and memory-efficient attention; the company reports a 45% reduction in median latency and 56% lower serving cost. Dashverse’s comic-generation system traded separate model endpoints for a unified layer, dropping generation time from about ten seconds to six and GPU cost by roughly 47%.
The strangest case may be Physics Wallah. Long lectures - sometimes twelve hours - arrive in synchronized bursts when classes finish. Simplismart split speech from silence, packed similar segments together and ran asynchronous Whisper jobs. The reported throughput reached roughly 900 times real time, with every video transcribed inside a fifteen-minute service target. When the classrooms went quiet, the GPUs scaled to zero.
What it costs, and what you are buying
The public price card makes Simplismart unusually legible for enterprise infrastructure. Shared endpoints are metered by tokens or audio minutes. Dedicated and training jobs are priced by GPU hour. Private-cloud and on-premises deployments are customized. In September 2026, listed hourly prices ran as follows:
| GPU | Public rate | Typical logic |
|---|---|---|
| NVIDIA T4 | $1.20 / hour | Economical, lighter work |
| NVIDIA A10G | $2.00 / hour | Voice and tuned models |
| NVIDIA A100 | $3.00 / hour | Heavier production loads |
| NVIDIA H100 | $4.00 / hour | High-throughput inference |
| NVIDIA H200 | $5.20 / hour | Larger memory demands |
The hourly price is only the cover charge. The real purchase is utilization: how much useful work happens before the clock advances. AWS says warm pools and custom autoscaling brought Simplismart customer scale-up time down from five or six minutes to 60–70 seconds and reduced infrastructure cost by as much as 40%. It also reported that Simplismart’s business grew sixfold in six months and deployed GPU hours rose eightfold in three months.
Its place in a crowded machine room
Baseten, Replicate, Modal, Together AI, Fireworks, TrueFoundry and self-managed open-source stacks all offer parts of the same future. Simplismart’s position is control: visible Kubernetes infrastructure, many clouds, on-premises installation and engineers willing to alter the serving stack for a particular workload. A fully managed rival may hide more machinery and get a simple application live faster. Simplismart becomes persuasive when the machinery itself is the problem - compliance, bursty demand, custom weights, odd pipelines or a GPU bill large enough to deserve its own meeting.
That condition matters. A small team serving modest traffic through a standard model can sensibly choose a plain API. Tailoring has overhead. Running in your own cloud brings permissions, clusters and procurement back into the room. Published customer results are also case studies, not controlled trials; workloads differ, and a 56% saving in one video pipeline is not a coupon for the next. The right comparison is measured on your traffic, accuracy requirement and latency target.
The quiet layer becomes a market
Simplismart is still small enough that its claims must travel farther than its name. Yet the market is moving in its direction. Open models multiply. Enterprises want them in private environments. Every new modality adds another workload shape. The company’s partnerships with AWS and Microsoft place it inside major cloud channels; its 2026 white-label platform for NVIDIA Cloud Partners turns the product into infrastructure for infrastructure providers. Forbes Asia put Simplismart on its 100 to Watch list the same year.
There is a Wildean joke hiding in the stack: the industry knows the value of every model and the cost of no inference call. Simplismart’s advantage is that it finds the cost fascinating. Its customers do not need the fastest benchmark in the abstract. They need a prescription approved, a voice answered and a lecture transcribed before anyone notices the machines thinking. The less visible the infrastructure becomes to the end user, the more successful it has been.