In 2024, Sudarshan Kamath and Akshat Mandloi tried to build a voice agent from the parts already on the market. They kept assembling models, testing combinations and listening to the results. According to Kamath's own account, the agents were slow, expensive and prone to breaking. The speech generator was often the weak link: wait seconds for a good voice, or get a quicker one that sounded mechanical. So the two founders built their own. They called the first text-to-speech model Lightning.
The choice says a good deal about Smallest.ai. Most people think of an AI phone agent as a clever answer machine. A caller has a different measure. After a question, how long is the silence? When the caller interrupts, does the machine stop? If an accent or a noisy room changes a word, does the conversation recover? The intelligence may be impressive, but the timing is what makes the encounter feel like a conversation.
- Smallest.ai sells speech-model APIs and a platform for building phone and web voice agents.
- Its own models handle transcription, speech generation and, in beta, direct speech-to-speech interaction.
- Its public customers include Kogta Financial and Pocket; the company raised a $13 million Series A in July 2026.
- The useful test is an entire call: speed, accuracy, handoff and auditability together.

01 / The originFirst, the voice failed
Smallest.ai lists 2023 as its founding year. The decisive product turn came in 2024, when Kamath says the founders moved from a broader interest in small AI models to the immediate problem of live voice. They were classmates at IIT Guwahati, and their early work looks like a researcher's response to an awkward product test: identify the slow part, then replace it. There is a neat irony in calling a company Smallest.ai while taking on a problem that consumes so much infrastructure. The name describes the intended method, not the scale of the ambition.
A text chatbot can take a beat without offending anyone. A telephone caller hears that beat as hesitation. The company therefore built Lightning to start producing audio quickly, and later added Pulse for speech recognition, Electron for fast conversational reasoning and Hydra, now in beta, for speech-to-speech interaction. Atoms is the layer that turns the models into deployable agents, with calls, workflows and analytics. Waves exposes the component models through APIs. A buyer can take a single part or use much of the stack.
The typical chain. Hydra's beta speech-to-speech approach aims to combine more of these steps in one model.
The company advertises time to first audio under 100 milliseconds for Lightning v3.1. That is a measure of the speech model beginning to speak, not the time from a customer's question to a useful answer. Speech recognition, reasoning, network travel and telephony all take their share. Smallest.ai's distinctive claim is that owning several of those layers lets it tune the whole exchange, including interruptions and handoffs, instead of asking customers to stitch them together.
02 / In productionFifteen million chances to get a call wrong
Kogta Financial offers a less polished test than a product demo. The Indian lender needed to make debt-collection calls across regional languages, accents and borrower situations. A fixed script was a poor fit. Kogta wanted to edit call flows itself and see what happened as changes went live. According to Smallest.ai's customer account, the company moved from proof of concept to production in two weeks and eventually placed more than 15 million collection calls using the platform.
The reported scale deserves its own arithmetic. Kogta says its system reached more than 130 simultaneous calls at peak, with a 0.2% failure rate. Switching some agents from a larger general model to Electron, Smallest.ai's small language model, reportedly cut latency by about half. Those are customer-case figures, not an independent audit. Still, they reveal what the buyer was paying for: a call that gets through, understands a number spoken in a local language, follows the lender's rules and leaves a record that can be checked later.
“With Smallest, it never felt like a vendor and a client.”Vishal Handa, CTO, Kogta Financial
The lender's team could build agents through APIs, inspect performance in platform analytics and change a flow in hours, the case study says. Its compliance team kept collection rules in-house and logged calls for audit. That arrangement illustrates why voice AI in finance is harder than merely sounding natural. An eloquent machine that promises a borrower the wrong thing is worse than a clumsy one. The company also offers private and on-premise deployment for customers with tighter data requirements; its public pricing puts those features in custom enterprise plans.
03 / The businessBuy a voice, or buy the conversation
Smallest.ai has two routes to market. Developers can call its models directly: Pulse turns speech into text, Lightning turns text into speech, and both can be used inside someone else's application. A business that wants an agent can use Atoms for telephony, knowledge, campaigns, workflows and call review. The company also builds and supports enterprise deployments. That places it between model specialists such as ElevenLabs, Deepgram and Cartesia, and agent platforms such as Retell AI or Bland AI. A customer could combine separate vendors for each layer. Smallest.ai argues for the speed and control of building across them.
The public price list makes the split visible. As of this profile, Pulse batch transcription is listed from about $0.003 per minute, with real-time transcription around $0.004. Lightning v3.1 is listed around $0.175 per 10,000 characters. A self-serve agent call is shown around $0.09 to $0.21 per minute depending on its models. Telephony, knowledge-base use, privacy features and enterprise support may add costs. These figures are useful for a pilot budget, not a substitute for pricing a complete production workflow.
Published self-serve rates in September 2026. Enterprise terms are custom; end-to-end cost varies with the call design.
That model business has a different proof point in Pocket, a physical AI notetaker. Smallest.ai says Pocket tested more than ten speech-to-text models over three months before choosing Pulse. A notetaker needs to survive distance from the microphone, noise, changing speakers and mid-sentence language switches. Here Smallest.ai is just the transcription layer under someone else's product. It is a quieter kind of success than a branded voice agent, but arguably a cleaner measure of whether a model earns its place.
04 / The wagerSmall enough to keep up
In July 2026, the company announced a $13 million Series A led by Seligman Ventures, with Sierra Ventures and 3one4 Capital participating. Together with its earlier $8 million seed round, disclosed funding passed $21 million. It also introduced Hydra and what it calls a Voice 4.0 architecture. The bet is that a voice system should listen, reason and speak with some overlap, rather than treating each sentence as a completed form that must be processed before the next can begin.
Kamath has described a division of labor: let a small model handle the immediate conversation, and call a larger model when a question needs deeper work. It resembles a good receptionist who can answer the familiar question at once and knows when to ask someone else. Whether Hydra can make that division reliable across messy, high-stakes calls remains the more interesting question than whether a sample voice sounds human.
There is also a useful lesson for buyers outside AI. Start with one call type, define the permitted actions, connect the right records, test accents and interruptions, and measure complete-call outcomes. Kogta's case suggests that control over scripts and audit trails mattered as much as the voice itself. The method will struggle where the underlying records are wrong, the workflow has no clear handoff, or consent and data rules are left until launch. No latency figure can rescue an agent that confidently reaches the wrong account.
Smallest.ai began by chasing silence measured in seconds. The next contest is less easily measured: whether a caller finishes with the right answer, a completed action and no unpleasant surprise. That is the difference between a fast voice and a useful one.