In 2023, Inworld AI released a detective game called Origins. The player arrived after an explosion in a future city and questioned its inhabitants. A police officer might discuss breakfast. A suspicious man might confess. Ask about somebody's height, and the performance could wobble. This was the charm and the hazard of an AI character: it could answer a question the writer had never scripted, including a question the writer would rather it ignored.
- What it makes: realtime speech, model routing and orchestration tools for interactive apps.
- Who buys it: game studios and developers of language tutors, companions and voice agents.
- What changed: a character demo became a business about speed, control and cost at scale.
The company, founded in 2021 by Ilya Gelfenbeyn, Michael Ermolenko and Kylan Gibbs, had a reasonable pedigree for this experiment. Gelfenbeyn and Ermolenko worked on API.AI, which became Google's Dialogflow. Gibbs had worked on conversational and generative language products at DeepMind. Inworld's first pitch was that game characters could possess a memory, a personality and some freedom to improvise. For players raised on dialogue trees, the idea was irresistible: perhaps the tavern keeper could finally remember your name.

The bill arrives before the sequel
A studio can make one dazzling encounter. A live game must make the next million encounters affordable. Every reply asks the system to hear or read the player, consult a language model, honor the fictional world and speak in a voice that belongs there. If each stage uses a different supplier, the developer must manage the joins, delays and bills. A character who takes too long to answer feels less like a person than a call center queue.
The arithmetic became vivid in Death by AI, a social survival game from Playroom and Little Umbrella. Its creators said it reached 700,000 daily users three days after launch and 20 million players within three months. A success of that size is also an unusually efficient machine for generating invoices. CEO Tabish Ahmed said the team had underestimated how much people would play, and therefore how much its original OpenAI and ElevenLabs setup would cost. When it tested cheaper models, players noticed changes in the game master's humor almost immediately. Some alternatives created safety problems. The product could not simply become cheaper by becoming duller.
We needed to be able to offer this at scale and remain profitable.Tabish Ahmed, Playroom and Little Umbrella
Inworld worked with the team on a custom model service and integration. The creators say the switch helped the game reach profitability. That is a customer account, not a public audit of its margins, but the mechanism is clear: model choice, latency and usage cost were now game-design concerns. A funny AI host was the feature. Keeping that host solvent was the business.
From character engine to conversation engine
Inworld still works in games. It has described collaborations with Xbox on narrative tools, with NVIDIA on responsive characters, and with Ubisoft on a prototype called NEO. But its catalog now reads like a parts list for any application that talks: streaming speech recognition, text-to-speech, a persistent realtime API, a router for language models, hosted inference and dedicated compute. Its Runtime coordinates models, memory, knowledge and tools. The customer chooses the application; Inworld wants to supply the moving parts that let it respond.
Each stage adds delay and cost. The trick is keeping the whole turn useful, quick and affordable.
The move is less surprising when you consider the original job. An NPC has to react in real time, respect a world, stay in character and perform at whatever scale the game attracts. Those are extreme versions of the problems faced by a language tutor or a customer-service voice agent. A language lesson fails when the tutor pauses so long that the student forgets the sentence. A companion app fails when a habit of long conversations turns retention into a cost problem.
Talkpal, which says it serves more than 10 million language learners, chose Inworld's voice technology for AI teachers. In a four-week A/B test reported by the two companies, text-to-speech cost fell 40%, feature use rose 7% and retention rose 4%. The case is more interesting than a demo in which a voice merely sounds nice. It suggests the voice had to work across languages, at a rate the company could keep paying, while learners kept showing up.
A voice with a meter attached
Inworld's public prices make its wager unusually legible. Its on-demand Realtime TTS-2 rate is listed at $25 per million characters; streaming speech recognition is $0.15 an hour. Larger plans lower the unit rate, and enterprise terms can lower it further. The LLM router advertises access to more than 220 models with no markup on routed third-party model prices. These are published rates as of this profile, not a promise about anyone's final bill. The amount of talking, model selection, plan and compute arrangement still decide that.
The product is also being tuned for the details a script cannot settle. Realtime TTS-2 can take plain-language direction about delivery and use previous audio turns as context. It can preserve a voice identity while switching languages. The company offers voice cloning from a short clean recording and a voice-design option built from a written description. For a game studio, that can mean a character who sounds worried when the scene requires it. For a tutor, it can mean a consistent teacher voice across several languages. Both use cases require permission and sound judgment when cloning a person's voice.

Inworld's distinction is therefore its attempt to own and coordinate more of the conversation than a voice-only vendor. A developer could assemble a stack from a speech provider, a model API and an orchestration framework. That approach may offer more freedom or better economics for a particular workload. Inworld's pitch is that integration and optimization across the stages can save engineering time and smooth the jagged edges that users hear. It is a claim best tested with real traffic, because the most beautiful sample voice is useless if the application cannot afford its best customers.
There are places where this play is unnecessary. A game with carefully written dialogue and no free-form conversation may gain little from a live model. A low-volume app with a good existing speech stack might not recover the effort of changing it. A product that must run entirely offline, or keep every component under its own control, needs a different architecture. The useful lesson to copy is narrower and sharper: measure the entire conversation, not only the voice sample. Test cost per active user, first-audio delay, response quality after many turns and the time engineers spend keeping the pieces together.
What the detective really discovered
The Origins detective had a simple tool: ask another question. That is also how a developer should approach this market. How long does the first audio take on a slow connection? What does a heavy user cost per month? Does the character maintain its voice, humor and boundaries after fifty turns? Can a model be swapped without rewriting the application? Inworld's own customer stories show why those questions matter more than the magic trick of a single unscripted sentence.
The company's route from AI characters to AI infrastructure is a small parable about software. A demo teaches the audience to want the impossible. A business has to make the impossible repeatable. Inworld started by giving a game character something to say. Its current challenge is to make sure that when millions of people answer back, the conversation can go on.