Profile Ankur Edkie builds the speech layer for conversational software Murf AI began in 2020 with three founders and ten voices

Person / Founder / Engineer

Ankur Edkie Is Teaching Software How to Speak

Ankur Edkie spent a decade building systems people rarely noticed. At Murf AI, he is applying that engineer's instinct to a more intimate interface - the human voice.

Ankur Edkie chose a strange problem to make simpler. A voiceover lasts a few seconds or minutes, yet producing one has traditionally required a small parade of decisions: script, actor, microphone, recording room, direction, retakes, editing, music, timing, export. Change one sentence and the parade may have to march again. Edkie and his future co-founders had each met some version of that friction in their work. They wondered why sound remained so stubborn when text, images, and video were becoming fluid.

Their answer became Murf AI. Edkie founded the company in October 2020 with Sneha Roy and Divyanshu Pandey, friends from the Indian Institute of Technology Kharagpur. Murf's own account of the beginning is unusually compact: three founders, ten voices, three early clients. The proposition was compact too. Type a script, choose a synthetic voice, adjust how it speaks, and produce audio without organizing a recording session.

This was before the current generative AI rush. Edkie later described speech as the “most untapped modality” of the moment. Images and video had mature editing languages. Speech was harder to modify, personalize, and scale. The founders saw a medium awaiting software.

3College friends who founded Murf
10Voices at the October 2020 start
$11.5mSeed and Series A funding announced by 2022

The education of an invisible-systems builder

Edkie's route to synthetic speech ran through products designed to disappear into ordinary life. He studied at IIT Kharagpur from 2006 to 2010. He then started his career at Travelocity, where he worked on a hotels-only product released across 20 countries and seven languages. A traveler clicking through rooms does not see the localization, catalog logic, or technical coordination beneath the page. The system earns its keep by making the complexity feel uneventful.

At Goldman Sachs, Edkie moved deeper into large-scale engineering. He led technology projects and became Vice President for AI and Blockchain Technology. The domains changed, but the habits traveled: build architecture that can carry serious traffic, account for users in different places, and make difficult machinery reliable enough to trust.

Murf made those habits audible. A generated sentence has to pronounce a product name, pause in the right place, and land with the intended pace. It may need to become a training module in several countries or an advertisement with multiple versions. The interface is simple only when the underlying system has absorbed the awkwardness.

“Voice AI needs to be linguistically and culturally aware.”Ankur Edkie, 2025

That sentence contains more engineering than it first appears to. A language is not a flat list of words. It carries region, rhythm, pronunciation, intonation, and the habit of switching mid-thought. In India, a conversation can move through English, Hindi, and a regional language without asking permission. A useful voice model has to follow the speaker, not force the speaker into a neat dropdown menu.

Murf AI product journeyA three-stage path from studio voiceovers to dubbing and real-time voice infrastructure. THE PRODUCT WIDENS 01 02 03 Voiceovers Dubbing Voice agents CREATE LOCALIZE CONVERSE
One stubborn workflow opened into a stack: create the voice, move it across languages, then make it respond in real time.

Follow the work, then follow the pull

Murf came out of beta in January 2021. Its early users made e-learning lessons, explainers, ads, audiobooks, podcasts, animation, and YouTube videos. These jobs shared a basic need: good audio without studio logistics. The company added ways to tune pitch, pauses, emphasis, speed, pronunciation, images, and background music. A voice generator grew into a workspace.

Capital followed the usage. Murf announced a $1.5 million seed round led by Elevation Capital, then a $10 million Series A led by Matrix Partners India in September 2022. At that point the company said its annual recurring revenue had grown more than twentyfold since the seed round and users had generated over a million voiceover projects. The money was intended for product development, more voices, and a larger team.

The early product also had an instructive constraint: the same tool had to serve a solo creator and a large communications team. The first user might want one polished narration before lunch. The second might need a library of approved voices, shared projects, review loops, and versions for several markets. Serving both pushed Murf beyond the usual text box and download button. Editing controls made the system approachable; collaboration and governance made it useful inside an organization.

That progression reveals Edkie's operating style more clearly than a list of features. He tends to describe AI through repeatable work rather than spectacle. In public conversations, his examples are concrete: updating a training script without recording the whole lesson again, producing local versions of an advertisement, or giving a developer a dependable response time. The pattern is to find where production gets stuck, remove the expensive handoff, and preserve room for a person to direct the result. The software handles repetition. The user keeps judgment.

The more consequential signal came from businesses. Enterprise customers needed distributed teams to review work together. They needed localization, security, deployment choices, and predictable performance. Murf developed Studio for production, Dubbing for localization, and APIs for developers. The humble audio file started behaving like a platform.

By 2025, Edkie was describing Murf as voice infrastructure. Its Falcon text-to-speech model was designed for real-time use, where a delay becomes part of the conversation. Murf reported model latency of 55 milliseconds and global time-to-first-audio around 130 milliseconds, with support for more than 35 languages. Those are company measurements, but the design priority is revealing: in a voice agent, timing is a product feature.

The conversational pause
InstantHuman patience starts noticing500 ms

A support agent that waits too long sounds lost. One that jumps in too quickly feels discourteous. Edkie's team avoided a large-language-model backbone for Falcon and chose a smaller, purpose-built architecture. The company has also discussed guaranteed latency for enterprise customers and infrastructure capable of handling many simultaneous calls. In this corner of AI, a millisecond is both a computer measure and a social cue.

The interface is the voice. The experience is everything that happens before the voice arrives.

Consent belongs in the stack

Synthetic speech carries a problem that hotel search and bank software do not: the raw material can sound like a person. A model may reproduce tone, accent, and identity with uncomfortable fidelity. Edkie's public position is that voice technology has to begin with permission. Murf says each voice is created with an artist's active consent, artists may withdraw that consent, and they earn royalties as their voice avatars are used.

“AI should augment, not replicate,” Edkie has said. The distinction guides Murf's voice-cloning rules and its pitch to enterprise buyers. Provenance, contracts, and control matter when generated audio can travel farther than its maker expects. For a company selling infrastructure, trust is less a slogan than a dependency. A customer integrating an API needs confidence that the service, the rights, and the data handling will remain stable.

01 / Consent

The artist participates and agrees before a synthetic voice is created.

02 / Control

The artist retains the ability to withdraw consent.

03 / Royalties

Usage of a voice avatar generates ongoing compensation.

This framework also explains Edkie's view that APIs can be unusually durable. Integrating voice into a product requires technical work, performance testing, legal review, and organizational confidence. Once price and performance fit, changing suppliers means reopening those decisions. Reliability creates switching costs, and responsible sourcing supports reliability.

The next layer is conversation

Murf's path mirrors a recurring pattern in software. A company begins by automating a visible task. Customers then expose adjacent needs. The product becomes a system, the system becomes an API, and the API becomes infrastructure. Edkie has said Murf is working toward an end-to-end speech-to-speech platform and an agent builder developed with enterprise teams. The aim is a programmable speech layer that can serve different regions, run inside private environments, and move naturally between languages.

The challenge is no longer only whether a synthetic voice can pass for natural in a finished clip. It is whether it can listen and answer at conversational speed, pronounce the local place name, respect the permissions behind the voice, and keep working through thousands of simultaneous exchanges. Each requirement pulls Edkie back to his familiar territory: architecture hidden beneath an experience that should feel ordinary.

There is a pleasing loop in that career. At Travelocity, he helped software travel across countries. At Goldman Sachs, he built systems meant to hold up under institutional demands. At Murf, he is asking software to carry something more expressive across the same distances. The output is audible. The craft remains mostly invisible.

For years, computers waited for people to learn their interfaces. Voice reverses the bargain. The machine has to learn timing, language, tone, and restraint. Edkie's bet is that this transition will be built through patient technical choices rather than a theatrical imitation of humanity. Smaller models. Faster replies. Clear permissions. More languages. A platform that speaks only after a great deal of quiet engineering.