Breaking
YC W25 · Mundo AI builds the world's largest multilingual training-data library 75% of the world speaks a language AI struggles with 4 founders from UBC, ex-Cohere, Hugging Face, AWS, Binance ~$660K revenue reported in 2025 with a small team Native speakers, not machine translation YC W25 · Mundo AI builds the world's largest multilingual training-data library 75% of the world speaks a language AI struggles with 4 founders from UBC, ex-Cohere, Hugging Face, AWS, Binance ~$660K revenue reported in 2025 with a small team Native speakers, not machine translation
Company Profile · Artificial Intelligence

Mundo AI Wants to Teach Machines the Other 6,000 Languages

AI writes flawless English and stumbles over almost everything else. Mundo AI, a Y Combinator W25 company, thinks the fix isn't a smarter model - it's better data, gathered by the people who actually speak the language.

Ask a modern AI model to draft a wedding toast in English and it hums along like a seasoned speechwriter. Ask the same model to do it in Uyghur, or Sundanese, or one of the several thousand languages that most of humanity actually speaks, and the confidence drains out of it. The words come out stiff, wrong, or invented. This is the gap Mundo AI decided to build a company around.

The premise is almost stubbornly simple. Models are only as good as what they read, and the internet's written record leans overwhelmingly English. Everything else - the grammar of a market town in Indonesia, the idioms a grandmother uses in Kazakh - is thin on the ground, badly transcribed, or missing entirely. Mundo AI's answer is to go get it, in person, from the people who speak it.

01 / The ProblemThe languages AI forgot

The company frames the stakes in one blunt line: AI models are great at English, but struggle with almost every other language. By Mundo AI's own reckoning, that leaves roughly three-quarters of the world's non-English speakers standing outside the AI economy, waiting for tools that half-work in their own tongue.

75%
of people speak a non-English language AI underserves
2024
Year Mundo AI was founded
W25
Y Combinator batch
4
Co-founders, all UBC alumni

Most labs paper over the gap with two shortcuts: synthetic data spun up by other models, and machine translation of English text. Both are cheap. Both, Mundo AI argues, leave a residue - the flat, slightly-off quality of language that was never really spoken by a person. Train on it long enough and the model learns a version of a language nobody actually uses.

"AI models are great at English, but struggle with almost every other language." Mundo AI, on why it exists

02 / The FixNative speakers, not machine translation

Instead of translating its way around the problem, Mundo AI sets up operations inside the countries where a target language actually lives. Native speakers collect, write, and annotate original material; the company's own software runs the collection, generation, annotation, and quality checks on top. The output is a library of datasets built to be culturally accurate rather than merely grammatically plausible.

Coverage: how AI training data is distributed today (illustrative)
EnglishHigh
Top 10 langsMid
Everything elseLow
The long tail of human language is where models go quiet. Mundo AI is building for the bottom bar. Chart is illustrative of the well-documented imbalance, not exact figures.

03 / How It WorksA supply chain for language

Think of it less as a website and more as a supply chain that ends in a dataset. The pipeline is deliberately unglamorous, which is rather the point - moats in AI tend to be built out of the work everyone else would rather skip.

STEP 01
Collect
Native speakers gather original material in-country.
STEP 02
Generate
New, novel data created rather than scraped or translated.
STEP 03
Annotate
Labeling and structure added via proprietary software.
STEP 04
QA
Quality assurance for cultural and linguistic accuracy.

04 / The BuyersWho's actually paying

Mundo AI's customers are AI labs and companies building foundation and language models - the teams that live or die on the quality of what they feed their systems. When those teams want a model that works in Bahasa Indonesia or Hausa rather than one that fakes it, buying purpose-built data is faster and cleaner than trying to scrape a language that was never fully written down. Specific customer names aren't public, but the shape of the buyer is clear: applied ML and research teams with a non-English problem.

The frontier labs get the headlines. The data companies get the leverage.

05 / The FoundersFour people who left good jobs

The founding team met at the University of British Columbia and arrives with an unusually literal fit for the mission - a mix of data engineering, quant finance, and, notably, native fluency in the languages the company serves.

Co-Founder & CEO
Jason Liao
Helped build a record-breaking fraud-detection model at Tsinghua University; led quant research at a $60B hedge fund. Went looking for non-English data, found none, and built the supply.
Co-Founder
Garreth Lee
The first Indonesian at Hugging Face; worked on pretraining data at Cohere and tokenization at Hugging Face.
Co-Founder
Kenneth Wu
Ex-quant at Canada's largest quant fund; former SWE at AWS and analyst at Ontario Teachers' Pension Plan.
Co-Founder
Naijide Anwaer
Among the youngest Platform PMs at Binance US; speaks four languages.

06 / The MarketPicks and shovels, not another model

Mundo AI has deliberately positioned itself one layer below the models everyone argues about. It doesn't want to win the race to build a smarter brain; it wants to sell what every brain eats. That puts it in the neighborhood of data companies like Scale AI, Surge AI, Appen, and Sama - but with a sharper wedge: language coverage as the product, not a feature.

Founded
2024
HQ
San Francisco / Vancouver
Funding
Seed · YC W25
Model
B2B data as a product

There are early signs the bet is landing. Third-party trackers put the company's 2025 revenue around $660K on a small team - modest in absolute terms, but a real number for a data business barely a year old. More telling is the shift in language on its own front page: from "multilingual training data" toward "the data layer for perceptual intelligence," a hint that the roadmap runs past text and into the messier, multimodal territory of images, audio, and everything a model perceives.

If you want to know where AI is underbuilt, look at the languages it ignores.

07 / What You Can Take From ItThe idea worth stealing

The founder story hides a lesson that travels well beyond AI: Jason Liao didn't set out to start a data company. He went looking for something he needed - non-English training data - couldn't buy it anywhere, and realized the absence was the opportunity. The product was just the thing he wished existed. For anyone hunting a neglected market, Mundo AI is a clean example that the biggest openings often look like the boring, unserved corners everyone else drove past.

#multilingual-ai#training-data#yc-w25 #non-english-nlp#data-annotation#ai-labs #native-speakers#machine-learning#data-infrastructure