"Data is no longer a barrier to good AI." A Paris startup's quiet bet that the unglamorous work - clean, expert-judged data - is what actually ships intelligent systems.
In the years before generative AI made "training data" a boardroom phrase, two engineers in Paris were already convinced the whole industry had its priorities backward. Everyone, they noticed, was arguing about models. Almost no one was arguing about the data feeding them. In 2018 they built a company on the difference. They called it Kili Technology, after Mount Kilimanjaro - the climb toward the summit, where the air is thin and only the best-prepared arrive.
Kili Technology is, in plain terms, a data company for the age of artificial intelligence. Its platform helps large organizations turn raw material - contracts, customer emails, medical images, video, drone and satellite imagery, chat logs - into labeled, structured, high-quality datasets that AI models can actually learn from. The pitch is deceptively simple: an AI system is only as good as the data behind it, and most enterprise data is messy, inconsistent, and untrustworthy until someone does the careful work of cleaning and annotating it.
That "someone," in the Kili model, is a combination of software and human expertise. The platform provides the annotation tools, the collaboration layer, and the quality controls. Businesses - or Kili's own managed workforce - provide the judgment. Together they produce the datasets that train, fine-tune, and, increasingly, evaluate machine-learning systems.
The two founders came at the problem from complementary angles. Edouard d'Archimbaud, the company's chief technology officer, had built one of Europe's most advanced corporate AI labs inside the bank BNP Paribas. He had seen firsthand how promising models stalled not because the algorithms were weak, but because the data was. François-Xavier Leduc, the chief executive, focused on turning that insight into a business large enterprises would pay for.
Their timing was patient rather than lucky. Kili launched its platform in July 2020 and, by the company's account, had already secured customer renewals by the end of that year - an early signal that the tool solved a real, recurring problem rather than a one-off curiosity. When generative AI erupted into public consciousness in 2023, the argument Kili had been making since 2018 - that data quality is the true bottleneck - suddenly sounded like conventional wisdom.
Kili's customer list reads like a directory of European industry: the software group SAP, aerospace names Airbus and Safran, the luxury house Louis Vuitton, the tiremaker Michelin, carmakers Stellantis and Renault, the consultancy Capgemini, and financial institutions including Crédit Agricole and LCL Bank, plus the insurer Groupe Covéa. These are not organizations that experiment casually. They are regulated, cautious, and protective of their data - which is precisely the point.
What draws them is less a single feature than a posture. Kili leans hard into enterprise requirements that flashier competitors sometimes treat as afterthoughts: security certifications (SOC 2 Type II, ISO 27001, HIPAA-aligned controls), role-based access, project-level data isolation, and the option to deploy on-premise or in a private environment so sensitive information never leaves the customer's control. For a bank or an aerospace supplier, that combination is often the deciding factor.
The problem Kili addresses is easy to state and hard to fix: enterprise AI projects fail more often in the unglamorous middle than at either end. The models exist. The ambition exists. What frequently goes missing is a reliable pipeline for producing labeled data that experts trust - and, more recently, for evaluating whether a model's outputs are actually correct. Kili positions itself in that middle, providing the workflows, the quality gates, and the real-time performance tracking that let hundreds of annotators work on a project without the whole thing dissolving into inconsistency.
For image and video work, that means bounding boxes, segmentation, and classification, with AI-assisted pre-annotation to speed the tedious first pass - including an integration with Meta's Segment Anything Model. For text, it means named-entity recognition, classification, and the kind of nuanced language annotation that trains conversational AI. For documents, it means OCR and layout analysis. And increasingly, it means something newer.
As the AI market shifted from training bespoke models to deploying large language models, Kili followed the money and the difficulty. The company expanded into LLM evaluation and reinforcement learning from human feedback, or RLHF - workflows where subject-matter experts rank model outputs, provide preference judgments, and rewrite responses to demonstrate correct behavior. In regulated fields, that expert judgment is not a nicety; it is the only way to certify that a legal, financial, or medical AI is behaving responsibly.
Kili's recent research and product work orbit a single argument: you cannot ship what you cannot evaluate. The company has published material on evaluation awareness, benchmark contamination, and "sandbagging" - the ways models can look better on tests than they are in practice - and released an LLM reasoning benchmark of its own. Its proposed answer is a "trust layer" that pairs automated triage (LLM-as-a-Judge) with structured human oversight, aimed at the enterprises whose GenAI projects keep stalling on the way to production.
Kili's credibility rests less on marketing than on where its people come from. The founding insight was forged inside a bank's AI lab, where the gap between a promising prototype and a production system is measured in regulatory scrutiny rather than demo applause. That heritage shows up in the product: audit-ready workflows, granular access controls, and a preference for traceable human judgment over black-box automation. The company frames its own operating principles around six values - candor-based teamwork, ambitious learning, ownership, data-driven decisions, pragmatic efficiency, and an obsessive focus on the customer.
The technical depth also shows in how Kili talks about evaluation. Rather than treating benchmarks as scoreboard trophies, the company publishes on their failure modes - how a model can appear to reason well while merely pattern-matching a contaminated test set, or how systems can "sandbag" by underperforming when they sense they are being examined. For a subject-matter expert reviewing a legal or medical model, that skepticism is the entire value proposition: someone has to catch the plausible-sounding answer that is quietly wrong.
The business model is straightforward B2B software-as-a-service: enterprises subscribe to the annotation and evaluation platform, with enterprise licensing and private-deployment options. That is complemented by managed workforce services, in which Kili takes over the entire annotation lifecycle - sourcing and training annotators, running quality assurance, and delivering finished datasets - for customers who would rather outsource the operational burden. It is a mix of recurring software revenue and higher-touch services, sold to a relatively small number of large accounts.
Kili competes in a crowded field. The best-known name is the American giant Scale AI; other rivals include Labelbox, SuperAnnotate, CloudFactory, Snorkel AI, and V7, alongside the ever-present option of enterprises building their own tools. Against that backdrop, Kili's differentiation is deliberately narrow: European by base and sensibility, security-first by design, and increasingly oriented toward expert-driven evaluation rather than commodity labeling. Where some competitors chase sheer scale and volume, Kili's argument is about trust, control, and the quality of human judgment in the loop.
It is, notably, a small company making that case - roughly 45 employees, with a third-party revenue estimate around $2.1 million and total funding north of $31 million. That modest headcount is part of the story. Kili has not tried to out-spend the field. It has tried to be right about where the field was heading, and to sell that conviction to customers who value being careful over being first.
The competitive geography matters here too. American labeling firms have raised and deployed vastly more capital, and they dominate headlines about frontier-model contracts. Kili's response has not been to match that scale but to occupy the ground those firms serve less naturally: European enterprises with strict data-residency needs, industries where an auditor may one day ask exactly how a model was trained and checked, and teams that would rather keep sensitive data inside their own walls than ship it to a third-party cloud. It is a smaller market by headcount of prospects, but a stickier one, and it rewards the security-first posture Kili chose early.
Whether the expert-data thesis carries Kili into its next chapter is an open question - one tied to how enterprises ultimately decide to govern the AI they deploy. But the founders' original wager looks less contrarian every year. The models kept getting bigger. The data kept mattering more. And a small team in the 12th arrondissement of Paris had said so from the start.
Collaborative annotation for image, video, text, PDF, and geospatial data with configurable ontologies, AI-assisted pre-annotation, and automated quality controls.
Named-entity recognition, classification, and language annotation to build datasets for LLMs and conversational AI.
Quality-focused RLHF workflows and RAG evaluation, letting subject-matter experts rank, judge, and rewrite model outputs.
Document layout analysis and OCR/IDP annotation for intelligent document processing.
End-to-end managed annotation projects: sourcing, training, and QA of expert annotators, delivered as a service.
SOC 2 Type II, ISO 27001, and HIPAA-aligned platform with role-based access, data isolation, and on-premise deployment.
Large enterprises across financial services, manufacturing, aerospace, luxury, and insurance:
Total raised: $31M+. Series A led by Balderton Capital.
Led by Balderton Capital, joined by existing backers Serena Capital and Headline, plus angels Olivier Pailhès (Aircall), Dimitri Sirota (BigID), and Chris Schagen (Contentful). Closed July 27, 2021.
"AI models dominated the conversation - but data quality was the real foundation for success."
d'Archimbaud and Leduc launch Kili on the thesis that data quality, not models, is AI's true bottleneck.
Kili launches its annotation platform in July and raises roughly $7M from Serena Capital and Headline.
Balderton Capital leads a $25M Series A to accelerate enterprise AI adoption; total funding passes $31M.
As generative AI takes off, Kili adds LLM evaluation and RLHF workflows to its platform.
Kili releases an LLM reasoning benchmark and repositions around expert-driven data curation.
Publishes a 2026 enterprise data-labeling guide championing expert human review for high-stakes AI.
The name. Kili is short for Kilimanjaro - a nod to the climb toward summit-grade data quality.
Banking roots. CTO Edouard d'Archimbaud built one of Europe's most advanced corporate AI labs at BNP Paribas.
Fast validation. Kili launched in July 2020 and had already secured customer renewals by year's end.
From luxury to orbit. The platform annotates everything from luxury-goods imagery to satellite and drone footage.
Video: browse Kili Technology's YouTube channel for product demos and founder interviews.