Company profile micro1 turns expert work into training data • $35M Series A • Realm, Cortex and Robotics •

AI infrastructure / Company profile

micro1 Bet the AI Boom Would Run Out of Human Judgment

The young company began with an AI interviewer. It is now selling the scarce ingredient frontier models, workplace agents and robots all need: structured evidence of how capable people think and work.

The cleverest thing about micro1 is also the easiest to miss. The company uses an artificial-intelligence interviewer to find the humans who make artificial intelligence better. A software agent named Zara asks a candidate to explain a skill, solve a problem and talk through real work. The result is not merely a hiring score. It is the entrance exam for a distributed workforce of engineers, doctors, lawyers, financial analysts, writers and other specialists whose judgment can become training data.

That circular machine has carried micro1 far beyond its original recruiting pitch. Founded in 2022 by Ali Ansari, the San Francisco company now describes itself as a data lab for frontier AI. It recruits experts, organizes their assignments, checks the quality of their output, builds realistic environments for model training, evaluates workplace agents after they launch and captures human demonstrations for robots. The common ingredient is not a particular model. It is a system for turning difficult human work into structured, measurable evidence.

The bet reflects a change in the AI market. Early machine-learning pipelines often needed oceans of simple labels: identify the stop sign, draw a box around the pedestrian, mark a sentence positive or negative. Modern systems have already swallowed much of the public internet. Their remaining failures show up in longer, murkier tasks - reconciling a financial workbook, reasoning within the limits of a pathology report, following a company's exception-filled refund policy, or manipulating an unfamiliar object without dropping it. Those jobs require context and taste, not just clicks.

$35MSeries A announced in September 2025
$500MReported valuation attached to that round
100+Expert fields supported by its data engine

Three numbers and one necessary footnote: valuation and revenue claims in young private companies are snapshots, not audited destinies.

Zara interviews the interviewers

Zara remains the supply-side engine. Candidates apply for a role and complete an on-demand, conversational interview, usually lasting 20 to 40 minutes. Questions change with the job. A technical applicant may face a coding exercise; an annotator may complete a human-data task. The session is recorded, a skills report is produced, and human recruiters can review the result before matching the candidate to work. One skill generally gets about seven minutes of interview time.

The practical problem is speed. A frontier lab may suddenly need hundreds of people who understand Go, tax law or radiology, then need an entirely different group weeks later. Conventional recruiting is expensive and inconsistent at that tempo. Zara standardizes the first interview and lets micro1 search globally. The company says the system has recruited thousands of experts, including university professors, while adding new specialists each week.

micro1 has tried to measure whether the loop works. In a randomized field test involving about 37,000 applicants for junior-developer jobs, candidates who passed Zara later passed a blind final human interview 54 percent of the time. The comparable figure for a conventional résumé-screened group was 34 percent. Recruiters needed 44 percent fewer human interviews to find each hirable applicant. The study was produced by micro1 researchers, so independent replication would strengthen the claim, but its scale makes the result more useful than a polished demo.

A 20-point gap in micro1's randomized junior-developer field test. Humans still made the final decision.

Recruiting became a data factory

Finding an expert is only the first step. AI labs need the work specified, routed, reviewed and delivered in a format a training pipeline can consume. micro1's Realm offering bundles that operation. Flow is the workspace where experts create and review datasets. Merit tracks quality, velocity, error rates and cost. Zara supplies the people. Together they form a managed production line for finance, medicine, law, coding, science, visual language and audio projects.

The human-data loop
ZaraSource and assess domain experts
FlowCreate, review and deliver expert work
MeritMeasure quality, speed and reliability
RealmTurn work into evaluations and RL environments

A recruiter, a workbench, a scoreboard and an environment. Very Silicon Valley, but also rather like a well-run kitchen.

Realm points at another important shift: customers increasingly want environments, not static question sets. An environment gives an AI agent a task, source files, tools and room to act over many steps. The model may browse, calculate, edit a spreadsheet and produce a decision. That reveals whether it can finish the work, not merely select the right answer. micro1's public finance benchmark, for example, places models inside long-horizon assignments with workbooks, reports and open-ended briefs. Extraction can look competent while autonomous judgment collapses later in the chain.

For enterprises deploying their own agents, Cortex applies the same idea after the laboratory. A team defines a real workflow and what success means. Experts inspect failures, classify why they happened and create targeted examples to improve the agent. The process repeats as prompts, policies and underlying models change. One published case study describes an unnamed conversational-AI customer whose expert trust score reached 83 percent and whose perceived-intelligence rating rose from 6.0 to 7.5 over eight weeks. The anonymity limits what outsiders can verify, but the workflow is the real product: evaluation becomes an operating rhythm instead of a launch-day test.

What customers are buying

Not a generic pool of résumés. They are buying a completed loop: recruit specialists, define a hard task, collect their judgment, check the output, feed it into a model and measure what changed.

Then the data stood up and walked

Robotics makes the data problem wonderfully literal. A useful robot must learn how hands approach a drawer, how tools move through clutter and how a practiced worker notices that something feels wrong. micro1 collects first-person and third-person demonstrations in homes, factories and specialized workplaces. Its service can synchronize stereo video, inertial sensors and multiple cameras, then add temporal segments, object tracks, hand interactions and natural-language steps.

The company also operates controlled environments where robots can be deployed, observed and evaluated end to end. Its Flow annotation research describes a hybrid pipeline in which machine models propose labels and humans concentrate on corrections. In a small five-episode test published in July 2026, the configured pipeline required fewer human edits than a single-pass baseline. It is early evidence, but the design is sensible: automation handles the obvious frames while people resolve ambiguity and edge cases.

This is where micro1's recruiting history becomes useful. A generic crowd can label a cup. A mechanic can demonstrate a repair, explain what matters and spot an unsafe motion. A logistics worker knows which shortcut breaks the process. The more robots leave controlled demos, the more training data must reflect the long tail of actual work. micro1 is betting its expert network can supply that texture.

A marketplace wearing a lab coat

micro1 sells to AI labs, large technology companies, enterprises, government teams and robotics developers. Microsoft is the best-known publicly named customer; many engagements remain confidential. The business is a software-enabled managed service. micro1 takes responsibility for recruiting, project design, worker management, quality assurance and delivery, then earns revenue from the engagement and a margin on expert labor. Public pricing is not disclosed.

The model can scale quickly because contracts arrive before a permanent workforce has to be assembled. It can also be operationally heavy. Every new domain brings new rubrics, expert rates, compliance rules and quality disputes. A dashboard helps, but judgment is not a clean commodity. The company's advantage depends on matching the right person to the right task and learning, over time, whose output customers trust.

The alternatives

Scale AI, Surge AI, Mercor and Invisible Technologies overlap in managed human data. Snorkel AI and Labelbox sell tooling. Braintrust and Turing compete for skilled talent.

micro1's angle

One system spans AI-led vetting, expert work, quality telemetry, realistic agent evaluation and physical-world data. Its marketplace is the supply engine, not the whole product.

Investors rewarded the story in September 2025. micro1 raised a $35 million Series A led by 01 Advisors, with LG Technology Ventures participating, at a reported $500 million valuation. Former Twitter operating chief Adam Bain joined the board. Ansari said annual recurring revenue had risen from roughly $7 million at the start of 2025 to $50 million by September. By December, he told TechCrunch it had passed $100 million. Those are company-reported figures, and ARR is not the same thing as recognized or profitable revenue, but the acceleration shows how urgently labs were buying expert data.

That growth placed micro1 beside much larger rivals. Scale AI helped define the category. Surge AI and Mercor reported substantially higher revenue. The opening for micro1 is not that competitors ignored humans; they did not. It is that the unit of valuable work is moving upward, from annotation to evaluation, from generic workers to specialists, and from one-off datasets to continuous loops inside customer operations.

Can judgment become infrastructure?

Ansari frames micro1's mission with a grand question: where should humanity spend its time? The immediate business is more concrete. AI still needs people to show it what good work looks like, especially when the work is new, consequential or hard to grade. micro1 wants to identify those people faster, capture what they know and route the evidence to the systems that need it.

There are tensions built into that mission. An AI interview can standardize screening, but candidates still need transparency and recourse. Operational workflows can become valuable training data, but privacy and consent must survive anonymization. Expert work can command high rates today, while the stated purpose of the work may be to automate parts of the expert's profession tomorrow. micro1 says candidate data is not used to train its internal models and that clients control retention periods. The credibility of such controls will matter as the company handles more sensitive domains.

What makes micro1 interesting is not a promise that humans remain permanently in charge. It is the observation that capable AI creates demand for harder human input. Each new agent opens another evaluation problem. Each robot encounters another unfamiliar room. Each industry adds rules that were never written down. The company has built an expanding toolkit for collecting those answers. Its future depends on whether that toolkit can make judgment repeatable without sanding away the very context that makes it valuable.