The most useful part of a programmer’s day may be the part they wish nobody saw. A prompt misses the point. A tool returns an error. A test fails. Then someone changes two lines, tries again, and the thing works. The clean commit records the ending; the sequence records the intelligence. PublicAI, a San Francisco company that made its name buying human contributions for AI training, now wants to pay for that sequence.
Its new product, Trajector, collects coding sessions from developers using supported AI coding agents. The company’s case is plain: real work produces prompts, tool calls, diffs, tests, failures and repairs that staged exercises struggle to imitate. For an AI lab trying to teach agents how software work actually unfolds, the ugly middle may be more useful than the immaculate finish.
- PublicAI sells data collection and annotation services to AI builders, with contributors paid for accepted work.
- Its Mother Tongue campaigns gathered verified speech across 12 languages; Trajector now focuses on real coding sessions.
- The company says it raised $10 million and its Data Hub network includes more than 3.5 million workers.
- The practical test is whether consent, quality checks and rewards can survive at scale.
The microphone came first
Co-founder Steven Wong traces PublicAI’s beginnings to a research effort in 2022: his lab was assembling a large 3D image dataset and needed enough real-world material to make it useful. His answer was to recruit people and reward their contributions, with validators and financial penalties meant to discourage bad submissions. Wong and co-founder Kenji Narushima formed PublicAI in 2023. In the business that followed, the human contributor became part of the data supply chain rather than a nameless source behind it.

Mother Tongue made that supply chain easy to picture. A participant read phrases aloud in a native language, uploaded recordings through the Data Hub, and waited for approval. In its second season, the campaign listed 12 languages, from Arabic and Indonesian to French and Korean. Contributors could submit up to 240 phrases a day. Reviewers voted on recording quality. Those accepted clips could help customers build speech recognition and voice systems that cope with more than one standard accent.

The quality problem appeared in the rules. Season 3 offered $1 a day for 140 recordings in highlighted languages, roughly half an hour of work by PublicAI’s estimate. By Season 4, the process had tightened: uploaders staked 100 PUBLIC tokens for a 28-day lock, only designated judges voted, and the company warned that malicious work could lose the stake. Season 5 added a more detailed member survey and accent checks. PublicAI has not published a simple before-and-after quality score for those changes, but the rulebook itself shows where a crowdsourced data market feels pressure: false voices, careless labels and people chasing payouts without producing useful data.
“Real human work is the best training data and the best guardrail there is.”PublicAI, announcing Trajector in August 2026
A market has to clear three times
PublicAI describes an enterprise marketplace. A buyer asks for specialized data; contributors collect or label it; validators decide what passes. The company’s site offers a pay-as-you-go marketplace tier and enterprise services, charging for completed annotation work rather than each annotator on a team. Wong says datasets from the earlier research work reached customers including Adobe and LumaAI. PublicAI has also named AbakaAI in its materials. These are company-reported customer relationships, not disclosed contract terms.
An AI buyer specifies the data or evaluation it needs.
People record, annotate or contribute work they already made.
Screening and validation decide what is accepted and paid.
This is where its Web3 machinery earns its keep, if it earns it at all. Staking gives a contributor something to lose; crypto payments can reach workers across borders without traditional payroll rails. But tokens do not make a bad accent label accurate, and a vote is only as sound as the voters. PublicAI still has to do the unglamorous work of screening, review and dispute handling. Its own campaign revisions are a useful reminder that incentives need maintenance.
The money behind the venture is clearer than the price of any one enterprise dataset. PublicAI announced a $2 million initial raise and an $8 million Series A in June 2025. It named Saudi Telecom Group, Blockchain Builders Fund, NEAR Foundation and others among the later backers. It also joined Chainlink Build that year. Funding is a measure of investor appetite, however, not proof that every submitted voice or code trace found a buyer.
Then the data moved into the terminal
In August 2026, PublicAI announced the turn from voices to code. The premise of Trajector is that a developer should keep working normally while a local tool captures eligible coding-agent sessions. The developer opts in project by project, reviews what is uploaded, and earns rewards on sessions that pass checks for substance, duplication and sensitive data. Its public page says accepted rewards can be withdrawn as stablecoins after a $10 minimum. During the beta, the supported agents and models are limited, so a session outside those rules earns nothing.

The cost to a contributor is subtler than a token stake. Trajector says its client scrubs secrets before upload, yet its privacy policy says file paths, file contents and command output can remain in the session; paths may reveal a user name. The contributor keeps rights to their code, while the license granted for delivered data is perpetual and irrevocable. Deleting a session later cannot retrieve copies already sent to customers. These terms make the offer legible: a developer is being paid for a durable data right, not merely lending PublicAI a temporary glance at a work log.
That matters for buyers, too. A company looking for authentic coding traces wants evidence that the work was real and that the seller had the right to share it. Trajector’s rules restrict contributions to code the developer owns or eligible open-source work. Its acceptance checks reject trivial or synthetic sessions. None of this makes every legal or privacy question vanish, but it explains the product’s shape: the collection tool, the authorization and the quality gate are all part of what the customer is buying.
The scoreboard next door
PublicAI launched a different kind of product in September: the PublicAI Index. It does not run its own model tests. It pulls scores from public leaderboards, puts unlike measures on a common scale, discounts uncertain results and requires two independent publishers before a model enters the overall ranking. The site exposes its weights and offers a JSON API and MCP interface for agents. A new Decisions category measures smaller routing and classification tasks, the kinds of calls an application makes repeatedly while nobody writes a triumphant launch post about them.

The Index is adjacent to the data business rather than identical to it. It gives AI teams a way to compare models, while Data Hub and Trajector supply the human material those models learn from. Together they place PublicAI between the builders of AI systems and the people whose work teaches them. Its rivals in data work include Scale AI, Appen and Amazon Mechanical Turk. PublicAI’s distinction is the combination of a distributed contributor network, explicit rewards and an increasingly specific hunt for ordinary, authentic work.
What another builder can borrow
There is a practical idea here for anyone building a data product: recruit around a task people can understand, publish acceptance rules, pay for verified work, and revise the rules when participants discover their edges. Mother Tongue did this with a microphone; Trajector tries it with a terminal. A buyer should still ask for quality evidence, rights documentation and a clear account of what contributors were paid. A contributor should read what leaves the machine and what rights the payment purchases.
The experiment has limits. It works best where human experience is hard to synthesize, contributors can show they have permission to share it, and a reviewer can tell useful work from theater. It gets harder when the data is full of employer secrets, when verification costs more than the dataset is worth, or when rewards attract imitation faster than honest work. PublicAI’s own history is a small lesson in that difficulty. The company began with voices and found itself writing ever more detailed rules for who could speak. Now it wants the error before the fix. The next challenge is knowing whose error it really was.