ON THE RECORD
DATACURVE · $15M SERIES A ANNOUNCED OCTOBER 2025DEEPSWE · 113 ORIGINAL ENGINEERING TASKSSHIPD · BOUNTIES FOR TECHNICAL WORK

AI / THE HUMAN INPUT

Datacurve and the price of a difficult answer

The internet has plenty of code. Datacurve pays people to produce the kind that teaches an AI what to do when the easy answers run out.

Imagine hiring a programmer and giving them an exam whose answer is hidden in the desk drawer. They find it, copy it, and receive a splendid score. You have learned something about the programmer. You have also learned rather more about the exam.

Datacurve has made a business out of the distance between those two conclusions. The San Francisco company supplies data for training and evaluating AI models, with coding as its original specialty. Its premise is deceptively practical: if you want a machine to handle difficult work, someone must first decide what difficult work looks like, do it properly, and establish whether the result deserves a pass.

THE STORY IN FOUR POINTS
  • The buyer: foundation model labs and companies building AI developer tools.
  • The supply: expert-created tasks, demonstrations, environments and evaluations.
  • The incentive: Shipd pays technical contributors bounties for selected work.
  • The public experiment: DeepSWE tests coding agents on original, longer engineering tasks.

There is a pleasing reversal here. The company is selling human expertise to businesses trying to automate human expertise. To get the supply right, it has to understand what makes a person want to spend an afternoon wrestling with a problem the machine cannot yet solve.

A planning app discovers its buyer

Serena Ge and Charley Lee met through computer-science classes and AI reading groups at the University of Waterloo. Ge had already built a climbing-training app in high school. Lee interned at Google. Together, they made UncleGPT, a planning tool that attracted 930 users, according to their Y Combinator launch account.

That is a respectable little audience. It is also a reminder that building something people try and discovering something businesses buy are separate accomplishments. Ge’s YC biography records three pivots before Datacurve. The useful detail is where the next problem came from: she had worked at Cohere on improving model reasoning and coding through synthetic data.

Waterloo’s account describes the shortage she encountered: companies needed better code data. Datacurve, founded in 2024 and part of YC’s Winter 2024 batch, would supply it. The founders moved toward a customer whose need they had seen from inside model-training work.

Datacurve co-founders Serena Ge and Charley Lee
Two founders, several pivots, one rather demanding customer: the next coding model. Serena Ge and Charley Lee. Photograph: Chemistry.

The developer is also a customer

Calling a skilled engineer a “data annotator” is an efficient way to make a potentially interesting job sound like detention. Datacurve’s contributor platform, Shipd, approaches recruitment differently. It organizes technical projects as quests and offers bounties. Its current public site presents software engineering, machine learning and data science under the names Mars, Eris and Monte.

Contributors choose work, while individual projects invite suitable people from the pool. Shipd describes a community reaching more than 40 countries. The product pitch combines a challenge, an opportunity to learn and a payout. A person supplying training data should have a reason to enjoy producing it.

“We treat this as a consumer product, not a data labelling operation.”Serena Ge · October 2025

The distinction affects what the company builds. In an October 2025 TBPN interview, Ge emphasized the experience that keeps contributors working through complex tasks, along with validations that steer them. Datacurve sells the final data rather than the labor. That leaves the company responsible for the machinery between a talented person accepting a task and a model lab receiving a useful deliverable.

For another product builder, the lesson is wonderfully ordinary: the people who create your inventory need a product too. Their experience affects what you have available to sell.

A receipt for expertise

By October 2025, Datacurve had distributed $1 million in contributor bounties, Waterloo reported. That figure is a tangible cost of collecting expertise, rather than a theoretical claim about its value. It does not represent the company’s total operating expense, and it does not tell a buyer what a dataset will cost.

The financing is clearer. A $15 million Series A led by Chemistry, announced October 9, followed a $2.7 million seed round. Together, those rounds totaled $17.7 million. Chemistry’s public investment announcement described multimillion-dollar contracts with leading AI labs. That is an investor’s account of commercial traction; it is not an audited revenue figure.

EXPERT WORK HAS A BILL$1M

Contributor bounties distributed by October 2025.

Datacurve occupies the supplier layer beneath the coding assistant. A foundation model lab can commission material to improve debugging or code explanation. A developer-tool company can seek examples tailored to code editing, design-to-code or repository changes. The company’s launch described these as tasks its data could support, rather than a list of consumer applications it had shipped.

The broader market includes expert-data providers such as Scale AI, Surge AI and Mercor. Datacurve’s distinctive proposition is its technical contributor product and coding focus. Internal data teams and synthetic generation are also options for a buyer. The decision turns on which approach produces the right evidence for the capability being trained.

Who grades the grader?

By 2026, Datacurve’s public offering included reinforcement-learning environments, off-the-shelf datasets, supervised fine-tuning demonstrations and agent trajectories. A trajectory records the route through a task: tool calls, checks, changes of direction and recoveries. The final answer alone can conceal most of the work.

Its public research product, DeepSWE, puts the same attention on evaluation. The benchmark contains 113 original tasks across 91 active open-source repositories and five programming languages. The accompanying paper argues that mining already-merged fixes creates two hazards: a model may have encountered the answer during training, and the original tests may reward one particular implementation rather than every correct solution.

Consider two programmers who implement identical visible behavior using different helper functions. A grader tied to one private helper can reject the other programmer’s perfectly serviceable patch. Conversely, a thin test can accept a patch that supplies the expected shape without supplying the requested functionality. A score becomes a measurement of the test’s preferences.

SIZE OF THE REFERENCE SOLUTION
SWE-Bench Pro
120
DeepSWE
668
Mean lines added in reference solutions, as reported in Datacurve’s May 2026 comparison. More code indicates scope here; it is not a score for better programming.

DeepSWE uses purpose-written behavioral verifiers. In the paper’s audit, an independent LLM judge disagreed with DeepSWE’s verifier on 1.4% of reviewed runs, compared with 32.4% for SWE-Bench Pro. These are company-reported findings from an automated judge, which can itself be wrong. Still, the question they raise belongs in every purchasing conversation: how much confidence should you place in the instrument measuring progress?

The second draft of the exam

Datacurve also revised its own instrument. DeepSWE v1.1, described in June 2026, grades a committed patch in a fresh, isolated container. It adds structured test reports, removes flaky tests on some tasks and fixes dependency drift. A model’s working environment and the environment deciding whether it succeeded are now separate.

That separation has practical consequences. Changes to the test framework in the agent’s workspace cannot simply manufacture a passing verdict. Missing tests become visible in the report. Future Git history is removed. The company’s update reports that aggregate scores and the ordering of leading models stayed close to the earlier version.

The limitations matter just as much. The benchmark uses a standardized agent harness, rather than reproducing every vendor’s native developer product. It selects maintained open-source repositories with at least 500 stars. Its five-language coverage excludes Java and C++. A team buying an assistant for a private Java monolith should test it on that monolith before treating a public ranking as a forecast.

Copy the question, then check the answer

There are two practical ways into Datacurve. A model team can use its datasets, demonstrations or environments to address a particular capability. A developer can explore Shipd’s contributor projects. An engineering buyer can examine DeepSWE’s tasks and trajectories to make a more informed evaluation, even without purchasing data.

The operating lesson travels further. Begin with the work you need done. Recruit people who can do it. Make participation attractive. Preserve their checks and recoveries. Then test whether your grader accepts a correct answer that happens to look different from your own.

Datacurve’s careers page describes a small team working long days with considerable autonomy. Behind the leisurely language of quests sits a demanding production business: contributors must stay engaged, examples must remain useful, and verification must survive inspection. Expertise only helps a model when the pipeline carries it intact.

The desk drawer is an excellent place to keep stationery. It is a terrible place to keep the answer to an exam you intend to trust.