ON THE RECORD
●G2i / AI training meets software engineering●May 2026 / Skills augmentation study published●210 TypeScript tasks reviewed / 87 met both top quality standards
Company / Artificial intelligence

G2i’s bet: AI needs better teachers

The company that once matched React developers with employers now builds the tasks, tests, and training environments behind AI models. Its most revealing finding: sometimes the exam is the problem.

Suppose a coding agent fails a software task. The verdict seems obvious: the agent needs to get smarter. But there is another possibility, rather embarrassing for the examiner. The instructions might be muddled. The test might reject a perfectly reasonable solution. The machine could be taking an exam that its human authors have not quite finished writing.

The useful bits
  • G2i turns experienced engineers’ judgment into AI training data, evaluations, and software environments.
  • Its TypeScript audit found 87 of 210 tasks had both clear specifications and top-rated tests.
  • Its published fee is 30% of contributor payments. The work still needs a well-defined scope.

That possibility sits near the heart of G2i, an engineering company whose business has moved from helping employers choose developers to helping AI labs choose better training signals. In both cases, the difficulty is deciding what competence looks like. A résumé can flatter. A benchmark can flatter, too. Somebody has to inspect the evidence.

The exam gets an exam

In its late-2025 research, G2i reviewed the TypeScript portion of Multi-SWE-bench. The team assessed 210 tasks, checking both their specifications and their tests. Only 87 received the strongest assessment on both dimensions. That is 41.4% of the collection. The other 123 did not all belong in the bin; they simply missed that combined standard.

A benchmark quality check / 210 tasks
87/210
41.4% met both standards58.6% missed at least one
A test of the tests. Clear instructions and good coverage must arrive together.

The process was deliberately human: two developers assessed each task independently, disagreements received further review, and senior engineers performed sampled quality checks. G2i estimated that restricting the benchmark to its verified subset could raise leading agents’ resolution rates by about seven percentage points. It was an estimate about changing the exam, not evidence that a model had improved.

Here is the lesson worth stealing. Before upgrading a model after a disappointing score, ask whether a competent engineer could understand the assignment and whether the grading accepts legitimate approaches. Difficulty and ambiguity are different expenses. Buying more intelligence to compensate for poor instructions is a peculiar way to balance the books.

From interviewing engineers to teaching models

G2i’s older business offers a clue to how it arrived here. It grew around React and JavaScript developers, using technical assessment and recorded interviews to help companies hire. Gabe Greenberg, its founder and CEO, made the problem sound less like recruitment than information retrieval: too many résumés, too little useful evidence.

A manager could inspect a candidate’s technical interview rather than schedule every conversation from scratch. Greenberg has described asynchronous responses, access to recorded interviews, and custom assessments. The useful asset was a network of engineers whose work and reasoning could be examined. That asset became relevant to a second market when AI companies needed people who could judge generated code.

G2i founder Gabe Greenberg speaking behind a laptop at AI Engineer Miami
The teacher has a microphone. Gabe Greenberg at AI Engineer Miami, where the software community meets the models learning its trade.

The company dates its frontier-lab RLHF work to February 2024. Reinforcement learning from human feedback turns human preferences into a signal for training. In coding, the person supplying that preference needs to understand more than whether the output looks plausible. They may need to recognise a fragile design, an omitted edge case, or an answer that satisfies the literal request while creating trouble elsewhere.

The connection between the two businesses is an interpretation, but a useful one. G2i already had practice assessing engineers. Its AI work asks those engineers to assess software and the tasks used to train software agents. The customer changed; the appetite for reliable evidence remained.

What the lab actually buys

G2i’s services cover several stages of post-training. Its RLHF pipelines collect comparisons with rationales and confidence scores. Reviewers work against shared rubrics, with golden examples and agreement checks intended to keep judgments consistent. A preference without a reason is a vote. A preference with a reason can become something a research team investigates.

“RLHF is only as good as the humans giving the feedback.”G2i’s product description

Supervised fine-tuning datasets supply worked demonstrations. G2i says authors receive the task, rubric, and known failure modes before writing; paired reviewers sign off before export. Delivery formats include JSONL and ChatML. For a lab, that is a practical distinction: material must fit a training pipeline as well as satisfy an editorial standard.

Evaluations address another question: how will the team recognise failure before production? G2i describes domain-specific tasks, replayable context, calibrated scoring, and traceable decisions. Its named domains include software, infrastructure, design, machine learning, and security. This is where a general-purpose answer meets the particulars of a real system.

The engineering work behind training
  1. 01Recreate the workTasks, interfaces, APIs
  2. 02Define successRubrics and validation
  3. 03Review the resultReasons and calibration
A useful loop starts with a task people can explain and ends with a decision people can inspect.

Reinforcement learning environments go further, recreating application behaviour so agents can act inside a repeatable setting. G2i describes realistic product surfaces and APIs, plus resets, checkpoints, telemetry, and batch execution. The point is to let an agent practise a workflow repeatedly while keeping the result measurable. A convincing screenshot alone cannot do that job.

Its seed-data service fills these environments with enterprise patterns and scrubbed application data. The company says it preserves useful structure and relationships while removing sensitive fields. This puts G2i between a talent supplier and an engineering delivery partner. Buyers can engage a specialist, an embedded team, or a managed project whose delivery G2i owns.

A fee you can put on a napkin

The pricing page offers an unusually concrete starting point: G2i charges a fixed 30% fee on contributor payments. If contributor payments total $10,000, applying that stated fee produces $3,000 for G2i and $13,000 combined. That is arithmetic illustrating the published fee, not a quotation for a particular environment or dataset.

Illustrative contributor-based calculation
$10,000Contributor payments
+$3,000G2i’s 30% fee
$13,000Combined illustration

For engineers, the talent FAQ lists typical rates of $50 to $200-plus an hour, varying with experience, skill set, location, and project. It also says prior AI or annotation experience is not required. Strong development experience, careful judgment, and clear writing matter. Passing vetting puts an engineer on the roster for upcoming projects; it does not itself specify a start date.

For a buyer, the real budget question is how many expert hours the task deserves. The published percentage makes one part of the bill legible. Scope still determines the work: how many tasks, which specialties, how many review rounds, and how much application behaviour must be reproduced. An elegant fee cannot rescue a shapeless brief.

The customer who wanted the contractors to stay

A published Shopmonkey case study makes the earlier service tangible. After an 18-month codebase rewrite, the automotive software company needed engineers to tackle accumulated bugs and stabilise its platform. G2i says it supplied 40 profiles that led to 20 hires, with recorded technical interviews helping managers reduce screening time.

The case study reports 443 pull requests in one third quarter, without identifying the year. More interesting than the volume is the change of intention: Shopmonkey initially sought short-term contractors, then moved toward integrating the engineers permanently after early results. This is G2i’s account, rather than an independent productivity audit, but it explains what the customer was buying: people who could join an existing workflow and contribute.

Lattice’s older mobile engagement provides a different unit of cost. G2i reports 880 hours over ten weeks for a React Native application project. The account describes shared development tools, automated testing, and contributions back to open-source modules. There is a family resemblance to the current AI work: delivery lives in the details of the system.

Teach the vocabulary before judging the answer

In May 2026, G2i published a skills-augmentation study covering 20 selected engineering and DevOps tasks and 1,820 scored runs across four models. For Claude Opus 4.6, the reported average pass rate rose from 37% without skill documents to 89% with them. The strongest pattern involved investigation tasks where documents supplied a specific vocabulary for classifying failures.

The selection matters. G2i chose tasks for reliable lift, consistency, and clean run histories. These results therefore describe that portfolio, not the gain any enterprise should expect. Still, the experiment suggests a sensible habit: give an agent the operational terms and rules your team uses, then test whether those instructions help on your own tasks.

G2i’s differentiation is its emphasis on experienced engineers, technical communities, and calibrated review. Its community work includes React Miami and AI Engineer Miami; its earlier developer-health initiatives addressed burnout and sustainable work. In an interview with ConTejas Code, Greenberg discussed autonomy, ownership, and the structure needed to run a remote business.

The model has conditions. Specialists need a task in their domain, reviewers need criteria they can agree on, and the environment needs to preserve the behaviour that matters. Otherwise, the buyer risks collecting expensive opinions about an ill-defined problem. The practical lesson travels well beyond AI labs: make the assignment clear, make the evidence inspectable, and ask whether the test deserves its authority.