Patronus AI began by catching a chatbot's bad answers. Now it is building practice worlds where agents can make their mistakes before a customer has to live with them.
A failed payment can masquerade as a change of heart. Global App Testing sends people into the markets, devices and conversations where software quietly lets users down.
The British testing company puts real people on real devices to find the faults that derail digital shopping. Its proposition: more eyes on a release, fewer surprises at checkout.
Nitrode is a San Francisco AI research company building the evaluation, data, and post-training layer for game development. It started in Y Combinator's Winter 2025 batch as a Godot-based AI game engine that let anyone prototype a 3D game in a day, then pivoted toward frontier research: it runs GameEngineBench, a benchmark that scores how well AI models handle real game-engine coding tasks, and produces expert-labeled datasets and reinforcement-learning environments (including Unreal Engine 5) for AI labs and enterprises. Nitrode's bet is that game development will be handled by systems of specialized models working together, not one general-purpose model.
Haize Labs is a New York-based AI safety and reliability startup that automates the red-teaming, stress-testing, and evaluation of large language models. Founded in 2024 by a trio of Harvard-trained researchers, the company builds algorithms that hunt for the inputs that make AI models misbehave - jailbreaks, failure modes, and edge cases - so they can be fixed before real users find them. Its 'haizing suite' and multi-turn attack engine Cascade are used by frontier model labs including OpenAI and Anthropic, alongside enterprises like Deloitte and MongoDB.
HoneyHive is a New York-based AI observability and evaluation platform that helps engineering teams ship reliable AI agents. Built OpenTelemetry-native, it unifies distributed tracing, offline experiments, real-time monitoring, and human-in-the-loop annotation into a single continuous-improvement loop so teams can see not just what an agent output, but how it got there and where it went wrong. Founded in 2022 by Columbia roommates Mohak Sharma and Dhruv Singh, HoneyHive is used in production by enterprises including Commonwealth Bank, NVIDIA, MongoDB, and Pinecone, and raised $7.4M in total seed funding led by Insight Partners.
Vellum is a New York-based AI development platform that helps companies build, evaluate, deploy, and monitor production-grade LLM applications and AI agents. Founded in early 2023 by three engineers who kept hitting the same wall - prototypes that demo well but break in production - Vellum gives teams a workflow builder, an evaluation suite, version-controlled deployments, and live monitoring so they can move AI from proof-of-concept to reliable production systems. The platform works with more than 150 companies including Drata, Redfin, Swisscom, and Headspace, and raised a $20M Series A led by Leaders Fund in July 2025.
Braintrust is the end-to-end developer platform for shipping AI products. It connects evaluations, observability, prompt iteration, and an agentic optimizer (Loop) into one workflow so engineering teams can measure, debug, and improve LLM applications in production. Teams at Notion, Stripe, Vercel, Airtable, Instacart, Zapier, Ramp, Dropbox, Cloudflare and BILL use it to ship AI that doesn't silently regress.
Snorkel AI is a Redwood City-based enterprise AI company spun out of the Stanford AI Lab in 2019. Its Data Development Platform lets enterprises programmatically label, curate, and evaluate the training and evaluation data that powers custom LLMs, RAG systems, and AI agents - turning expensive manual annotation into a software-engineering discipline.