ai-reliability

(12)
Legend
Carl Bate Wants AI to Show Its Work
Founder · Engineer · Executive

Carl Bate Wants AI to Show Its Work

After three decades translating between software, strategy and human behavior, the Evidium founder is pursuing a stubborn idea: consequential AI should make its reasoning inspectable.

carl-bate · evidiumRead →
Legend
Dan Roth
Founder · Executive · Operator

Dan Roth

Dan Roth is the co-founder and CEO of Scaled Cognition, a Mountain View AI company building reliable, hallucination-free models for enterprises. Its flagship Agentic Pretrained Transformer (APT) tops agentic benchmarks and powers customer-service agents that follow policy and take action where accuracy matters, such as banking, insurance, and telecom. In June 2026 the company raised a $100 million Series A led by Khosla Ventures at a roughly $750 million valuation. Roth is a repeat founder: he co-founded Semantic Machines, an early conversational-AI company acquired by Microsoft in 2018, where he then served as corporate vice president of conversational AI, and earlier built Voice Signal Technologies, acquired by Nuance.

dan-roth · scaled-cognitionRead →
Company
Scaled Cognition
Ai · Enterprise · Saas

Scaled Cognition

Scaled Cognition is a Mountain View AI company building what it calls the first Agentic Pretrained Transformer (APT), a model designed and trained specifically for taking reliable, policy-adherent actions rather than just predicting language. Founded by the team behind Semantic Machines (acquired by Microsoft in 2018), it targets enterprise customer-experience workflows where mistakes and hallucinations are unacceptable - financial services, healthcare, telecom, and insurance. In June 2026 the company raised a $100M Series A led by Khosla Ventures to expand its research team and accelerate enterprise deployments.

agentic-ai · agentic-pretrained-transformerRead →
Company
ActionAI
Ai · Enterprise · Saas

ActionAI

ActionAI is a New York- and Israel-based enterprise software company building reliability infrastructure for AI. Its platform lets companies build, run, evaluate and improve mission-critical AI workflows with a human in the loop, mapping data across the AI stack for granular testing and using an 'Explainable Exceptions' system to catch model errors before they cause damage. Founded by Stanford-trained engineer Miriam Haart, the company raised a $10M seed round in April 2026 to serve regulated sectors like finance, insurance, legal and logistics where accuracy is non-negotiable.

ai · enterprise-aiRead →
Company
Haize Labs
Ai · Developer Tools · Enterprise

Haize Labs

Haize Labs is a New York-based AI safety and reliability startup that automates the red-teaming, stress-testing, and evaluation of large language models. Founded in 2024 by a trio of Harvard-trained researchers, the company builds algorithms that hunt for the inputs that make AI models misbehave - jailbreaks, failure modes, and edge cases - so they can be fixed before real users find them. Its 'haizing suite' and multi-turn attack engine Cascade are used by frontier model labs including OpenAI and Anthropic, alongside enterprises like Deloitte and MongoDB.

ai-safety · red-teamingRead →
Company
HoneyHive
Ai · Developer Tools · Saas

HoneyHive

HoneyHive is a New York-based AI observability and evaluation platform that helps engineering teams ship reliable AI agents. Built OpenTelemetry-native, it unifies distributed tracing, offline experiments, real-time monitoring, and human-in-the-loop annotation into a single continuous-improvement loop so teams can see not just what an agent output, but how it got there and where it went wrong. Founded in 2022 by Columbia roommates Mohak Sharma and Dhruv Singh, HoneyHive is used in production by enterprises including Commonwealth Bank, NVIDIA, MongoDB, and Pinecone, and raised $7.4M in total seed funding led by Insight Partners.

ai-observability · agent-observabilityRead →
Company
Distributional
Ai · Enterprise · Saas

Distributional

Distributional is an enterprise AI testing platform that uses adaptive, statistical analysis of production traces to catch behavioral drift and regressions in AI agents and LLM applications. Founded in 2023 by former Intel AI software GM Scott Clark, the company helps Fortune 500 teams ship generative AI with confidence by treating reliability as a continuous test problem rather than a one-shot benchmark.

ai-testing · ai-reliabilityRead →
Company
Openlayer
Ai · Saas · Enterprise

Openlayer

Openlayer is an AI governance and observability platform that helps enterprise teams test, monitor, and trust their machine learning and LLM systems across the full lifecycle - from prototype to production. Founded by ex-Apple, Amazon and Harvard alumni and backed by Y Combinator and Race Capital, the company raised a $14.5M Series A in May 2025.

ai-observability · ml-monitoringRead →
Legend
Leonard Tang
Founder · Executive · Engineer

Leonard Tang

Leonard Tang is the co-founder and CEO of Haize Labs, a New York AI safety startup that stress-tests frontier language models by automatically jailbreaking them before they reach the public. A Harvard math and computer science graduate who turned down a Stanford PhD, Tang built Haize into a venture that frontier labs like OpenAI and Anthropic pay to break their models on purpose. By 2024 the company had raised a $12.5M seed led by General Catalyst at a $100M valuation, roughly eight months after founding, and Tang landed on Forbes' 2025 30 Under 30 list for AI.

leonard-tang · haize-labsRead →
Legend
Mohak Sharma
Founder · Executive · Operator

Mohak Sharma

Mohak Sharma is the co-founder and CEO of HoneyHive, a New York-based AI observability and evaluation platform that helps companies ship reliable AI agents to production. Built OpenTelemetry-native, HoneyHive lets engineering teams trace, test, and monitor multi-step AI pipelines and answer the deceptively hard question of what their agents are doing and why. Sharma started the company in 2022 with his former Columbia roommate Dhruv Singh, and by 2025 had raised $7.4M in total funding, including a $5.5M seed led by Insight Partners, while serving Fortune 100 customers in insurance and banking.

mohak-sharma · honeyhiveRead →
Legend
Scott Clark
Founder · Executive · Engineer

Scott Clark

Scott Clark is the Co-founder and CEO of Distributional, an enterprise AI testing platform that helps companies identify behavioral drift and unknown failures in AI systems. A triple-degree graduate of Oregon State University with a PhD in Applied Mathematics from Cornell, he co-founded and sold SigOpt (a Bayesian optimization platform backed by Y Combinator and a16z) to Intel in 2020, then led 200 engineers as VP & GM of AI/HPC Supercomputing at Intel before founding Distributional in September 2023. The company raised $30M in under a year, including a $19M Series A led by Two Sigma Ventures in October 2024.

ai · machine-learningRead →
Company
ZeroEval
Ai · Developer Tools · Saas

ZeroEval

ZeroEval is a New York-based AI startup from Y Combinator's Summer 2025 batch building an auto-optimizer for AI agents. Founded by Jonathan Chavez and Sebastian Crossa - two friends who met in college in Mexico - the platform captures every interaction your AI agent makes, scores quality with custom LLM judges, and automatically turns real production data into better prompts. The result: agents that get smarter after launch without manual intervention. Trusted by DoorDash, Datadog, Hugging Face, and Harvard Medical School, ZeroEval closes what the founders call 'the last mile reliability gap' in agentic AI.

ai-agents · llm-evaluationRead →