# ZeroEval

> ZeroEval is a New York-based AI startup from Y Combinator's Summer 2025 batch building an auto-optimizer for AI agents. Founded by Jonathan Chavez and Sebastian Crossa - two friends who met in college in Mexico - the platform captures every interaction your AI agent makes, scores quality with custom LLM judges, and automatically turns real production data into better prompts. The result: agents that get smarter after launch without manual intervention. Trusted by DoorDash, Datadog, Hugging Face, and Harvard Medical School, ZeroEval closes what the founders call 'the last mile reliability gap' in agentic AI.

- **Founded:** 2025
- **Headquarters:** New York, NY, USA
- **Founders:** Jonathan Chavez (Co-Founder), Sebastian Crossa (Co-Founder)
- **Team size:** 2
- **Products:** ZeroEval Platform, Calibrated LLM Judges, Autotune, MCP Server, Tracing & Monitoring
- **Notable:** Y Combinator Summer 2025 batch (YC S25), Built llm-stats.com - reached 60,000 MAU and 333,000 total unique users, SOC 2 Type II compliant

## Products & services

- **ZeroEval Platform** — An auto-optimizer for AI agents that captures every LLM call, tool call, and session in real time, scores quality with calibrated LLM judges, and automatically generates prompt, model, and code improvements from failure patterns.
- **Calibrated LLM Judges** — Custom evaluation judges that learn from human feedback corrections and improve alignment scores over time - unlike static LLM judges, these are task-aware and continuously calibrated.
- **Autotune** — Automatic evaluation across multiple models and prompt optimization based on minimal human feedback samples. Runs evals on dozens of variants simultaneously and deploys the best version.
- **MCP Server** — A Model Context Protocol server compatible with Cursor, Claude Code, and 30+ AI coding agents - lets agents query ZeroEval directly to understand where they are failing and improve their own behavior.
- **Tracing & Monitoring** — Real-time logging of LLM calls with minimal overhead, automatic PII redaction, built-in judges for hallucinations, safety, and user frustration detection.

## Achievements

- Y Combinator Summer 2025 batch (YC S25)
- Built llm-stats.com - reached 60,000 MAU and 333,000 total unique users
- SOC 2 Type II compliant
- Trusted by DoorDash, Datadog, Hugging Face, and Harvard Medical School
- 18% accuracy improvement cited across 2,400 traced runs with 127 feedback instances
- MCP server integrated with Claude Code, Cursor, and 30+ AI agents

## Latest updates

- **2025-08** — Admitted to Y Combinator Summer 2025 batch; Jonathan and Sebastian photographed at YC campus in August 2025
- **2025-07** — Launched on YC with product announcement: 'Build Self-Improving Agents'
- **2025** — Launched MCP server enabling coding agents like Claude Code and Cursor to query ZeroEval directly for self-improvement

## Links

- Website: https://zeroeval.com
- LinkedIn: https://www.linkedin.com/company/zeroeval
- Twitter/X: https://x.com/zeroeval
- GitHub: https://github.com/zeroeval

---

Profile page: https://yespress.io/zeroeval
Published by YesPress — https://yespress.io
Last updated: 2026-04-15
