llm-ops

(4)
Company
HoneyHive
Ai · Developer Tools · Saas

HoneyHive

HoneyHive is a New York-based AI observability and evaluation platform that helps engineering teams ship reliable AI agents. Built OpenTelemetry-native, it unifies distributed tracing, offline experiments, real-time monitoring, and human-in-the-loop annotation into a single continuous-improvement loop so teams can see not just what an agent output, but how it got there and where it went wrong. Founded in 2022 by Columbia roommates Mohak Sharma and Dhruv Singh, HoneyHive is used in production by enterprises including Commonwealth Bank, NVIDIA, MongoDB, and Pinecone, and raised $7.4M in total seed funding led by Insight Partners.

ai-observability · agent-observabilityRead →
Company
WAi Technologies
Ai · Saas · Enterprise

WAi Technologies

WAi Technologies is a Santa Clara-based digital engineering and AI services company that helps enterprises modernize legacy .NET applications and ship software faster. As a long-standing Microsoft partner, it pairs Azure, Dynamics 365 and Power Platform expertise with its own product line: the open-source Raaghu Design System and AI Pundit, an AI toolchain that converts Figma designs into production-ready front-end code. Its current bet is 'spec-driven' AI delivery - keeping humans in control of strategy while AI does the building.

wai-technologies · ai-punditRead →
Company
Datasaur
Ai · Saas · Enterprise

Datasaur

Datasaur builds secure, private AI infrastructure for regulated enterprises - starting with one of the most loved NLP data labeling platforms and evolving into LLM Labs, a workbench for building custom, on-prem ChatGPT-style assistants on a company's own data.

nlp · data-labelingRead →
Company
ZeroEval
Ai · Developer Tools · Saas

ZeroEval

ZeroEval is a New York-based AI startup from Y Combinator's Summer 2025 batch building an auto-optimizer for AI agents. Founded by Jonathan Chavez and Sebastian Crossa - two friends who met in college in Mexico - the platform captures every interaction your AI agent makes, scores quality with custom LLM judges, and automatically turns real production data into better prompts. The result: agents that get smarter after launch without manual intervention. Trusted by DoorDash, Datadog, Hugging Face, and Harvard Medical School, ZeroEval closes what the founders call 'the last mile reliability gap' in agentic AI.

ai-agents · llm-evaluationRead →