In the stream
01 From LogDNA to data in motion02 A 20% target became 30%03 The pipeline met production04 Now the agent reads the evidence

Company profile / Observability

Mezmo and the Art of Keeping Less

The company that once made logs easier to find now asks a more awkward question: why keep so many in the first place?

The most expensive sentence in software may be, “Keep it, just in case.” It sounds prudent. In a cloud system, it can mean keeping millions of lines that say a service is healthy, a DNS lookup succeeded, or a container did exactly what it did a second ago. The rare failure hides in this abundance. The storage bill arrives with admirable punctuality.

Mezmo built its first business by helping engineers search that pile. As LogDNA, founded in 2015 by Chris Nguyen and Lee Liu, it gave developers a central place to read logs from increasingly scattered applications. The company later made a different bet: the valuable decision happens before the pile is stored. Its telemetry pipeline can filter, sample, enrich, transform and route logs, metrics and traces while they are moving. Its newer AI tools attempt to use that prepared data to investigate incidents.

The short version
  • Mezmo started as LogDNA, a cloud log management product born from a startup pivot.
  • Its pipeline gives teams rules for what telemetry to keep, summarize and send to each tool.
  • In its own deployment, Mezmo reported cutting stored logs by an average of 30% after aiming for 20%.
  • AURA, its open source agent harness, extends the argument from data handling to incident investigation.

A logging tool escaped its original job

Nguyen and Liu were working on a different startup during Y Combinator’s winter 2015 batch when they built a logging system they needed themselves. Other founders wanted it. The side tool became the company. That origin matters because logging is one of those problems that looks solved until a system grows: a few servers become clusters, clusters become services, and a useful record becomes a torrent. What developers need at 2 a.m. is a clue. What the infrastructure produces is an autobiography.

LogDNA found an audience. IBM Cloud used its technology for log analysis and activity tracking, and the company named customers including Asics, Lime and Sysdig. By 2021, LogDNA reported about 3,000 customers. A $25 million Series C in 2020 brought Tucker Callaway into the CEO role; a $50 million Series D led by NightDragon followed in December 2021. The latter round explicitly backed a move toward data in motion. When LogDNA became Mezmo in 2022, the name announced a product decision: a log search box was no longer a sufficient description.

2015LogDNA founded
$50mSeries D, 2021
30%Average stored log reduction in Mezmo's own test

The editor between the machine and the bill

A telemetry pipeline is an editor with a very tight deadline. It sees a stream from applications, Kubernetes clusters, collectors and cloud services. It can recognize a repeated health check, strip a needless field, turn a log event into a metric, attach context such as service ownership, and send different results to different destinations. The destination might be Mezmo’s own Log Analysis, Datadog, Prometheus, a security system or storage. That makes Mezmo both an observability vendor and an infrastructure layer for other observability vendors.

The distinction is practical. If a team pays a downstream service to ingest each gigabyte, removing duplicate chatter before ingestion can lower that charge. If an engineer needs live detail during an incident, the pipeline can route that signal elsewhere. Mezmo says pricing for its own platform depends on data volume and retention needs, and that its AI root cause analysis is included in the platform license. Exact customer prices are negotiated or otherwise unpublished, so the more revealing numbers are the ones customers and engineers report from actual deployments.

AuditBoard, for example, is featured on Mezmo’s customer page with a reported 52% reduction in Datadog costs after filtering and aggregating logs before Datadog ingestion. The same page reports a 40% faster incident resolution and 90% fewer noisy alerts for that deployment. Those are customer results presented by Mezmo, not universal savings forecasts. Netlink Voice reports a 50% drop in overall telemetry volume after filtering and parsing data. Different systems give the editor different drafts.

Mezmo governance interface showing a proposed production action and an approval requirement
Even a machine with a plan has to ask permission. Mezmo's governance interface shows a proposed canary restart waiting for approval.

The dog food bit back

Mezmo’s most persuasive example is less tidy than a customer testimonial. Its own SRE team routed logs from a global Log Analysis account through the new pipeline. The goal was to cut stored data by at least 20% and move existing exclusion rules. Then the team looked at what was arriving. CoreDNS alone represented roughly half the traffic in one flow. They sampled that source down by 90% and converted some DNS log information into Prometheus metrics, preserving a useful count without keeping every line. Mezmo reported an average 30% reduction in stored logs after the work.

That was not the end of the internal lesson. In a separate move to collect metrics through OpenTelemetry, Mezmo said its nodes were sending about 600 to 700 MB per second into Pipeline. Lower environments had behaved. Production did not. Queues filled, and ordering guarantees concentrated data in a few Kafka partitions, blocking the flow. Engineers worked through the issue and continued the migration. The account is useful precisely because the company names the first failure. A pipeline that governs everyone else’s traffic has to survive its own.

“The data consumer must be able to capture the real-time value of data in motion, not just data at rest in storage.”Tucker Callaway, 2021 company announcement

Now the first reader is an agent

In 2026 Mezmo pushed the same premise into AI operations. AURA is an open source, Rust-based harness for agents that can use tools and investigate production systems. Mezmo describes configuration files that define models, prompts and tools, alongside records of execution that engineers can inspect. Its commercial platform supplies prepared telemetry to the agent. The theory is plain: a model cannot reason well about an incident if the evidence is scattered, noisy or missing.

There is an important line between a demo and a shipped workflow. In an August 2026 walkthrough, Mezmo’s VP of Research said its combined OpenTelemetry source, service map, anomalous trace view and Explain Trace were shipping. He also said AI Investigations, which starts from a pipeline alert, was running behind a per-account feature flag. A September post described that investigation flow: an alert triggers AURA to assemble evidence, confidence and a probable cause before a human starts from a blank screen. This is the newest part of the story, and customers should ask to see it in their own accounts.

Mezmo’s competitive position is unusual. Datadog, Dynatrace, Elastic and Splunk can be places to analyze data. Cribl is a familiar alternative for routing it. Mezmo wants to sit at the traffic junction and, increasingly, provide an agent to read the cleaned stream. It also sends data to tools a buyer already owns. That makes it attractive when a team has several destinations and wants to control their inputs. A small, simple stack may have little reason to add a pipeline. A team that cannot say which events it can safely sample should measure carefully before discarding them.

A rule worth copying

The lesson travels beyond Mezmo’s software. Start with one noisy source. Count its volume. Separate failure signals from routine repetition. Keep a way to test the decision during a real incident. Only then make the filter permanent. Mezmo’s CoreDNS episode supplies the memorable number, but the method is more useful: edit with evidence, and remember that the system that edits the data is itself a production system. At 2 a.m., a smaller, better annotated record can be worth more than a warehouse of everything.