Latest

Company / Developer tools

The Bug That Refused to Reappear

Lightrun built a business around an awkward truth of software: the failure you most need to understand often appears only after the code is live. Its answer is to ask the running program for evidence, one line at a time.

There is a peculiarly cruel variety of software bug: the one that occurs only when real customers, real data, and several impatient services meet in production. At a desk, the engineer can read the code. In staging, the tests can pass. In the live application, something goes wrong anyway. The usual response is to add a log line, build a new version, deploy it, and wait to see whether the clue appears. This is detective work conducted by mailing a question to the witness and hoping for a reply tomorrow.

The short version

  • Lightrun adds logs, snapshots, metrics, and traces to running software on demand.
  • Engineers use it from an IDE; newer AI tools use the same runtime evidence to investigate incidents.
  • Its value is clearest when a bug cannot be reproduced locally and another diagnostic deployment is expensive.

Lightrun exists because Ilan Peleg and Leonid Blouvshtein had endured that ritual themselves. The two developers founded the company in 2019, after an earlier joint venture in homemade food. It is a curious career turn, though perhaps not as strange as it sounds: software becomes most troublesome when it leaves the kitchen. Their new company gave developers a way to attach a question to a specific line of code while the application was still running.

Lightrun co-founders Ilan Peleg and Leonid Blouvshtein together
01 / PeopleThe two founders built a food venture together before turning their attention to a different kind of mystery: what software does after it leaves the developer's desk.

The missing sentence in the log

Most monitoring systems work with information chosen in advance. A team decides which events to log, which metrics to collect, and which traces to sample. That is sensible engineering until the incident asks a question nobody anticipated. A dashboard may show that requests are slowing down; it cannot reveal the value of a variable that was never recorded. Lightrun’s proposition is to create that missing observation at the moment it matters.

A developer installs an agent alongside an application and a plugin in the IDE. From the source file, they can add a dynamic log, a metric, or a snapshot at a chosen line. A snapshot captures state such as variables and the call stack when its condition is met. These actions can have conditions and expiry times; they can be removed when the investigation is over. The company supports JVM, Python, Node.js, and .NET runtimes, and its paid plans include multi-instance environments such as Kubernetes and serverless deployments.

Lightrun snapshot comparison inside Visual Studio Code
02 / The instrumentA screenshot from Lightrun’s VS Code documentation: the interesting object is the running program’s state, not another screenshot of an error message.

This distinction explains where Lightrun sits in the market. Datadog, New Relic, Dynatrace, Sentry, and other monitoring products can warn a team that something is amiss and supply the evidence they already collect. Lightrun is useful when that evidence stops short of the code path in question. Its integration with IBM Instana makes the division of labor plain: Instana identifies a problem across applications and infrastructure; Lightrun adds code-level telemetry to investigate it. An alert and a microscope are different instruments.

“A day’s work turned into just one hour.”Rami Stern / R&D infrastructure team leader, Taboola

A half-hour wait, repeated

Taboola offers a particularly vivid account of the cost. Its customer story describes services handling roughly half a million requests a second and diagnostic deployments that could take 30 minutes to an hour. Some servers carried as much as 2 TB of RAM, making local reproduction a rather optimistic hobby. The company used Lightrun’s conditional snapshots to isolate particular requests, and temporary timing metrics to locate slow code. Taboola says the change saved more than 260 debugging hours per month. The figure is a customer-reported outcome, and the mechanism is more useful than treating it as a universal promise.

At PPG, the obstacle was different. Its engineers had careful reviews, QA, and multiple environments, yet production failures could still depend on data and conditions unavailable lower down the pipeline. To answer a question about a variable or a request path, they had to edit code, open a pull request, and redeploy just to get a better log. With Lightrun, PPG says its engineers can add logs and snapshots in the live environment and triage without that cycle. The common enemy in both stories is not simply a bug. It is the delay between asking a precise question and receiving evidence.

Teaching an agent to show its work

The company’s path since then has followed the same logic into AI. In 2024, it introduced a Runtime Autonomous AI Debugger in private beta. In February 2026, it launched Lightrun AI SRE, which combines incident signals with on-demand runtime evidence to form and test root-cause hypotheses, suggest fixes, and prepare postmortems. Its newer runtime-aware PR review work applies the sensor before a change is merged. The product is moving from a developer asking the running program a question to an agent that can ask, inspect, and report back.

01 / NoticeAn alert arrives

A monitoring tool or customer report describes the symptom.

02 / ProbeCollect the missing state

A conditional log or snapshot queries the live code path.

03 / VerifyTest the explanation

Engineers examine evidence before accepting a proposed fix.

That progression requires care. A suggested fix is not a verified repair merely because an AI wrote it; the important part is the evidence trail and the human decision around it. Lightrun’s enterprise features - role-based access, audit logs, single sign-on, privacy controls, and read-only instrumentation - speak to the obvious concern: production systems are not a playground. The current pricing page lists a limited free AI SRE plan, while Pro and Enterprise plans have custom prices. Runtime Context is also sold through custom-priced Pro and Enterprise tiers.

The company says its customers include engineers at enterprises across software, finance, retail, and other sectors. Public case studies name Automation Anywhere, Drata, Gong, Inditex, Clarivate, PPG, Priceline, and Taboola. Lightrun has raised substantial capital for that enterprise push, including a $23 million Series A in 2021 and a $70 million Series B in 2025 led by Accel and Insight Partners. The latter announcement said total funding had reached $110 million; published totals elsewhere differ, so the round itself is the sturdier fact.

The idea worth borrowing

Before adding more permanent logging, ask which production question remains unanswered. Instrument that code path narrowly, set a condition and an expiry, capture the evidence, and remove the probe when the answer is known.

The witness is still running

There are limits to this method. It depends on an agent installed in a supported runtime, permission to observe the relevant environment, and a failure that recurs while the probe is active. It cannot magically reconstruct a past variable value that nobody collected. For a routine, reproducible defect, a local debugger may remain the simplest tool. But when an incident lives only amid production traffic, the ability to ask one new question without shipping one new build changes the economics of investigation. Lightrun’s argument is that software should be allowed to answer for itself while it is still on the stand.