Imagine a retailer on the eve of its busiest weekend. The checkout works. The dashboards are green. Somewhere between the storefront and the database, however, a dependency takes a little longer to answer than anyone expected. No one has tested what happens next. The first honest measurement may arrive with a queue of angry customers.
Gremlin’s business is built around moving that moment forward. It gives engineers a controlled way to slow a connection, remove a host, exhaust a resource, or mimic a larger outage, then watch how an application responds. The company began in 2016 with the provocative language of chaos engineering. Today, it sells something more prosaic and more useful: a repeatable account of which services withstand trouble and which ones merely appear to.
The short version
- Gremlin tests real failure modes, from slow dependencies to a lost cloud region.
- Its platform turns recurring test results into service-level reliability scores.
- Large engineering teams buy it by service; pricing requires a custom quote.
- The useful habit is simple: find a weakness, fix it, and run the same test again.
The pager was the market research
Co-founder Kolton Andrus had been a call leader at Amazon and Netflix, responsible for helping resolve high-pressure incidents. Gremlin’s 2017 launch post made the founding logic plain: if you have spent enough nights finding failures at 3 a.m., you begin to wonder why the test could not have happened on a quiet Tuesday. Andrus and co-founder Matthew Fornaciari packaged that idea as a hosted service with controlled attacks, an undo button, and security controls intended to make experiments acceptable inside companies whose systems cannot simply be treated as toys.

At launch, Gremlin named Twilio, Expedia, Confluent, and Remind among its users and announced a $7.5 million Series A. An $18 million Series B followed in 2018, alongside application-level fault injection. That mattered because modern failures do not always look like a machine turning off. A service may return the wrong response, answer too slowly, or keep running while its dependencies quietly collapse.
“Gremlin exists to help companies prevent downtime before it happens.”KOLTON ANDRUS / GREMLIN’S 2017 LAUNCH POST
A score for what has actually been tried
The company’s more consequential shift came when it stopped treating each experiment as an isolated event. Gremlin’s Reliability Management platform defines services, runs standard tests for scalability, redundancy, and failed or slow dependencies, then assigns scores based on the outcomes. Its agents can discover some dependencies from DNS and socket data; teams can add others and test them deliberately. A dashboard shows the scores across an organization and how they move over time.
The distinction is sharp. An uptime graph can tell a manager that a system stayed available last month. It cannot show whether a failover will work tomorrow. Gremlin’s score is useful only because someone has asked the system an awkward question and recorded its answer. The number is an aid to prioritization, not a promise of immunity. A service can pass the tested scenarios and still fail in a way no one thought to simulate.

That loop is Gremlin’s place in the market. Monitoring vendors tell engineers what is happening; incident-response tools help after an alert; cloud providers and open-source projects offer ways to inject faults. Gremlin’s enterprise pitch joins fault injection to standardized tests, dependency maps, scoring, governance, and reports. It supports AWS, Azure, Google Cloud, Kubernetes, Linux, Windows, and on-premises environments, with a self-hosted Private Edition for organizations that require it. Its 2025 Dynatrace integration uses discovered Kubernetes services and observability signals to shorten test setup.
The checkout test
Sephora offers a useful example because the story has a before and an after. During a move from a monolithic system to Kubernetes microservices, its performance engineering team used Gremlin in a production-like performance environment. They recreated earlier issues, checked circuit breakers, introduced latency, and tested how services behaved when nodes or dependencies became unavailable. The tests exposed severe issues before release. Each sprint then acquired a regular set of resilience tests.
The reported result was a 2024 holiday season on the new platform without a major service interruption, including Black Friday and Cyber Monday. It would be foolish to assign that outcome to one product: Sephora’s engineering work, migration planning, and fixes were essential. What Gremlin supplied was a way to turn worrying possibilities into observable behavior before the traffic arrived. The first thing to fail in this story was the assumption that a green dashboard meant the new architecture was ready.
Visa Cross-Border Solutions chose a different route to the same habit. It built a staging environment designed to mirror production traffic, dependencies, and data, then ran automated tests on critical services every two weeks. Scores below 90 called for investigation in the same sprint. Friday GameDays made failure response a group exercise rather than a specialist’s private skill. A score threshold could have become theater; ownership by service teams, reruns after fixes, and a realistic staging environment gave it teeth.
How big can the rehearsal get?
In February 2026 Gremlin launched Disaster Recovery Testing, which can run coordinated experiments across services to simulate a lost zone, region, or datacenter. Organization-level health checks can halt the exercise if key metrics cross a threshold, and the resulting report records which services passed. Gremlin says it ran a zone evacuation on its own production infrastructure before launch without interrupting service. The claim is narrow, but instructive: the company submitted its own product to the sort of exam it sells.
The practical lesson for a reader is smaller than a full regional drill. List the services people would miss first. Name one failure each service ought to survive. Decide which customer-facing metric would tell you the experiment is going badly. Run a limited test, repair the weakness, and rerun it. If staging bears little resemblance to production, as Visa’s team understood, a pass there proves less than it seems. If there is no safe stop condition, the experiment needs better preparation before it runs.
Its newest wager, Foresight AI, is listed as early access. Gremlin says supervised agents can analyze test failures, recommend concrete changes, and rerun tests to verify repairs. Whether that makes reliability programs easier to maintain will depend on the quality of the tests and the care with which engineers review the proposed fixes. The older lesson remains the sturdier one: software confidence becomes more credible when the failure has been invited in, watched closely, and made to leave on schedule.