site-reliability-engineering

(18)
Company
PagerDuty Built the Fire Alarm for the Internet. Now It Wants to Put Out the Fire.
Saas · Enterprise · Developer Tools

PagerDuty Built the Fire Alarm for the Internet. Now It Wants to Put Out the Fire.

The software that made the right engineer's phone ring at 3 a.m. has grown into a $500 million-a-year operations business. Its harder assignment is teaching machines what to do before a human even answers.

pagerduty · incident-managementRead →
Company
Honeycomb Bet Engineers Would Rather Ask Better Questions Than Stare at More Dashboards
Saas · Enterprise · Developer Tools

Honeycomb Bet Engineers Would Rather Ask Better Questions Than Stare at More Dashboards

Born from the production chaos at Parse, Honeycomb built an observability business around a stubborn idea: keep the context, query it fast, and let engineers investigate the failure they did not know to predict. Ten years later, that architecture is becoming the evidence layer for AI agents too.

observability · developer-toolsRead →
Legend
Nishant Modak Is Teaching Software to Notice What Humans Miss
Founder · Engineer · Executive

Nishant Modak Is Teaching Software to Notice What Humans Miss

The Last9 founder spent years learning that failure is rarely a single dramatic moment. His answer is an observability company built to close the gap between the systems engineers imagine and the systems actually running.

nishant-modak · last9Read →
Legend
Pankaj Thakkar Wants the Cloud to Keep Its Receipts
Founder · Engineer · Executive

Pankaj Thakkar Wants the Cloud to Keep Its Receipts

Before he asked companies to own their observability data, Pankaj Thakkar helped make networks programmable. Now the Kloudfuse co-founder and CEO is applying the same instinct to telemetry: open the system, expose the machinery and make the bill predictable.

pankaj-thakkar · kloudfuseRead →
Company
BayOne Solutions
Enterprise · Ai · Saas

BayOne Solutions

BayOne built a sizable business by selling two things enterprises usually buy separately: the technology work and the people who make it run. Now its bet on AI, managed delivery and diverse talent is turning a staffing heritage into a broader transformation play.

bayone-solutions · enterprise-technologyRead →
Company
Just After Midnight
Enterprise · Saas · Developer Tools

Just After Midnight

Just After Midnight (JAM) is a London-founded, globally distributed managed services company that keeps mission-critical, customer-facing digital platforms running around the clock. It provides 24/7/365 application, website and cloud support, managed cloud services across AWS, Azure, GCP and Alibaba Cloud, and DevOps consultancy - acting as the after-hours and full-stack support partner behind major brands and digital agencies. Founded in 2016 to solve the problem of out-of-hours support that agencies struggle to staff, JAM now operates across four continents and is part of the Ultima group.

24/7-support · managed-cloudRead →
Company
EverOps
Developer Tools · Enterprise · Saas

EverOps

EverOps is a San Francisco-based IT and cloud managed services provider that embeds small teams of senior engineers - called TechPods - directly inside client organizations to run DevOps, SRE, ITOps, and security work. Founded in 2012 and re-capitalized in 2022, the firm pitches itself as an alternative to advice-only consulting and staffing shops: pods sit in the client's toolchain, attend standups, and take ownership of outcomes across cloud infrastructure, platform engineering, observability, and cloud cost reduction.

devops · sreRead →
Company
Vibranium Labs
Ai · Saas · Enterprise

Vibranium Labs

Vibranium Labs is a New York-based AI company building Vibe AI, an agentic 'AI Site Reliability Engineer' that acts as the first responder to IT incidents. Its Vibe OnCall platform pages AI agents to detect, triage, and resolve outages before human engineers are woken up, integrating with tools like Slack, PagerDuty, Datadog, Jira, and ServiceNow. The company says its agents cut mean time to resolution by up to 85%, and positions itself as an AI-native alternative to PagerDuty and Opsgenie. Founded by Sang Lee (ex-Google, ex-AWS) with FiscalNote co-founder Tim Hwang, it raised $4.6M in seed funding in September 2025.

ai-incident-response · ai-sreRead →
Legend
Auden Ehringer
Engineer · Executive · Operator

Auden Ehringer

Auden Ehringer is the Chief Technology Officer of Pave Finance and CEO of its technology subsidiary, Pave Labs. A Stanford-trained electrical engineer and computer scientist, he spent the early part of his career keeping some of the world's largest systems running - building Alexa sports features at Amazon Lab126 and serving as a Site Reliability Engineer on YouTube Discovery at Google. In 2023 he traded planet-scale infrastructure for the messier world of money, joining the New York fintech Pave Finance to build software that lets independent advisers automate and personalize portfolio management across thousands of client accounts. The platform he leads now helps oversee billions in assets, and in 2025 the company closed an oversubscribed $14M seed round.

auden-ehringer · pave-financeRead →
Legend
Sang Lee
Founder · Engineer · Executive

Sang Lee

Sang Lee is the co-founder and CEO of Vibranium Labs, a New York startup building Vibe AI, billed as the first 24/7 AI site reliability engineer. A former Google and AWS engineer who lived through too many 2 a.m. outage firefights, Lee is building an agentic AI teammate that detects, triages, and resolves IT incidents before a human ever wakes up. In September 2025 the company raised $4.6M in seed funding led by Calibrate Ventures and Mirae Asset, with backing from a16z, Franklin Templeton, and others, claiming up to an 85% cut in mean time to resolution and counting Shutterstock among its clients.

sang-lee · vibranium-labsRead →
Company
Raw Engineering
Enterprise · Saas · Developer Tools

Raw Engineering

Raw Engineering is a San Francisco software products-and-services firm founded in 2007 that builds custom mobile apps, websites, SaaS products and cloud infrastructure for Fortune 500 enterprises, sports teams and startups. Best known as the lab that incubated and spun out the headless CMS pioneer Contentstack and the enterprise mobile platform Built.io (acquired by Software AG in 2018), the company today runs five centers of excellence and powers personalized Digital Fan Experience platforms for clients including the Sacramento Kings, Miami Heat and the NBA.

raw-engineering · rawengRead →
Company
SORINT.lab
Enterprise · Developer Tools · Ai

SORINT.lab

SORINT.lab is a multinational, vendor-neutral 'Next Generation System Integrator' founded in Bergamo, Italy in 1985. With roughly 1,500 engineers across 17 offices in Europe, the USA and Africa, it helps over 100 large organizations run and modernize their IT through DevOps, CI/CD, cloud adoption, application modernization, next-generation IT operations and site reliability engineering. The company is a notable open-source contributor (Agola, Stolon, Ercole, Sircles) and runs on a flat, holacracy-style 'Sircles' organizational model.

devops · cicdRead →
Company
Zelar
Developer Tools · Enterprise · Ai

Zelar

Zelar (ZelarSoft) is a cloud-native engineering and consulting firm that helps banks, telcos, oil & gas, and government teams adopt Kubernetes, DevOps, and AI without the usual pain. Founded in 2018 and led by CEO Vasu Maganti, the company pairs hands-on services - cloud migration, SRE, security, and data engineering - with its own platforms: Klusternetes for self-service Kubernetes multi-tenancy, OpenOps for production-ready GKE stacks, and Cokpit for agentic AI DevOps. A Google Cloud Premier Partner with offices across the US, Canada, India, and the UAE, Zelar bets that most teams want cloud-native outcomes, not cloud-native homework.

kubernetes · devopsRead →
Legend
Jesse Robbins
Investor · Founder · Operator

Jesse Robbins

Jesse Robbins is General Partner at Heavybit, the San Francisco-based venture firm focused exclusively on developer-first companies. He co-founded Chef (sold to Progress Software for $200M+), invented GameDay chaos engineering at Amazon where he held the title 'Master of Disaster', and co-created the O'Reilly Velocity Conference that seeded the global DevOps movement. A former volunteer firefighter and EMT, he brings a crisis-responder's instincts to early-stage investing, backing companies like Snyk, PagerDuty, Fastly, LaunchDarkly, and Tailscale. His portfolio spans 60+ companies with five IPOs.

venture-capital · developer-toolsRead →
Legend
Freddie Heygate
Executive · Operator · Founder

Freddie Heygate

Freddie Heygate is CEO, North America at Just After Midnight, a London-founded managed cloud services and 24/7 digital support company now part of the Ultima group. After spending five years building JAM's Asia-Pacific operations from Singapore, he relocated to Austin, Texas in January 2024 to lead the company's US expansion. With a background spanning digital consultancy, marketing, and business development at agencies including Reading Room and The Drum, Freddie brings over a decade of cross-continental experience helping global brands like Ford, Abbott, and Vodafone keep their digital infrastructure running around the clock.

ceo · north-americaRead →
Legend
Laura Nolan
Engineer · Author · Activist

Laura Nolan

Laura Nolan is a Principal Engineer at Stanza Systems, a veteran Site Reliability Engineer, and one of tech's most credible voices on autonomous weapons ethics. After five years at Google - where she contributed to the seminal O'Reilly SRE book and resigned over Project Maven - and seven years at Slack as Senior Staff Engineer, she brings deep technical authority to both the reliability engineering world and the global debate over killer robots. Based in rural Ireland, she speaks at SREcon, QCon, and TED stages alike, writes the Responsible Computing newsletter on Substack, and holds seats on the USENIX Board and the SREcon Steering Committee.

sre · site-reliability-engineeringRead →
Legend
Liz Fong-Jones
Engineer · Activist · Author

Liz Fong-Jones

Liz Fong-Jones is a Technical Fellow at Honeycomb.io, renowned SRE practitioner, co-author of 'Observability Engineering' (O'Reilly), and one of the most influential voices in the observability and platform engineering space. With 18+ years in software engineering spanning Google (11 years) and Honeycomb, she bridges deep technical expertise with fierce advocacy for labor rights, trans inclusion, and workplace equity. She led the Google Walkout Strike Fund in 2018, founded the Solidarity Fund by Coworker, and sits on the OpenTelemetry governance committee - all while speaking at every major SRE and DevOps conference on Earth.

observability · sreRead →
Legend
Roy Rapoport
Engineer · Executive · Author

Roy Rapoport

Roy Rapoport is a veteran engineering leader and writer whose work has quietly reshaped how tech companies think about people, reliability, and operational culture. Best known for his two stints at Netflix (where he built Insight Engineering and its operational platform) and a stint at Slack, Roy popularized the Manager README format and authored influential frameworks on feedback, trust, and performance improvement. He writes on Medium about the subtler mechanics of leadership, raises goats in California, and insists he will never retire.

sre · site-reliability-engineeringRead →