At Roku, engineers reserved computing power for peak demand. Resource settings could persist between infrequent reviews, according to its ScaleOps case study. Imagine keeping every restaurant table for the dinner rush, all afternoon. The reservations are sensible. The empty dining room is expensive.
- ScaleOps adjusts the computing resources applications reserve.
- Its software coordinates workloads, replicas and nodes.
- The buying decision turns on reliability as much as savings.
The insurance premium hidden in a request
A Kubernetes “request” tells the scheduler how much CPU or memory to reserve for a container. Consumption can be much lower. That distinction sounds clerical until the scheduler treats a machine as full, and the company rents another one. The first machine may have plenty of breathing room. Its reservation book says otherwise.
There is a sensible reason for the padding. Too little memory can kill a process. Too little usable compute can hurt response times. An oversized request usually produces a less dramatic consequence: another expense on a bill someone else reviews. To an engineer responsible for uptime, spare capacity can look like inexpensive insurance. Multiply that preference across applications and teams, and insurance becomes infrastructure.
ScaleOps was founded in 2022 by Yodar Shafrir and Guy Baron. By its November 2024 funding announcement, it was selling automated resource management to customers including Wiz, SentinelOne and Cato Networks. Its expertise sits in the operational details between an application’s changing needs and the cluster’s ability to satisfy them.
“Most cloud management solutions merely provide optimization recommendations, leaving developers to implement them.”
Yodar Shafrir, speaking to CTech in 2024
Recommendations create another task. ScaleOps sells execution: software that observes application behavior and adjusts allocations continuously. The distinction matters to a platform team already supporting hundreds of services. Another dashboard can make the backlog better informed without making it any shorter.
Three ways to rent less computer
The Core Platform approaches the problem at several levels. Pod rightsizing changes CPU and memory requests according to workload behavior and live cluster conditions. Replica optimization adjusts how many copies of an application run, including minimums, maximums and scaling triggers. Node optimization works on the machines beneath those copies, consolidating capacity and selecting node types that better match the workloads.
Smart Pod Placement addresses a particularly stubborn detail: pods that cannot readily be evicted can leave machines sparsely occupied and prevent efficient packing. ScaleOps places these workloads together to reduce fragmentation. It is a useful reminder that idle capacity is sometimes a placement problem, rather than a shortage of scaling rules.
The software also supports existing HPA and KEDA definitions. Coordinating horizontal scaling with resource requests matters because percentage-based CPU utilization depends on the request used as its denominator. Change that denominator and a horizontal autoscaler can see a different utilization percentage without any change in actual demand. Two individually reasonable decisions need someone to keep the arithmetic straight.
ScaleOps occupies the market between FinOps, which asks where the money goes, and platform engineering, which must safely change where it goes. Cast AI offers competing workload automation, including rightsizing and autoscaler integration. Automation itself is therefore no exclusive distinction. Buyers need to compare deployment requirements, workload behavior, policy controls and the way savings are measured.
Fourteen days before the switch
Roku began with fourteen days of observation before automation. The feared scaling conflict did not materialize, and rollout progressed toward production.
The account reports no production incidents during rollout. Its larger savings projections remain projections. Engineers had evidence for a recurring decision they could delegate.
There is a practical sequence readers can borrow: establish an allocation and cost baseline, observe the recommendations, test representative workloads, then expand automation while watching service performance. A flattering savings percentage is incomplete if traffic changed, latency worsened or the capacity remained on the invoice.
A GPU can be occupied without being busy
ScaleOps has extended its approach to AI infrastructure. A GPU reserved for one workload need not be fully used by it. Memory allocation, compute utilization and model behavior can tell different stories. The company’s AI Infra tools use those signals to decide how workloads can share GPU capacity dynamically.
Its GPU platform includes fractional allocation, replica optimization, inference observability and memory optimization. The promise is to fit useful work into capacity already purchased. That depends on the workload: sharing is valuable when there is shareable headroom and performance can be preserved. A genuinely saturated accelerator offers less room for this particular saving.

Deployment is another part of the pitch. ScaleOps describes its self-hosted option as keeping collected data inside the customer’s environment, and offers air-gapped operation. For a restricted enterprise, the location of the control software can matter before a savings calculation ever reaches procurement.
The company describes its working culture in similarly operational terms: attention to customer problems, execution and ownership of an issue through its resolution. Those are declared values, rather than an employee verdict. For buyers, the useful question is whether support can explain a resource decision when an unfamiliar workload behaves badly.
The price of changing the habit
ScaleOps sells through custom quotes rather than a universal public price. Its pricing form asks about clusters and workloads. A credible purchase calculation therefore needs the quoted software cost, the operational effort and recoverable infrastructure spending. Lower requests alone cannot repay a contract.
Investors have financed a broader ambition. After a $15 million Series A and $58 million Series B, ScaleOps announced a $130 million Series C on March 30, 2026, led by Insight Partners at a valuation above $800 million. Its stated vision is a “Cloud Operating System for the AI era.”
The everyday test is smaller and more exacting. Can the system recover capacity while respecting how an application actually behaves? Memory leaks still need fixing. Workloads with strict placement or disruption requirements still need appropriate policies. Engineers still need evidence. ScaleOps earns its place when their insurance premium can fall without their confidence falling with it.