Skip to main content

FinOps in 2026: How to Cut Cloud Costs by 30% Without Slowing Down Engineering

FinOps in 2026: How to Cut Cloud Costs by 30% Without Slowing Down Engineering
Photo by Markus Winkler on Pexels

Photo by Pixabay on Pexels

In 2023, our cloud bill hit a number that finally got the CFO's attention. Not because it was unexpected — engineering had been warning about the trajectory for two quarters — but because it had crossed a threshold that made it visible in board-level reporting. The conversation that followed was uncomfortable, but it led to the most productive cost reduction effort I've been part of: a six-month FinOps program that cut our AWS spend by 34% without a single feature being delayed.

This is what I learned from that program, updated with what's changed in 2026. I'll be specific about the tactics that worked, the ones that sounded good but didn't, and the organizational dynamics that determine whether FinOps succeeds or stalls at the planning stage.

What FinOps Actually Is (and Isn't)

FinOps is the practice of bringing financial accountability to the variable spend model of cloud infrastructure. The FinOps Foundation — the nonprofit that governs the practice — defines it as "an operational framework and cultural practice which maximizes the business value of cloud, enables timely data-driven decision making, and creates financial accountability through collaboration between engineering, finance, and business teams."

Note what that definition doesn't say: it doesn't say "cut spending." FinOps is not a cost-cutting mandate. It's a visibility and accountability practice. Sometimes FinOps leads you to spend more — in a specific area where underinvestment is causing outages or slow response times that cost more than the compute would. More often, it leads to real reductions. But the mechanism is always the same: make the cost visible, attribute it to the people making the decisions, and give them the information they need to make better choices.

The distinction matters because "we need to cut cloud costs" and "we need FinOps" require different organizational approaches. Cost-cutting is a project with an end date. FinOps is a cultural change that becomes permanent operating practice.

The Three Stages of FinOps Maturity

The FinOps Foundation's maturity model — Crawl, Walk, Run — is a useful map of the journey. I've seen organizations try to skip stages and fail; the progression is meaningful for a reason.

Crawl: Just Get Visibility

At the Crawl stage, the goal is basic visibility. You're answering three questions: what are we spending, on what, and who owns it? This requires tagging infrastructure consistently — every resource tagged with environment, team, service, and cost center. It requires a cost reporting mechanism that non-finance people can access (AWS Cost Explorer, a simple Grafana dashboard, whatever works). And it requires someone responsible for looking at the data regularly.

Most organizations think they're past Crawl when they're not. The tell: can you tell me within five minutes how much a specific team or service spent last month? If not, you're still Crawling. In my experience, 60–70% of organizations claiming to be further along are not.

Walk and Run: Where the Savings Show Up

Walk is where you start acting on the visibility — right-sizing instances, buying commitments, killing idle resources. Run is where cost becomes a first-class engineering metric, factored into design reviews the same way latency or reliability is. You don't need to reach Run to see results. Most of our 34% reduction came from moving cleanly through Walk.

The Tactics That Actually Moved the Number

TacticEffortTypical Saving
Kill idle/orphaned resourcesLow5–10% of bill
Right-size over-provisioned instancesMedium10–20% of compute
Savings Plans / Reserved commitmentsLowup to ~30% on steady workloads
Storage lifecycle policiesMedium20–40% of storage spend

Here's how to translate that into money. On a $1,000/month bill, deleting idle resources alone typically returns $50–$100 — for a few hours of work. Applying a Savings Plan to steady compute that runs 24/7 can turn a $1,000 workload into roughly $700, saving $300 a month, or $3,600 a year, with no engineering change at all. That last one is the most under-used lever I see: teams chase clever architectural rewrites while ignoring a discount they could apply this afternoon.

What Sounded Good but Didn't Work

Two things burned time for us. First, spot instances for anything stateful — the savings looked attractive on paper, but the engineering effort to handle interruptions gracefully cost more than we saved. We kept spot for batch and CI only. Second, chasing micro-optimizations on small services. Trimming a $40/month service to $30 feels productive, but it's noise against a bill measured in tens of thousands. Spend your attention where the dollars are.

Why FinOps Stalls

The failures I've watched weren't technical. They were organizational. FinOps stalls when cost data has no owner, when engineers get a spreadsheet but no authority to act on it, or when finance treats it as a one-time cleanup rather than an ongoing practice. The programs that work give a named person a recurring 30-minute review, tie service cost to the team that ships the service, and make the numbers visible without anyone having to ask.

Bottom Line

If I were starting again with a $1,000-a-month bill (or a $100,000 one — the math scales), I'd do three things this week: tag everything so you know who owns what, delete anything idle, and apply commitments to workloads that run around the clock. Those three moves alone routinely recover 25–30% before you touch a single architectural decision.

FinOps isn't glamorous and it isn't a one-time project. It's the discipline of knowing what you spend and why — and making that knowledge sit with the people who can change it. Done that way, cutting 30% doesn't require slowing engineering down. In our case, it made engineering faster, because teams finally understood the cost of the systems they built.

Comments

Popular posts from this blog

AWS vs Azure vs GCP in 2026: Which Cloud Platform Should You Choose?

The cloud platform decision is one of the most consequential technology choices an organization makes, and in 2026 it's also one of the most misunderstood. Most of the debate I see in enterprise architecture forums reduces to "we're an AWS shop" or "we go Azure because of Microsoft" — neither of which is a strategy. A platform choice made primarily on inertia or existing vendor relationships is a choice that will cost you for years. I've spent significant time in all three major cloud environments — AWS for scale workloads and data engineering, Azure for enterprise SAP and Microsoft-integrated architectures, and GCP for AI-intensive and analytics-heavy use cases. My goal in this guide is to give you a genuine, nuanced comparison that goes beyond feature lists and into the practical realities of choosing and running a cloud platform in 2026. I'll cover market position, each platform's honest strengths and weaknesses, how to match workloads t...

EU AI Act Compliance in 2026: What Every Enterprise Needs to Do Now

EU AI Act Compliance in 2026: What Every Enterprise Needs to Do Now The EU AI Act entered into force on August 1, 2024. The first provisions took effect six months later, and the full implementation timeline runs through 2027. If you're building, deploying, or using AI systems in or for the European Union, this law applies to you — and the window for being caught unprepared is closing fast. I've spent the past year working with enterprise clients on AI governance programs, and one pattern shows up again and again: organizations badly underestimate how much operational work compliance actually takes. It's not a checkbox exercise. It's a rethink of how you develop, document, deploy, and monitor AI systems. This guide is what I wish someone had handed me when I started — the substance of the law, the practical requirements, the deadlines that matter, and the mistakes I keep watching enterprises make. Photo by Petrit Nikolli on Pexels Photo by Karolina Gra...

GPT-6 Astra Is Here: $10/M Tokens, 100% on ExploitBench, and What It Actually Means for Developers

Photo by Michał Robak on Pexels Photo by Tara Winstead on Pexels OpenAI launched GPT-6 Astra on September 3rd, and unlike the usual cadence of incremental updates, this release ships with benchmarks that are hard to look past: 100% on ExploitBench, 98–99.9% on FrontierMath Tier 4 and ARC-AGI-3, and 72.6% on OSWorld 2.0. OpenAI is calling it their most capable model yet — and specifically their best for computer use, coding, and professional work. If you manage an API budget, the real question isn't whether the benchmarks look impressive. It's whether switching your workloads over saves money or burns it. Here's how the numbers actually shake out. The Numbers Behind the Launch Here are the headline specs from OpenAI's announcement: FrontierMath Tier 4: 98–99.9% (previous frontier models sat in the 60–70% range) ARC-AGI-3: 98–99.9% — a benchmark built specifically to resist memorization ExploitBench: 100% — which tripped OpenAI's Preparedness F...