
Photo by Pixabay on Pexels
In 2023, our cloud bill hit a number that finally got the CFO's attention. Not because it was unexpected — engineering had been warning about the trajectory for two quarters — but because it had crossed a threshold that made it visible in board-level reporting. The conversation that followed was uncomfortable, but it led to the most productive cost reduction effort I've been part of: a six-month FinOps program that cut our AWS spend by 34% without a single feature being delayed.
This is what I learned from that program, updated with what's changed in 2026. I'll be specific about the tactics that worked, the ones that sounded good but didn't, and the organizational dynamics that determine whether FinOps succeeds or stalls at the planning stage.
What FinOps Actually Is (and Isn't)
FinOps is the practice of bringing financial accountability to the variable spend model of cloud infrastructure. The FinOps Foundation — the nonprofit that governs the practice — defines it as "an operational framework and cultural practice which maximizes the business value of cloud, enables timely data-driven decision making, and creates financial accountability through collaboration between engineering, finance, and business teams."
Note what that definition doesn't say: it doesn't say "cut spending." FinOps is not a cost-cutting mandate. It's a visibility and accountability practice. Sometimes FinOps leads you to spend more — in a specific area where underinvestment is causing outages or slow response times that cost more than the compute would. More often, it leads to real reductions. But the mechanism is always the same: make the cost visible, attribute it to the people making the decisions, and give them the information they need to make better choices.
The distinction matters because "we need to cut cloud costs" and "we need FinOps" require different organizational approaches. Cost-cutting is a project with an end date. FinOps is a cultural change that becomes permanent operating practice.
The Three Stages of FinOps Maturity
The FinOps Foundation's maturity model — Crawl, Walk, Run — is a useful map of the journey. I've seen organizations try to skip stages and fail; the progression is meaningful for a reason.
Crawl: Just Get Visibility
At the Crawl stage, the goal is basic visibility. You're answering three questions: what are we spending, on what, and who owns it? This requires tagging infrastructure consistently — every resource tagged with environment, team, service, and cost center. It requires a cost reporting mechanism that non-finance people can access (AWS Cost Explorer, a simple Grafana dashboard, whatever works). And it requires someone responsible for looking at the data regularly.
Most organizations think they're past Crawl when they're not. The tell: can you tell me within five minutes how much a specific team or service spent last month? If not, you're still Crawling. In my experience, 60–70% of organizations claiming to be further along are not.
Walk and Run: Where the Savings Show Up
Walk is where you start acting on the visibility — right-sizing instances, buying commitments, killing idle resources. Run is where cost becomes a first-class engineering metric, factored into design reviews the same way latency or reliability is. You don't need to reach Run to see results. Most of our 34% reduction came from moving cleanly through Walk.
The Tactics That Actually Moved the Number
| Tactic | Effort | Typical Saving |
|---|---|---|
| Kill idle/orphaned resources | Low | 5–10% of bill |
| Right-size over-provisioned instances | Medium | 10–20% of compute |
| Savings Plans / Reserved commitments | Low | up to ~30% on steady workloads |
| Storage lifecycle policies | Medium | 20–40% of storage spend |
Here's how to translate that into money. On a $1,000/month bill, deleting idle resources alone typically returns $50–$100 — for a few hours of work. Applying a Savings Plan to steady compute that runs 24/7 can turn a $1,000 workload into roughly $700, saving $300 a month, or $3,600 a year, with no engineering change at all. That last one is the most under-used lever I see: teams chase clever architectural rewrites while ignoring a discount they could apply this afternoon.
What Sounded Good but Didn't Work
Two things burned time for us. First, spot instances for anything stateful — the savings looked attractive on paper, but the engineering effort to handle interruptions gracefully cost more than we saved. We kept spot for batch and CI only. Second, chasing micro-optimizations on small services. Trimming a $40/month service to $30 feels productive, but it's noise against a bill measured in tens of thousands. Spend your attention where the dollars are.
Why FinOps Stalls
The failures I've watched weren't technical. They were organizational. FinOps stalls when cost data has no owner, when engineers get a spreadsheet but no authority to act on it, or when finance treats it as a one-time cleanup rather than an ongoing practice. The programs that work give a named person a recurring 30-minute review, tie service cost to the team that ships the service, and make the numbers visible without anyone having to ask.
Bottom Line
If I were starting again with a $1,000-a-month bill (or a $100,000 one — the math scales), I'd do three things this week: tag everything so you know who owns what, delete anything idle, and apply commitments to workloads that run around the clock. Those three moves alone routinely recover 25–30% before you touch a single architectural decision.
FinOps isn't glamorous and it isn't a one-time project. It's the discipline of knowing what you spend and why — and making that knowledge sit with the people who can change it. Done that way, cutting 30% doesn't require slowing engineering down. In our case, it made engineering faster, because teams finally understood the cost of the systems they built.
Comments
Post a Comment