The email arrives on a Thursday afternoon.
Finance has reviewed last quarter’s cloud invoice. The number is significantly higher than the previous quarter. Nobody in the business can immediately explain why. Leadership wants answers before the next board meeting. Engineering gets pulled in. A scramble starts.
If any part of that sounds familiar, this guide is for you.
Cloud bills do not usually double because something catastrophic happened. They double because a handful of small, unnoticed changes stacked up over 90 days , and nobody was watching closely enough to catch them in the first two weeks when they would have been obvious.
This is a diagnostic guide. We are going to walk through exactly how to investigate a cloud bill that has spiked, what you are most likely to find, and what to do about it once you have found it. More importantly, we are going to cover how to make sure it does not happen again , because the businesses that go through one bill shock and make no structural changes are the businesses that go through the second one six months later.
Let’s get into it.
Before anything else, a quick reset.
A cloud bill that has doubled is not evidence that your business is out of control. It is evidence that something changed, and that the change was not flagged in time. That is a process problem, not a technology crisis.
Most of the time, the cause is not a single dramatic event. It is the cumulative effect of several small decisions made across different teams, across different weeks, without anyone connecting the dots.
The goal of this exercise is not to assign blame. It is to understand exactly what happened, so you can make sensible decisions about what to do next , and structural decisions about how to prevent the next one.
Before anything else, a quick reset.
A cloud bill that has doubled is not evidence that your business is out of control. It is evidence that something changed, and that the change was not flagged in time. That is a process problem, not a technology crisis.
Most of the time, the cause is not a single dramatic event. It is the cumulative effect of several small decisions made across different teams, across different weeks, without anyone connecting the dots.
The goal of this exercise is not to assign blame. It is to understand exactly what happened, so you can make sensible decisions about what to do next , and structural decisions about how to prevent the next one.
Across the hundreds of cloud environments FinOps practitioners look at every year, spikes in AWS spending almost always trace back to one or more of the same patterns.
Here they are, in rough order of frequency.
Cause One , A workload was scaled up and never scaled back down
Somebody temporarily increased capacity for a launch, a migration, a load test, or a marketing campaign. The traffic came. The event passed. The capacity was never reduced.
This is the single most common cause of quiet spend growth. It rarely looks dramatic in any individual week. It just adds up , and by the end of a quarter, the baseline has shifted without anyone noticing.
How to check: Compare your average compute utilisation this quarter against the previous quarter. If utilisation has dropped significantly but spend has risen, you have over-provisioned resources that nobody has walked back.
Cause Two , A new service went live and was never properly tagged or tracked
A product team launches a new feature, service, or microservice. The engineering work gets done, the deploy happens, the feature goes to production , and the associated cloud resources get provisioned without being assigned to a team, a product, or a cost centre.
The bill grows. Nobody can immediately tell who or what is driving the growth because the new spend is untagged.
How to check: Pull last quarter’s Cost and Usage Report and filter for resources without proper tagging. The untagged bucket is usually where the biggest surprises live.
Cause Three , A data transfer cost that was hidden until it was not
Data egress , moving data between availability zones, between regions, or out to the public internet , carries different costs at different volumes. Architectures designed at one scale can become unexpectedly expensive at a larger one.
The classic version of this pattern: a service that worked fine when it handled 10,000 events per day becomes a meaningful cost driver when it starts handling a million. Nothing in the architecture changed. Only the volume did.
How to check: Look at your data transfer costs as a percentage of total spend. If they have grown disproportionately compared to your compute or storage, data transfer is likely the culprit.
Cause Four , Storage grew, but nothing was archived
S3 bills tend to grow in one direction. Up.
Every product, every service, every team generates data. Log files. Backups. Snapshots. Media. Application state. Over time, storage-tier data that should be archived to cheaper tiers stays in premium storage because nobody has set up a lifecycle policy.
The bill climbs quietly. None of it triggers an alert because no single addition is dramatic. It just compounds.
How to check: Look at your S3 bill over the last six months. If storage cost has grown linearly or exponentially, and your business has not generated proportionally more user activity, you are paying premium prices for data nobody is reading.
Cause Five , A pricing arrangement expired or was never claimed
Savings Plans, Reserved Instances, and other commitment-based discounts have expiry dates. When they lapse and are not renewed, workloads that were running on discounted pricing silently revert to on-demand , and on-demand is dramatically more expensive.
The bill jumps. Nothing else changed. But the underlying pricing arrangement did.
The related version of this problem: the discounts were never claimed in the first place. The business has been paying on-demand prices for stable workloads that qualified for 30 to 72 percent discounts, and nobody ever did the maths.
How to check: Pull your Savings Plans and Reserved Instance coverage reports. If coverage dropped in the same period the bill spiked, expiry is your answer. If coverage is low or zero, you have been overpaying from the start.
Cause Six , An AI or data workload started running
This is the 2026 addition to the classic list.
Generative AI workloads , model training, inference, embeddings, RAG pipelines , consume cloud resources differently from traditional applications. A single experimental AI feature that went live in the middle of the quarter can drive significant spend on GPU compute, storage, and data transfer.
Businesses that had not built FinOps around AI are finding themselves surprised by bills they cannot immediately reconcile.
How to check: Look at your SageMaker, Bedrock, EC2 GPU instance, and inference endpoint spend. If any of these have grown meaningfully in the past quarter, an AI workload is part of your story.
Now let’s get practical. If your bill has doubled and you need to find out why, here is the investigation process we use.
Step One , Compare month over month, not just total
Do not just look at the quarterly total. Break it down by month. A bill that doubled over three months tells a different story depending on whether it doubled in month one, month three, or steadily across all three.
A sudden jump in a single month points to a specific event , a new service launch, a misconfigured resource, an expired discount. Gradual growth across all three months points to a drift problem , over-provisioning, storage accumulation, or unchecked scaling.
Knowing which pattern you are looking at immediately narrows the investigation.
Step Two , Break it down by service
Once you know the timeline, break the spend down by AWS service. Which services are responsible for the growth?
If it is EC2, you are looking at a compute issue. If it is S3, you are looking at a storage issue. If it is data transfer, you are looking at an architectural issue. If it is SageMaker, Bedrock, or GPU instances, you are looking at an AI workload. If it is multiple services simultaneously, you are looking at a broader scaling event.
The service-level view tells you which team to talk to, and what questions to ask them.
Step Three , Break it down by tag
If your tagging is in place, this is where the investigation gets specific. Break spend down by team, by product, by environment. Which combinations have grown the most?
This is also where you find out how much of your spend is untagged , and untagged spend is almost always where the surprises live.
If your tagging is inconsistent or incomplete, you are going to hit a wall here. That wall is itself a finding. Fixing the tagging is step one of preventing the next bill shock.
Step Four , Look at the resource inventory
Pull the list of resources that were created in the billing period. How many new EC2 instances were launched? How many new RDS databases? How many new Lambda functions? How many new storage buckets?
If the growth in resources does not match the growth in product features or user activity, something was provisioned that should not have been.
Step Five , Check for anomalies and alerts
AWS provides anomaly detection through Cost Anomaly Detection. If you had it turned on, go back and review the alerts from the period. If you did not have it turned on, turn it on now.
The anomalies that were flagged but ignored often explain the bulk of the growth.
Once you have identified the causes, the next question is what to do. Here is the short-term playbook.
Shut down what should not be running. If the investigation surfaces over-provisioned resources, idle workloads, or orphaned services , retire them. Immediately. Do not wait for the next quarterly review.
Schedule what does not need to run 24/7. Dev and staging environments, test workloads, and any non-production infrastructure should be on a schedule. Shutting them down outside working hours is usually a same-week fix.
Renew or claim your discounts. If Savings Plans or Reserved Instances expired, renew them. If they were never claimed, model out the commitment levels and claim what makes sense for your baseline usage.
Introduce tagging where it is missing. This is not a one-week project, but it is the most important medium-term fix. Without proper tagging, every future cost investigation will hit the same wall this one did.
Set up budgets and anomaly alerts. If you did not have them before, set them now. The whole point of alerting is to catch the next version of this problem in the first two weeks, not the twelfth.
Talk to leadership, clearly. Once you understand the causes, brief the leadership team plainly. What happened, why it happened, what you have already done about it, and what you are going to do structurally to prevent it from happening again. A clear-eyed post-mortem builds credibility. Vague reassurance does the opposite.
This is the part most businesses skip. They fix the immediate problem, move on, and then go through the same cycle again six months later.
If you want to be the business that goes through the bill shock once and never again, this is the work.
Assign a clear owner for cloud spend. A single person, or a small team, whose job it is to understand the bill, explain it, and improve it. Not a committee. Not a quarterly meeting. A standing responsibility.
Build the cross-functional monthly review. Engineering, finance, and at least one senior leader in the same room once a month, looking at the same numbers. The room itself is often the intervention.
Make cost visible to engineers. Cost data in the tools engineers already use. Cost impact on architecture review templates. Cost dashboards surfaced in team channels. When engineers can see what their decisions cost, they make better decisions.
Set governance guardrails. Budgets with automated alerts. Tag enforcement that prevents untagged resources from being created. Automated anomaly detection. These are not restrictions , they are the systems that make good decisions the default.
Start tracking cloud cost per customer or per unit of revenue. The single most important metric shift a business can make is moving from “how much did we spend on cloud” to “how much did it cost to serve each customer, or generate each transaction, or run each product.” That shift turns cloud spend from an expense to be controlled into a unit economic to be optimised.
This is what FinOps actually is , not a one-time audit, but the operational discipline that keeps a business from ending up in this situation in the first place.
For some businesses, running this investigation internally is straightforward. The data is clean, the team has bandwidth, and the patterns are recognisable.
For others, it is not. The tagging is incomplete. The team is stretched thin. The leadership team wants answers faster than internal capacity allows. Or the business knows there is waste but does not know where to start.
If that describes your situation, this is exactly where a proper FinOps engagement adds disproportionate value. The right partner will do the investigation for you, surface the causes with specificity, propose concrete fixes, and , if they are worth working with , help build the structural practice that prevents a repeat.
At Marlocks Technologies, this is what our FinOps practice is built for.
We are an AWS Advanced Consulting Partner with AWS-certified practitioners who have investigated cloud environments across fintech, banking, healthcare, and enterprise , including the specific operational realities of African markets, where currency exposure and infrastructure considerations add layers that global frameworks do not always account for.
Our FinOps Assessment is designed for exactly this moment. We look at your environment, identify your top three cost optimisation opportunities, explain what caused the current state, and give you a clear plan for both immediate recovery and long-term discipline. It runs at no cost and typically delivers in under a week.
If your AWS bill has doubled and you need answers, this is the fastest way to get them.
Every business we have ever worked with has, at some point, been surprised by a cloud bill.
The difference between the ones that stay surprised and the ones that do not is whether they treat the moment as a wake-up call or a one-off event. The businesses that treat it as a one-off pay the same lesson fee again within a year. The ones that treat it as a wake-up call end up running tighter, more predictable, and more confident cloud operations within six months.
Which one this ends up being for your business is up to you.
[Book your complimentary FinOps Assessment]
Or if you would prefer to have a broader conversation about your cloud strategy first:
[Book a discovery call with our team]
Marlocks Technologies is an AWS Advanced Consulting Partner helping enterprises and growing businesses across Africa and beyond build cloud, AI, and data practices that scale sustainably. Learn more at marlockstech.com.
We are building solutions and talents that transcend the future. We have over 15 years of experience in ICT services industry.
Join our newsletter for exciting updates and deals.
© 2025 Marlocks Technologies