A media analytics company running its entire data platform on AWS was spending $480,000 per month on cloud infrastructure. The bill had grown organically over three years as the platform expanded from a single Redshift cluster to a complex architecture spanning Redshift, S3, EMR, Kinesis, Lambda, and SageMaker. Each component had been added to solve a specific problem. Nobody had stepped back to evaluate the architecture as a whole.
The CFO issued a mandate: reduce the data platform spend by thirty percent without reducing the data products that the company delivered to its customers. The data team pushed back. Every component was in use. Every pipeline was serving production traffic. Removing anything would break something. The CFO was unmoved. The spend was growing faster than revenue, and the ratio was unsustainable.
We were brought in to find the waste. Not the features to cut: the architecture that was costing more than the value it produced.
The Billing Audit
The first step was decomposing the $480,000 monthly bill into its components and mapping each component to the business value it produced.
Redshift accounted for $186,000: thirty-nine percent of the bill. The company ran four clusters: a production cluster for customer-facing queries, a staging cluster for development, an analytics cluster for internal reporting, and a machine learning cluster for model training data preparation. The production cluster was appropriately sized. The staging cluster was running twenty-four hours a day despite only being used during business hours, a $28,000 monthly waste. The analytics cluster was oversized by approximately sixty percent based on actual query patterns. The ML cluster was used for three batch jobs per day that collectively ran for four hours, but the cluster was provisioned for peak capacity and sat idle for the remaining twenty hours.
EMR accounted for $94,000: twenty percent of the bill. The company ran twelve persistent EMR clusters for various Spark jobs. Nine of the twelve clusters were running continuously. Analysis of job schedules showed that the average cluster was active for three hours per day. The remaining twenty-one hours were idle compute that the company was paying for because cluster startup time discouraged on-demand provisioning. The team had sized clusters for peak memory requirements and kept them running because a cold start took eight minutes, which was too slow for jobs that needed to run on a tight schedule.
S3 accounted for $67,000: fourteen percent of the bill. The company stored 4.2 petabytes of data across all S3 buckets. A lifecycle analysis showed that 2.8 petabytes (sixty-seven percent) had not been accessed in over ninety days. This data was in S3 Standard storage at $0.023 per gigabyte per month. The same data in S3 Glacier Instant Retrieval would cost $0.004 per gigabyte: an eighty-three percent reduction for cold data.
Kinesis accounted for $52,000: eleven percent of the bill. The company ran eight Kinesis data streams for real-time event ingestion. Analysis of consumer throughput showed that four of the eight streams were over-provisioned by a factor of three. The streams had been sized for a projected traffic increase that never materialised.
Lambda, SageMaker, and other services accounted for the remaining $81,000.
This diagram requires JavaScript.
Enable JavaScript in your browser to use this feature.
The waste was not in any single component. It was distributed across every component. Each individual waste amount seemed defensible when viewed in isolation. The staging cluster needed to be available, the EMR clusters needed fast startup, the S3 data was potentially needed. Viewed together, the waste totalled $201,000 per month, forty-two percent of the bill.
Changes Made
Every change was made during a six-week sprint. No features were cut. No pipelines were decommissioned. No data was deleted.
Redshift: The staging cluster was moved to a scheduled start-stop. It started at 7 AM and stopped at 8 PM on weekdays. Weekend access was available on-demand through a script that started the cluster and stopped it after two hours of inactivity. Savings: $28,000 per month. The analytics cluster was right-sized based on actual query patterns. The largest tables were identified, and sort keys and distribution styles were optimised to reduce the cluster’s memory requirements. Savings: $22,000 per month. The ML cluster was replaced by Redshift Serverless for the three daily batch jobs. Serverless charged only for compute consumed during query execution. Savings: $31,000 per month.
EMR: Nine persistent clusters were converted to transient clusters triggered by the orchestration tool. The orchestration tool started an EMR cluster, ran the job, and terminated the cluster. To address the startup time concern, the team used EMR-managed scaling with a warm pool of instances that started in under two minutes. For the three jobs that truly required persistent clusters because they ran every thirty minutes, the clusters were right-sized using auto-scaling policies. Savings: $58,000 per month.
S3: The 2.8 petabytes of cold data was moved to S3 Glacier Instant Retrieval using a lifecycle policy. Data not accessed for sixty days was automatically transitioned. The transition was transparent to downstream systems because Glacier Instant Retrieval provides millisecond access latency: slower than Standard but fast enough for analytical queries. The team verified query performance before and after the transition and found no measurable degradation for the workload patterns in use. Savings: $44,000 per month.
Kinesis: The four over-provisioned streams were right-sized to match actual throughput plus a thirty percent buffer. The buffer was maintained because traffic spikes did occur (breaking news events drove sudden surges in the media analytics workload) but the spikes did not justify a three-times headroom. Savings: $18,000 per month.
What We Did Not Touch
Three cost centres were evaluated and left unchanged.
The production Redshift cluster was the largest single expense at $105,000 per month. It could have been reduced by offloading some queries to a cheaper alternative. But the cluster served customer-facing dashboards with strict latency requirements. Any performance degradation would have been visible to customers and would have triggered complaints. The team decided that customer-facing performance was not an acceptable trade-off for cost savings.
The SageMaker training jobs ran on GPU instances that were expensive: $47 per hour for a p3.2xlarge. The jobs ran for an average of six hours per day, totalling approximately $8,500 per month. The team evaluated spot instances, which would have reduced cost by sixty percent. But spot instances can be interrupted, and a training job that was interrupted after four hours wasted four hours of compute. For short training jobs (under two hours), spot instances were viable. For the longer jobs, the risk of interruption and retraining was not worth the savings. The team chose to keep on-demand pricing for long jobs and use spot instances only for short jobs, saving $2,100 per month: a modest amount that was not included in the headline number.
The Lambda functions consumed $6,800 per month. The team reviewed the functions for inefficiencies (cold starts, oversized memory allocations, unnecessary invocations) and found $400 per month of optimisation. The effort required to refactor the functions was estimated at three engineer-weeks, which was not justified for $400 monthly savings. The Lambda bill was left unchanged.
The Outcome
Monthly spend dropped from $480,000 to $290,000, a forty percent reduction. The savings of $190,000 per month, or $2.28 million per year, exceeded the CFO’s thirty percent target.
No customer-facing features were affected. No internal dashboards changed. No data was lost. The only user-visible change was that the staging cluster required an explicit start command outside business hours, which the engineering team accepted because they understood the cost trade-off.
The more important outcome was architectural awareness. The data team had been making provisioning decisions in isolation. Each team sized their own clusters, configured their own storage, and managed their own streams. Nobody had a cross-cutting view of the total cost. After the audit, the company implemented a monthly cost review where each component owner reported their spend against a budget and justified any increase. The review was not punitive. It was informational. The team needed to see the aggregate picture to make good provisioning decisions.
The Waste Detection Rule
Cloud waste follows a pattern: resources sized for peak demand running during non-peak hours, persistent resources that should be transient, and storage tiers that do not match access patterns. These three patterns account for seventy to eighty percent of cloud waste in most data platforms.
Run a billing audit before cutting features. Decompose the bill. Map each component to its utilisation. Look for resources that are running when nobody is using them, sized larger than their workload requires, or storing data in a tier that does not match its access frequency. The waste is almost always there. It accumulated because each provisioning decision was made in isolation and nobody audited the aggregate.
The audit takes two weeks. The fixes take four to six weeks. The savings are permanent. This is the highest-ROI engagement in cloud cost management because the waste is architectural, not behavioural. You are not asking people to use less. You are asking the infrastructure to match the workload.