Last quarter, we pulled up our AWS Cost Explorer for a routine review and found something uncomfortable: we were spending 40% more than we needed to — not because our system was under-provisioned or our architecture was wrong, but because of a quiet accumulation of bad defaults, forgotten resources, and pricing models we'd never revisited. No infrastructure migration. No refactoring sprint. Just a focused two-week audit that saved us nearly half our cloud spend. Here's the full breakdown of what we found and what we changed.
1. Right-Sizing EC2 Instances: The Biggest Win
Our first stop was EC2. We'd provisioned most of our instances during early scaling decisions — when in doubt, we went one size up. Reasonable at the time, but those decisions had never been revisited. Using AWS Compute Optimizer and CloudWatch metrics, we analysed 90 days of CPU, memory, and network utilisation across every instance. What we found was striking: roughly 60% of our instances were running at under 15% average CPU utilisation. Several production-adjacent staging boxes were sitting at under 5%.
We right-sized 14 instances — mostly dropping from m5.xlarge to m5.large, and from r5.2xlarge to r5.xlarge for our memory-optimised workloads. We ran each change through a 48-hour observation window before moving to the next. No incidents. No latency regressions. The combined saving on EC2 alone accounted for roughly 18% of our total monthly bill reduction.
2. Switching to Savings Plans and Reserved Instances
We'd been running the majority of our workloads on On-Demand pricing — which made sense when we were uncertain about traffic patterns, but our usage had been stable for over eight months. Switching our baseline compute to a 1-year Compute Savings Plan (no upfront) gave us an immediate 30–36% discount on covered usage with zero commitment to specific instance types or regions. For our RDS instances, which are even more predictable, we moved to 1-year Reserved Instances and locked in a 38% reduction on database costs.
The key insight here is that Savings Plans are flexible — they apply across instance families and regions automatically. We didn't need to predict exactly which instance type we'd be using six months from now. We just needed to commit to a consistent hourly spend, which our historical data made easy to calculate with confidence.
3. Eliminating Idle and Orphaned Resources
This was the most embarrassing category — and probably the most common. Our audit surfaced 11 unattached EBS volumes totalling 2.3 TB, 4 unused Elastic IPs (each billed hourly when not associated), 3 forgotten load balancers with zero targets, and 6 old snapshots that had outlived their parent volumes by over a year. None of these were doing anything. All of them were billing us every hour. We wrote a short Lambda function to flag unattached volumes and idle EIPs on a weekly schedule, so this category of waste can't quietly accumulate again.
import boto3
def lambda_handler(event, context):
ec2 = boto3.client('ec2')
findings = []
# Flag unattached EBS volumes
volumes = ec2.describe_volumes(
Filters=[{'Name': 'status', 'Values': ['available']}]
)['Volumes']
for vol in volumes:
findings.append({
'type': 'UnattachedEBS',
'id': vol['VolumeId'],
'size_gb': vol['Size'],
'created': str(vol['CreateTime'])
})
# Flag unassociated Elastic IPs
addresses = ec2.describe_addresses()['Addresses']
for addr in addresses:
if 'AssociationId' not in addr:
findings.append({
'type': 'IdleElasticIP',
'id': addr['AllocationId'],
'ip': addr['PublicIp']
})
print(f"Found {len(findings)} idle resource(s):")
for f in findings:
print(f)
return findings4. Optimising S3 Storage Classes
We store a significant volume of assets, backups, and log archives in S3. Everything was sitting in S3 Standard — including objects that hadn't been accessed in over 180 days. By enabling S3 Intelligent-Tiering on our primary asset buckets and applying Lifecycle Policies to move objects older than 90 days to S3 Standard-IA (Infrequent Access) and anything older than 365 days to S3 Glacier Instant Retrieval, we reduced our S3 costs by approximately 52% month-over-month. The Intelligent-Tiering monitoring fee is $0.0025 per 1,000 objects — negligible compared to the savings at our storage volume.
5. Data Transfer: The Hidden Cost We Almost Missed
Data transfer costs are notoriously easy to overlook because they don't appear on the main service line — they're buried in the EC2 bill under 'Data Transfer Out'. We discovered two services that were communicating across Availability Zones unnecessarily, generating inter-AZ transfer charges on every request. Pinning those services to the same AZ for non-redundant internal traffic, and routing large S3 responses through a CloudFront distribution instead of directly to clients, reduced our data transfer bill by 31% in the following month.
What We Track Now (So It Doesn't Creep Back)
Saving 40% is only meaningful if you don't spend the next six months drifting back up. We now have AWS Budgets configured with alerts at 80% and 100% of our monthly baseline, Cost Anomaly Detection enabled on all services (it caught an accidental NAT Gateway misconfiguration within hours last month), and a fortnightly cost review as a standing calendar item. Our idle resource Lambda runs every Sunday night and posts a Slack summary. The audit that generated these savings took about two weeks of part-time effort. The tooling we put in place to keep it this way took another three days. That's a worthwhile trade for every team running workloads on AWS.
