Amazon Web Services

Cost Optimisation & the Well-Architected Framework

Cut spend without cutting capability, and know the six pillars an architecture review is scored against.

The Well-Architected Framework is six pillars, and interviewers do ask you to name them.

A household bill. The savings are not in buying a cheaper kettle but in the immersion heater nobody remembered was on, the subscription nobody cancelled, and the standing charge nobody read.

Key Concepts

1
    Operational excellence   run and monitor, improve procedures
    Security                 identity, detection, protection
    Reliability              recover from failure, meet demand
    Performance efficiency   use the right resources
    Cost optimisation        avoid unnecessary spend
    Sustainability           reduce the environmental footprint
2
Pricing models, roughly in order of saving.
    On-Demand        no commitment, highest rate
    Savings Plans    commit to a $/hour for 1 or 3 years. Up to ~72%.
                     Compute Savings Plans apply across EC2, Fargate
                     and Lambda, and across instance families.
    Reserved         older, tied to an instance family
    Spot             spare capacity, up to 90% off, reclaimed with
                     a 2-minute warning
    Graviton         ARM instances, commonly 20-40% better price
                     for performance on a simple recompile
3
The standard shape of a real bill. Steady baseline on a Savings Plan, burst on On-Demand, batch and CI on Spot.
4
Where the money actually leaks.
    idle or oversized instances       rightsizing reports
    unattached EBS volumes            deleted instance, kept disk
    old snapshots                     nobody owns them
    NAT Gateway data processing       often a surprise line item
    cross-AZ traffic                  charged in both directions
    S3 without lifecycle rules        everything stays Standard forever
    forgotten non-production          dev running at the weekend
5
S3 lifecycle rules are the easiest large saving. Move objects to Infrequent Access after 30 days, Glacier after 90, expire after a retention period — and use Intelligent-Tiering when access patterns are unknown.
6
Tag everything and enforce it. Without team, env and service tags, Cost Explorer cannot tell anyone which team spent what, and nothing improves because nobody owns the number.
teamenvservice
7
Measure before optimising. Compute is usually the biggest line, but data transfer and NAT charges are where the surprises hide.
8
What the interviewer is probing.1. "Name the Well-Architected pillars." Probing: whether you know the framework by name. Stalls: "Cost and security." Moves up: operational excellence, security, reliability, performance efficiency, cost optimisation and sustainability.
9
2. "Where does cloud spend usually leak?" Probing: the unglamorous items. Stalls: "Oversized instances." Moves up: unattached EBS volumes, old snapshots, NAT Gateway data processing, cross-AZ traffic, buckets with no lifecycle rules, and non-production left running.
10
3. "Savings Plans or Spot?" Probing: matching the instrument to the workload. Stalls: "Whichever is cheaper." Moves up: Savings Plans for the steady baseline you are confident of running, Spot for interruptible burst and batch — usually both.
11
4. "Why does tagging matter for cost?" Probing: accountability. Stalls: "For organisation." Moves up: without team and environment tags, Cost Explorer cannot attribute spend, and nothing improves while nobody owns the number.