Amazon S3 storage classes and lifecycle rules
Every object in Amazon S3 has a storage class, and the class decides three things: what you pay to store a gigabyte, what you pay to read it back, and how quickly it comes back. The cheapest storage is never the cheapest bill if the data is read often, and the fastest class is wasted money for data nobody touches. Picking the right class is a matter of matching it to how the data is actually used.
Exam questions on this topic are nearly always cost questions with a catch. The scenario gives an access pattern ("read heavily for 30 days, then rarely"), a retrieval requirement ("within milliseconds" or "within 48 hours") and sometimes a resilience requirement ("can be regenerated"), and three of the four options break one of them. If you know each class's retrieval time, retrieval fee, minimum storage duration and number of Availability Zones, most of these questions answer themselves.
The storage classes
Frequently accessed data
- S3 Standard is the default. It stores data across at least three Availability Zones, has no retrieval fee, no minimum storage duration and no minimum object size. It costs the most per gigabyte stored, and that is the right trade for data that is read regularly.
- S3 Express One Zone is a high-performance class for latency-sensitive workloads. It gives consistent single-digit millisecond access from directory buckets in one Availability Zone that you choose. It is about speed, not saving money on storage.
Infrequently accessed data
- S3 Standard-IA keeps data in at least three zones with millisecond access, but charges a per-GB retrieval fee. It has a 30-day minimum storage duration and a 128 KB minimum billable object size: a 10 KB object is billed as 128 KB.
- S3 One Zone-IA has the same fees and minimums, stores data in one Availability Zone and costs less. The data is lost if that zone is physically destroyed, so it suits copies you can re-create, such as thumbnails or secondary replicas.
Archive data
- S3 Glacier Instant Retrieval is for data read about once a quarter that still needs millisecond access. It is built for colder data than Standard-IA (read about once a quarter rather than once a month), charges per-GB retrieval fees, and has a 90-day minimum storage duration. It has the same 128 KB minimum object size.
- S3 Glacier Flexible Retrieval is archived data. You can't read it directly; you first restore a temporary copy. Expedited restores typically take 1 to 5 minutes, Standard 3 to 5 hours, and Bulk 5 to 12 hours. Bulk retrievals are free. The minimum storage duration is 90 days.
- S3 Glacier Deep Archive is the lowest-cost storage class, for data read less than once a year. Standard restores finish within 12 hours and Bulk within 48 hours. There is no Expedited option. The minimum storage duration is 180 days.
Each object archived to Flexible Retrieval or Deep Archive also carries about 40 KB of metadata overhead, partly billed at S3 Standard rates. That is why archiving millions of tiny files can cost more than it saves.
S3 Intelligent-Tiering
Intelligent-Tiering is for data whose access pattern is unknown or changes over time. It watches each object and moves it between tiers for you:
- New objects start in the Frequent Access tier.
- After 30 consecutive days without access, an object moves to Infrequent Access.
- After 90 consecutive days without access, it moves to Archive Instant Access.
- You can also opt in to the Archive Access and Deep Archive Access tiers, for objects not read for at least 90 or 180 days. Objects in those two tiers must be restored before they can be read, so turn them on only if the application can wait.
As soon as an object in the first three tiers is read, it moves back to Frequent Access. There are no retrieval fees and no minimum storage duration. What you pay instead is a small monthly monitoring and automation fee per object. Objects smaller than 128 KB aren't monitored and stay in the Frequent Access tier, so Intelligent-Tiering does little for buckets full of tiny files.
Lifecycle rules
A lifecycle configuration is a set of rules on a bucket that S3 applies for you, with no code. A rule can be filtered by prefix, object tags or object size, and has two kinds of action:
- Transition actions move objects to a cheaper class after a number of days, for example to Standard-IA at 30 days and to Glacier Flexible Retrieval at 180 days.
- Expiration actions delete objects after a number of days. For versioned buckets, separate actions expire noncurrent versions. A rule can also abort incomplete multipart uploads, whose parts are billed but don't show up as objects in a listing.
Transitions only go "down" a waterfall: Standard to IA, Intelligent-Tiering or Glacier classes; IA to One Zone-IA or Glacier; Glacier Flexible Retrieval only to Deep Archive. A lifecycle rule can't move an object back up. To bring an archived object back to Standard, you restore it and copy it.
Two cost details matter. Each transition is a billed request, so S3 by default doesn't transition objects smaller than 128 KB. And moving an object out of a class before its minimum duration is charged as if it had stayed the full term.
How to choose
- Read often, or nobody knows how often: Standard, or Intelligent-Tiering when the pattern is unpredictable and the objects are bigger than 128 KB.
- A known pattern, such as "hot for 30 days and then cold": Standard with a lifecycle rule to the class that fits the cold period.
- Rarely read, must still come back in milliseconds, kept for months: Glacier Instant Retrieval. Kept for weeks and read monthly: Standard-IA.
- Re-creatable data that is rarely read: One Zone-IA.
- Kept for years, can wait hours: Glacier Flexible Retrieval. Can wait up to two days: Glacier Deep Archive, the cheapest.
- Data must be deleted after a set time: an expiration action. No storage class deletes anything on its own.
Common exam traps
- Ignoring the retrieval requirement. An answer that saves the most on storage but needs a restore fails a "milliseconds" requirement. Glacier Flexible Retrieval and Deep Archive are never right when the application issues a plain
GetObject. - Moving actively read data to an IA class. Per-GB retrieval fees on data that is still read every week can cost more than leaving it in Standard.
- Forgetting minimum durations. Deleting or replacing objects in Glacier Instant Retrieval at 30 days still bills 90 days, and in Deep Archive 180 days. Short-lived data belongs in Standard or Intelligent-Tiering.
- Choosing One Zone-IA for the only copy of irreplaceable data. It is as durable as Standard-IA but not resilient to losing its Availability Zone.
- Expecting Intelligent-Tiering to delete data. It only moves objects between tiers. Retention limits need a lifecycle expiration rule.
- Overlooking incomplete multipart uploads. A bill that is much bigger than the objects you can list often means abandoned upload parts, cleaned up with a lifecycle abort rule.
A worked example
A video platform stores uploaded source files and the web renditions made from them. Source files must be kept for seven years for licensing. They are read in the first two weeks while renditions are made, then almost never. If a dispute comes up, legal can wait two days for a file. Renditions are streamed heavily for a few weeks, but an old video can suddenly go viral again and must play instantly. Any rendition can be regenerated from its source.
The source files have a known pattern and a generous retrieval window. Keep them in Standard while they are being processed, then use a lifecycle rule to send them to Glacier Deep Archive after the processing window. A 48-hour Bulk restore fits "two days", the 180-day minimum is no problem for a seven-year retention, and an expiration action at seven years removes them on schedule. Glacier Instant Retrieval would also work, but it costs more to store, and the millisecond access it pays for isn't needed here.
The renditions are the opposite. Their pattern can't be predicted and they must always be readable in milliseconds. Intelligent-Tiering fits: idle renditions drift down to cheaper tiers, a viral one jumps back to Frequent Access with no retrieval fee, and the archive tiers stay switched off so nothing ever needs a restore. One Zone-IA is tempting because renditions can be regenerated, but a flat lifecycle transition would charge retrieval fees every time an old video comes back.
Practice questions
- Datasets with unpredictable access that must stay fast
- Ten-year contract retention with a 48-hour audit window
- Thumbnails that can be regenerated from the originals
- A bucket bill much larger than its listed objects
- Log files that are read for a month and deleted after a year
- Media objects whose popularity changes over time