Amazon Web Services

Caching — ElastiCache for Redis & Memcached

Put a managed in-memory cache in front of a database, and handle invalidation and failure properly.

ElastiCache runs managed Redis or Memcached, used to keep hot data out of the database.

A chef keeping prepared ingredients within arm's reach instead of walking to the cold store for every order — and throwing anything out that has been sitting there too long.

Key Concepts

1
    app -> cache  [hit]  -> return in under a millisecond
        -> cache  [miss] -> database -> write to cache -> return
2
Cache-aside is the pattern to describe.
    value = cache.get(key)
    if value is None:
        value = db.query(...)
        cache.set(key, value, ttl=300)
    return value
3
The application owns the cache, so a cache outage degrades latency rather than breaking correctness.
4
Redis or Memcached.
    Redis       data structures, persistence, replication, pub/sub,
                sorted sets, transactions, Lua. Almost always the choice.
    Memcached   plain key/value, multi-threaded, trivially simple.
                Only when you want a pure string cache and more cores.
5
Invalidation is the hard part, and every interviewer asks it.
6
    TTL only          simple, serves stale data for up to the TTL
    write-through     update the cache on every write -- always fresh,
                      slower writes, caches data nobody reads
    delete-on-write   evict the key and let the next read repopulate
7
Delete-on-write is the usual compromise.
8
The three failure modes worth naming.
    stampede      a hot key expires and a thousand requests hit the
                  database at once. Fix with a short lock or
                  staggered TTLs.
    penetration   repeated misses for a key that does not exist.
                  Cache the negative result.
    avalanche     many keys expire together. Add jitter to the TTL.
9
Eviction policy decides what goes when memory fills. allkeys-lru is the usual choice for a cache; noeviction turns a full cache into write errors, which surprises people.
allkeys-lrunoeviction
10
Multi-AZ with automatic failover promotes a replica when the primary fails. Without it, a node failure is a cold cache and a sudden full load on the database.
11
And know what not to cache. Data that must be exactly right at the moment it is read — account balances, stock counts at checkout — belongs in the database.
12
What the interviewer is probing.1. "Describe cache-aside." Probing: whether you can state the flow and its benefit. Stalls: "The cache sits in front." Moves up: read the cache, fall back to the database on a miss and populate it; the application owns the cache, so an outage costs latency rather than correctness.
13
2. "A popular key expires and the database spikes. What happened and how do you fix it?" Probing: the stampede. Stalls: "Too much traffic." Moves up: every request missed simultaneously; take a short lock so one request repopulates, and jitter TTLs so keys do not expire together.
14
3. "How do you handle repeated lookups for a key that does not exist?" Probing: cache penetration. Stalls: "Nothing to cache." Moves up: cache the negative result with a short TTL, or every miss reaches the database.
15
4. "Redis or Memcached?" Probing: whether the default is justified. Stalls: "Memcached is simpler." Moves up: Redis almost always — data structures, replication, persistence and atomic operations; Memcached only for a pure multi-threaded string cache.