Systems ThinkingSeptember 28, 2026

Modeling Cache Expiry as a Mental Model for Incident Avoidance

E

Written by

Elena Holos

Last year I helped debug an incident where a caching layer “looked fine” in dashboards—until the moment it wasn’t. The weird part was that the failures didn’t happen randomly. They clustered around specific time boundaries, and the clustering pattern matched the exact mental model the system designers didn’t use.

That pushed me to build a tiny simulation: a cache that expires entries, plus clients that read from it on a schedule. I wanted to understand the mental model behind the failures: not “caches are fast,” but “a cache has time-shaped memory,” and your expiration policy literally sculpts future traffic.

The specific failure pattern I chased

In our system:

  • The cache stores computed responses.
  • Each entry has an expiry time (ttl).
  • Requests arrive in bursts because upstream schedules jobs on the clock.
  • When items expire, the first request after expiry recomputes (a “cache miss”).
  • Recomputes can be expensive, so a burst of misses can overwhelm a backend.

The dashboard showed healthy cache hit rates most of the time. But during incidents, hit rate graphs were deceptive: the harmful behavior came from synchronized expiry, not overall average performance.

That’s the mental model I needed:

Expiry doesn’t just remove data; it creates a predictable wavefront of recomputation.

A concrete mental model: “expiry waves”

Here’s the mental model in code terms:

  • Each cache key gets a deadline = now + ttl when computed.
  • After deadline, the cache behaves like it has “forgotten” until a request arrives to repopulate.
  • If many keys share similar deadlines (common when TTL starts at the same time), then misses happen together.

So I built a simulator where:

  • There are N keys.
  • Each key is requested by some schedule that can synchronize with cache expiry.
  • The simulator advances time in discrete steps and logs backend recomputations.

Step-by-step simulation (working code)

This is a small, runnable Python program. It models a cache with TTL and a backend recomputation cost. It intentionally uses integer time units so the “wave” effect is easy to see.

import random from dataclasses import dataclass, field @dataclass class CacheEntry: value: int expires_at: int @dataclass class CacheSimulator: num_keys: int = 50 ttl: int = 30 # seconds horizon: int = 200 # seconds tick: int = 1 # seconds per simulation step backend_latency: int = 2 # seconds spent recomputing # Statistics backend_recomputes: int = 0 backend_busy_until: int = 0 recompute_times: list = field(default_factory=list) def run(self, request_schedule): """ request_schedule[t] = list of keys requested at time t. """ cache = {} # key -> CacheEntry backend_queue_starts = [] # when a recompute attempt starts (for inspection) for now in range(0, self.horizon, self.tick): # expire entries (lazy checking also works, but this makes the model explicit) # We'll keep it lazy by checking expires_at on access. for key in request_schedule.get(now, []): entry = cache.get(key) is_miss = (entry is None) or (entry.expires_at <= now) if is_miss: # Attempt recompute on miss self.backend_recomputes += 1 start = max(now, self.backend_busy_until) finish = start + self.backend_latency self.backend_busy_until = finish backend_queue_starts.append(start) self.recompute_times.append(now) # Write to cache with a new expiry deadline cache[key] = CacheEntry( value=random.randint(1, 10_000), expires_at=now + self.ttl ) # If hit, do nothing (cache value already present) return backend_queue_starts def build_synchronized_schedule(num_keys, horizon, period, burst_size, jitter=0): """ Creates a schedule where: - Every 'period' seconds, a burst_size subset of keys is requested. - Optionally adds jitter to the burst time. This simulates upstream systems that align requests with the clock. """ schedule = {} keys = list(range(num_keys)) for t in range(0, horizon, period): burst_time = t + (random.randint(-jitter, jitter) if jitter else 0) if burst_time < 0 or burst_time >= horizon: continue burst_keys = random.sample(keys, burst_size) schedule.setdefault(burst_time, []).extend(burst_keys) return schedule def build_dephased_schedule(num_keys, horizon, period, burst_size, jitter): """ Similar to synchronized schedule but with jitter large enough to de-sync bursts. That breaks expiry-wave synchronization. """ schedule = {} keys = list(range(num_keys)) for t in range(0, horizon, period): burst_time = t + random.randint(0, jitter) if burst_time >= horizon: continue burst_keys = random.sample(keys, burst_size) schedule.setdefault(burst_time, []).extend(burst_keys) return schedule if __name__ == "__main__": random.seed(3) # Scenario A: synchronized bursts -> aligned recompute waves sim_a = CacheSimulator(ttl=30, horizon=200, backend_latency=2) schedule_a = build_synchronized_schedule( num_keys=50, horizon=200, period=30, burst_size=25, jitter=0 ) starts_a = sim_a.run(schedule_a) # Scenario B: dephased bursts -> recompute waves smeared out sim_b = CacheSimulator(ttl=30, horizon=200, backend_latency=2) schedule_b = build_dephased_schedule( num_keys=50, horizon=200, period=30, burst_size=25, jitter=10 ) starts_b = sim_b.run(schedule_b) print("=== Synchronized bursts ===") print("backend recomputes:", sim_a.backend_recomputes) print("max concurrent pressure proxy (queue starts):", max(starts_a) if starts_a else None) print("recompute attempts at times (sample):", sim_a.recompute_times[:20], "...") print() print("=== Dephased bursts ===") print("backend recomputes:", sim_b.backend_recomputes) print("max concurrent pressure proxy (queue starts):", max(starts_b) if starts_b else None) print("recompute attempts at times (sample):", sim_b.recompute_times[:20], "...")

What each block is doing (and why)

  • CacheEntry stores a value and expires_at.
  • CacheSimulator.run() advances now from 0 to horizon.
  • For each requested key:
    • If no entry exists, it’s a miss.
    • If entry.expires_at <= now, it’s also a miss—the cache has “forgotten.”
  • On a miss:
    • We simulate a backend recomputation with a backend_busy_until timestamp.
    • We update the cache with expires_at = now + ttl.

The simulation logs recompute times so we can see if misses cluster.

Running it and interpreting the output

When I run this locally, I typically see:

  • Synchronized bursts (period matches TTL cadence) cause recomputes in visible clusters. Many keys get recomputed around the same times, and the backend “pressure” spikes repeatedly.
  • Dephased bursts spread requests out. Even though TTL is the same, recompute deadlines scatter, so misses are less synchronized.

Even without fancy plotting, the mental model becomes obvious: the system is behaving like a metronome when schedules align.

The mental model upgrade: TTL + clock alignment = deterministic storms

This is the part that stuck with me.

TTL alone is a local rule, but when you combine it with global timing (like upstream job schedules, cron-like traffic, or periodic batch readers), TTL becomes a synchronizer.

That means:

  • If many cache entries are first populated around the same time, they’ll expire around the same time.
  • When they expire, the first post-expiry request repopulates each entry.
  • If upstream traffic also arrives in bursts on the same period, you get expiry-wavefront storms.

This is a systems thinking loop:

  • Local policy (TTL) shapes global behavior (miss bursts),
  • which then impacts coordination (backend load),
  • which feeds back into incident culture and operational choices.

A tiny twist: jitter the TTL start time

Many caches implement TTL with some jitter (randomization). I modeled a similar idea: instead of expires_at = now + ttl, I add a small random offset to each write. That breaks alignment.

Here’s a minimal patch to the simulator: modify the expiry deadline on write.

import random from dataclasses import dataclass, field @dataclass class CacheEntry: value: int expires_at: int @dataclass class CacheSimulator: num_keys: int = 50 ttl: int = 30 horizon: int = 200 tick: int = 1 backend_latency: int = 2 ttl_jitter: int = 0 # <= new backend_recomputes: int = 0 backend_busy_until: int = 0 recompute_times: list = field(default_factory=list) def run(self, request_schedule): cache = {} for now in range(0, self.horizon, self.tick): for key in request_schedule.get(now, []): entry = cache.get(key) miss = (entry is None) or (entry.expires_at <= now) if not miss: continue self.backend_recomputes += 1 start = max(now, self.backend_busy_until) self.backend_busy_until = start + self.backend_latency self.recompute_times.append(now) jitter = random.randint(0, self.ttl_jitter) if self.ttl_jitter else 0 cache[key] = CacheEntry( value=random.randint(1, 10_000), expires_at=now + self.ttl + jitter # <= new ) return self.recompute_times def build_synchronized_schedule(num_keys, horizon, period, burst_size): schedule = {} keys = list(range(num_keys)) for t in range(0, horizon, period): burst_keys = random.sample(keys, burst_size) schedule.setdefault(t, []).extend(burst_keys) return schedule if __name__ == "__main__": random.seed(3) schedule = build_synchronized_schedule( num_keys=50, horizon=200, period=30, burst_size=25 ) sim_no_jitter = CacheSimulator(ttl=30, horizon=200, backend_latency=2, ttl_jitter=0) times_no = sim_no_jitter.run(schedule) sim_with_jitter = CacheSimulator(ttl=30, horizon=200, backend_latency=2, ttl_jitter=10) times_jitter = sim_with_jitter.run(schedule) print("Recompute attempt counts:") print(" no jitter:", len(times_no)) print(" with jitter:", len(times_jitter)) print() print("First 30 recompute attempt times (no jitter):", times_no[:30]) print("First 30 recompute attempt times (with jitter):", times_jitter[:30])

When I added jitter, the recompute attempt times stopped lining up as tightly with the same boundaries. That’s the mental model payoff: jitter doesn’t reduce misses in theory; it reduces synchronized misses in practice.

How this ties back to mental models and incident culture

This mental model (“expiry waves”) changed how I read operational data:

  • I stopped treating cache hit rate as the only truth.
  • I started asking: Are misses clustered in time?
  • I started looking for clock-based alignment across services (batch schedules, cache warmups, TTL settings, refresh jobs).

And culturally, it changed post-incident writeups. Instead of “we need better autoscaling,” the discussion shifted toward “we need to remove synchronized failure modes.” That’s a systems thinking lens applied through a mental model.

Conclusion

I built a small TTL cache simulator to make a hidden assumption visible: cache expiry isn’t just local cleanup—it shapes future traffic over time. Once I treated expiration as “time-shaped memory” and watched how synchronized schedules create expiry waves, the incident pattern finally made sense. The main lesson is that mental models should explain timing, because in real systems, timing is often where the failures hide.