Search... ⌘K
Write Log In Sign Up

Join DevSolved

Share war stories & investigate outages.

Sign Up for Free Log In to Account
ESC
CRITICAL Resolved

Redis Cache Stampede: The Thundering Herd Problem

Avatar Devsolved Feb 15, 2024 — 12:00 UTC MTTR: 8 hours 1 min read 135000 views
Executive Summary

When a highly requested, computationally expensive cache key expired, 5,000 concurrent requests instantly hit our primary database to recalculate the value, melting the database cluster.

The Symptom

Symptom

Exactly at midnight every day, our entire API would go down for exactly 3 minutes. The Postgres database connection pool would max out, CPU would hit 100%, and pgbouncer would start rejecting connections with `Pool exhausted`.

Root Cause Analysis

Root Cause

We generated a massive "Daily Global Leaderboard" that took 15 seconds to compute. We cached this in Redis with an exact TTL (Time To Live) of 24 hours, expiring right at midnight. Because our traffic is global, thousands of users hit the leaderboard endpoint at 00:00:01. They all checked Redis, found a cache miss, and *simultaneously* triggered the 15-second database query. This is a classic Cache Stampede (or Thundering Herd).

5 Whys Root Cause Drill-Down

1
Why #1
2
Why #2
3
Why #3
4
Why #4

Resolution

Resolution

We implemented "Probabilistic Early Expiration" (XFetch). Instead of a hard TTL, a background worker is responsible for calculating the leaderboard every 23 hours and 55 minutes, overwriting the cache *before* it ever expires. The cache essentially never expires from the perspective of user traffic.

Preventive Action Items

critical Move leaderboard generation to Celery background worker @Backend Team
normal Implement distributed locking using Redlock for heavy cache misses @Backend Team
|

Discussions 0

Most recent