The Symptom
Exactly at midnight every day, our entire API would go down for exactly 3 minutes. The Postgres database connection pool would max out, CPU would hit 100%, and pgbouncer would start rejecting connections with `Pool exhausted`.
Root Cause Analysis
We generated a massive "Daily Global Leaderboard" that took 15 seconds to compute. We cached this in Redis with an exact TTL (Time To Live) of 24 hours, expiring right at midnight. Because our traffic is global, thousands of users hit the leaderboard endpoint at 00:00:01. They all checked Redis, found a cache miss, and *simultaneously* triggered the 15-second database query. This is a classic Cache Stampede (or Thundering Herd).
Resolution
We implemented "Probabilistic Early Expiration" (XFetch). Instead of a hard TTL, a background worker is responsible for calculating the leaderboard every 23 hours and 55 minutes, overwriting the cache *before* it ever expires. The cache essentially never expires from the perspective of user traffic.
critical
Move leaderboard generation to Celery background worker
@Backend Team
normal
Implement distributed locking using Redlock for heavy cache misses
@Backend Team
Discussions 0