Fix the design · Incident triage
Silent Queue Backlog: The Job Queue That Grew for Six Hours
During Flipkart Big Billion Days prep, a deploy slowed workers to about 40% speed while producers kept publishing — consumer lag grew unnoticed for about six hours and the oldest jobs waited over four hours.
You are on Flipkart’s order-jobs team during Big Billion Days prep. A deploy slows workers to about 40% speed, but nobody notices for six hours. The oldest jobs wait over four hours.
We need your help. Identify the bottlenecks and failure modes in the current design, then redesign the path from “order accepted” → “job processed on time” so lag is visible, workers can shed or isolate poison messages, and producers slow down when consumers fall behind.
Problem
Redesign Flipkart’s existing order-fulfillment job path — it already publishes background jobs to Kafka for workers to consume and process.
Incident summary
- After a deploy, workers slowed to about 40% of normal throughput.
- Producers kept publishing at full speed.
- Consumer lag grew for about six hours before anyone noticed.
- The oldest job sat unprocessed for about 4.2 hours.
- Order fulfillment missed SLAs. A poison message on one partition blocked progress on that shard.
Assumptions made by the team
- Kafka lag is normal during traffic spikes.
- Workers will add capacity on their own eventually.
- HTTP 200 means the system is healthy end to end.
- Poison messages (jobs that always crash the worker) are rare.
Impacted services
- Kafka consumers (critical) — Fell behind for about six hours with no alert.
- Worker pool (critical) — Ran at about 40% throughput after the deploy.
- Order fulfillment (critical) — Missed SLAs while jobs sat unprocessed.
- Producer APIs (degraded) — Looked healthy because HTTP 200 only means accepted, not finished.
Triage questions
- Why did the job queue grow for six hours without anyone being alerted?
- How would you redesign the pipeline so producers slow down when workers fall behind?
- How does one bad message stall a whole partition — and how would you isolate it?
- What should you alert on first: queue depth, consumer lag, or age of the oldest job?