A Kafka Rebalance Storm Caused Our UPI Integration Failure
Last Thursday, we implemented a routine payment gateway config change, increasing the Kafka consumer max.poll.records from 500 to 2000 to improve evening-peak performance. The entire ingestion pipeline failed within 47 seconds.
The failure began with this error:
What caused the Kafka rebalance storm?
org.apache.kafka.clients.consumer.CommitFailedException: Commit cannot be completed since the group has already rebalanced and assigned the partitions to another member. This means that the time between subsequent calls to poll() was longer than the configured max.poll.interval.ms, which typically implies that the poll loop is spending too much time processing messages. You can address this by either increasing max.poll.interval.ms or by reducing max.poll.records.
This was a classic rebalance storm, but the surprising part was that our processing logic had not changed. Each transaction still took ~12ms end-to-end. Based on those numbers, the batch should have remained within bounds: 2000 records × 12ms = 24 seconds, comfortably below the default 5-minute max.poll.interval.ms.
How did GC pauses impact consumer performance?
GC pauses revealed a different reality. At peak load, when we reached ~7,200 TPS—not quite UPI's 8k, but close—the larger batches produced more frequent young GC cycles. Every pause added 200-400ms. Across 2000 records, that was enough for the poll loop to exceed the interval. The coordinator removed the consumer, partitions moved to another member, and the replacement consumer resumed from the last committed offset. Because the previous commit had failed, that offset was stale. Duplicate processing surged, downstream idempotency keys collided, and the database connection pool was exhausted.
Here is how the cascade unfolded:
- Consumer A polls 2000 records
- GC pause hits at record 1,400 (320ms)
max.poll.interval.msexceeded → LeaveGroup sent- Partition revocation → Consumer B takes over
- Consumer B polls from last committed offset (record 0)
- Duplicate processing on records 0-1,399
- Idempotency key conflicts on Redis → retries → more latency
- HikariCP pool exhausted → HTTP 500s to merchant APIs
- Circuit breakers trip → full outage
What changes were made to fix the issue?
We fixed the problem without reverting the configuration. We kept max.poll.records=2000, but made the following changes:
- Dropped
max.poll.interval.msto 180000 (3 min) — tighter guardrail - Added
max.poll.interval.msmonitoring with alert at 60% utilization - Switched to incremental processing: commit every 500 records inside the poll loop via
commitSync()instead of waiting for batch completion - Tuned G1GC:
-XX:MaxGCPauseMillis=100 -XX:InitiatingHeapOccupancyPercent=35 - Added a dead-letter topic for idempotency failures so they don't block the main flow
How long did recovery take and was data lost?
Once we terminated the stuck consumers and allowed the group to stabilize, recovery took 12 minutes. No data was lost; the idempotency layer caught exactly 347 duplicates.
The UPI article reports 8,000 TPS sustained. We are running at ~1/10th that scale and still encounter these coordination problems. That gives a clearer sense of the engineering required beneath NPCI's systems, particularly its partition strategy across 300+ banks. For anyone running comparable throughput on Kafka, which max.poll.records balance have you found?
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
OOM kills are the worst. Did increasing the heap size actually help your throughput? Last Thursday, we pushed a routine payment gateway config change, increasing the Kafka consumer max.poll.records from 500 to 2000 in the belief that higher throughput would improve evening-peak performance. The entire ingestion pipeline collapsed within 47 seconds.
The failure began with this error: ## What caused the Kafka rebalance storm?
org.apache.kafka.clients.consumer.CommitFailedException: Commit cannot be completed since the group has already rebalanced and assigned the partitions to another member. This means that the time between subsequent calls to poll() was longer than the configured max.poll.interval.ms, which typically implies that the poll loop is spending too much time processing messages. You can address this by either increasing max.poll.interval.ms or by reducing max.poll.records.
This was a classic rebalance storm, but the surprising part was that our processing logic had not changed. Each transaction still took ~12ms end-to-end. Based on those numbers, the batch should have remained within bounds: 2000 records × 12ms = 24 seconds, comfortably below the default 5-minute max.poll.interval.ms. ## How did GC pauses impact consumer performance? GC pauses revealed a different reality. At peak load, when we reached ~7,200 TPS—not quite UPI's 8k, but close—the larger batches produced more frequent young GC cycles. Every pause added 200-400ms. Across 2000 records, that was enough for the poll loop to exceed the interval. The coordinator removed
That session.timeout.ms nightmare is the worst. Did you have to tweak the heartbeat settings too? Last Thursday, we pushed a routine payment gateway config change, increasing the Kafka consumer max.poll.records from 500 to 2000 in the belief that higher throughput would improve evening-peak performance. The entire ingestion pipeline collapsed within 47 seconds.
The failure began with this error: ## What caused the Kafka rebalance storm?
org.apache.kafka.clients.consumer.CommitFailedException: Commit cannot be completed since the group has already rebalanced and assigned the partitions to another member. This means that the time between subsequent calls to poll() was longer than the configured max.poll.interval.ms, which typically implies that the poll loop is spending too much time processing messages. You can address this by either increasing max.poll.interval.ms or by reducing max.poll.records.
This was a classic rebalance storm, but the surprising part was that our processing logic had not changed. Each transaction still took ~12ms end-to-end. Based on those numbers, the batch should have remained within bounds: 2000 records × 12ms = 24 seconds, comfortably below the default 5-minute max.poll.interval.ms. ## How did GC pauses impact consumer performance? GC pauses revealed a different reality. At peak load, when we reached ~7,200 TPS—not quite UPI's 8k, but close—the larger batches produced more frequent young GC cycles. Every pause added 200-400ms. Across 2000 records, that was enough for the poll loop to exceed the interval. The coordinator removed
That timeout cascade is a nightmare. Did changing max.poll.interval.ms fix it permanently? You can address this by either increasing max.poll.interval.ms or by reducing max.poll.records.