A Kafka Rebalance Storm Caused Our UPI Integration Failure

JordanCat Expert 8/19/2026 452 views 0 likes 2 min read

Last Thursday, we implemented a routine payment gateway config change, increasing the Kafka consumer max.poll.records from 500 to 2000 to improve evening-peak performance. The entire ingestion pipeline failed within 47 seconds.

The failure began with this error:

What caused the Kafka rebalance storm?

org.apache.kafka.clients.consumer.CommitFailedException: Commit cannot be completed since the group has already rebalanced and assigned the partitions to another member. This means that the time between subsequent calls to poll() was longer than the configured max.poll.interval.ms, which typically implies that the poll loop is spending too much time processing messages. You can address this by either increasing max.poll.interval.ms or by reducing max.poll.records.

This was a classic rebalance storm, but the surprising part was that our processing logic had not changed. Each transaction still took ~12ms end-to-end. Based on those numbers, the batch should have remained within bounds: 2000 records × 12ms = 24 seconds, comfortably below the default 5-minute max.poll.interval.ms.

How did GC pauses impact consumer performance?

GC pauses revealed a different reality. At peak load, when we reached ~7,200 TPS—not quite UPI's 8k, but close—the larger batches produced more frequent young GC cycles. Every pause added 200-400ms. Across 2000 records, that was enough for the poll loop to exceed the interval. The coordinator removed the consumer, partitions moved to another member, and the replacement consumer resumed from the last committed offset. Because the previous commit had failed, that offset was stale. Duplicate processing surged, downstream idempotency keys collided, and the database connection pool was exhausted.

Here is how the cascade unfolded:

  1. Consumer A polls 2000 records
  2. GC pause hits at record 1,400 (320ms)
  3. max.poll.interval.ms exceeded → LeaveGroup sent
  4. Partition revocation → Consumer B takes over
  5. Consumer B polls from last committed offset (record 0)
  6. Duplicate processing on records 0-1,399
  7. Idempotency key conflicts on Redis → retries → more latency
  8. HikariCP pool exhausted → HTTP 500s to merchant APIs
  9. Circuit breakers trip → full outage

What changes were made to fix the issue?

We fixed the problem without reverting the configuration. We kept max.poll.records=2000, but made the following changes:

  • Dropped max.poll.interval.ms to 180000 (3 min) — tighter guardrail
  • Added max.poll.interval.ms monitoring with alert at 60% utilization
  • Switched to incremental processing: commit every 500 records inside the poll loop via commitSync() instead of waiting for batch completion
  • Tuned G1GC: -XX:MaxGCPauseMillis=100 -XX:InitiatingHeapOccupancyPercent=35
  • Added a dead-letter topic for idempotency failures so they don't block the main flow
A Kafka Rebalance Storm Caused Our UPI Integration Failure

How long did recovery take and was data lost?

Once we terminated the stuck consumers and allowed the group to stabilize, recovery took 12 minutes. No data was lost; the idempotency layer caught exactly 347 duplicates.

The UPI article reports 8,000 TPS sustained. We are running at ~1/10th that scale and still encounter these coordination problems. That gives a clearer sense of the engineering required beneath NPCI's systems, particularly its partition strategy across 300+ banks. For anyone running comparable throughput on Kafka, which max.poll.records balance have you found?

Help Wanted

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

N
Nova28 Advanced 8/19/2026

That timeout cascade is a nightmare. Did changing max.poll.interval.ms fix it permanently? You can address this by either increasing max.poll.interval.ms or by reducing max.poll.records.

0 Reply
T
TaylorDreamer Intermediate 8/19/2026

OOM kills are the worst. Did increasing the heap size actually help your throughput? Last Thursday, we pushed a routine payment gateway config change, increasing the Kafka consumer max.poll.records from 500 to 2000 in the belief that higher throughput would improve evening-peak performance. The entire ingestion pipeline collapsed within 47 seconds. Debugging the timeout cascade that killed our UPI integration The failure began with this error: ## What caused the Kafka rebalance storm?

 org.apache.kafka.clients.consumer.CommitFailedException: Commit cannot be completed since the group has already rebalanced and assigned the partitions to another member. This means that the time between subsequent calls to poll() was longer than the configured max.poll.interval.ms, which typically implies that the poll loop is spending too much time processing messages. You can address this by either increasing max.poll.interval.ms or by reducing max.poll.records.

This was a classic rebalance storm, but the surprising part was that our processing logic had not changed. Each transaction still took ~12ms end-to-end. Based on those numbers, the batch should have remained within bounds: 2000 records × 12ms = 24 seconds, comfortably below the default 5-minute max.poll.interval.ms. ## How did GC pauses impact consumer performance? GC pauses revealed a different reality. At peak load, when we reached ~7,200 TPS—not quite UPI's 8k, but close—the larger batches produced more frequent young GC cycles. Every pause added 200-400ms. Across 2000 records, that was enough for the poll loop to exceed the interval. The coordinator removed

0 Reply
L
Leo37 Novice 8/19/2026

That session.timeout.ms nightmare is the worst. Did you have to tweak the heartbeat settings too? Last Thursday, we pushed a routine payment gateway config change, increasing the Kafka consumer max.poll.records from 500 to 2000 in the belief that higher throughput would improve evening-peak performance. The entire ingestion pipeline collapsed within 47 seconds. Debugging the timeout cascade that killed our UPI integration The failure began with this error: ## What caused the Kafka rebalance storm?

 org.apache.kafka.clients.consumer.CommitFailedException: Commit cannot be completed since the group has already rebalanced and assigned the partitions to another member. This means that the time between subsequent calls to poll() was longer than the configured max.poll.interval.ms, which typically implies that the poll loop is spending too much time processing messages. You can address this by either increasing max.poll.interval.ms or by reducing max.poll.records.

This was a classic rebalance storm, but the surprising part was that our processing logic had not changed. Each transaction still took ~12ms end-to-end. Based on those numbers, the batch should have remained within bounds: 2000 records × 12ms = 24 seconds, comfortably below the default 5-minute max.poll.interval.ms. ## How did GC pauses impact consumer performance? GC pauses revealed a different reality. At peak load, when we reached ~7,200 TPS—not quite UPI's 8k, but close—the larger batches produced more frequent young GC cycles. Every pause added 200-400ms. Across 2000 records, that was enough for the poll loop to exceed the interval. The coordinator removed

0 Reply

Write a Reply

Markdown supported