Preventing Duplicate Charges: My Idempotency Workflow

Quinn20 Expert 7/24/2026 506 views 1 likes 2 min read

Network retries in LLM agents are a silent killer. When an agent calls a tool over a network and doesn't get a response in time, it often triggers a retry. If that tool isn't idempotent, you end up with duplicate charges, duplicate emails, or duplicate orders—all while your logs insist everything is functioning normally. This is a classic distributed systems problem appearing in a new AI context.

The Failure Scenario

Most agent-tool failures happen because of a mismatch between the tool's execution time and the agent's client-side timeout. Here is a concrete example of how a "naive" retry loop creates a double-charge bug:

import threading
import time

ledger = {"order_123": 0}

def charge_card(order_id, amount):
    # Simulate a slow payment API response
    time.sleep(0.4) 
    ledger[order_id] += 1
    return {"order_id": order_id, "status": "charged"}

def agent_charge_with_naive_retry(order_id, amount, max_retries=2):
    for attempt in range(max_retries):
        thread = threading.Thread(target=lambda: charge_card(order_id, amount))
        thread.start()
        # The agent's client-side timeout is shorter than the API response time
        thread.join(timeout=0.2) 
        if thread.is_alive():
            # Agent assumes failure and retries, but the first thread is still running
            continue 
        return 

# Execution
agent_charge_with_naive_retry("order_123", 100.0)
print(f"Total charges in ledger: {ledger['order_123']}") 
# Result: Total charges in ledger: 2

In this snippet, the payment call takes 0.4s, but the agent gives up after 0.2s. Because Python threads cannot be safely killed mid-execution, the first request completes in the background while the second request starts. The agent eventually sees a "success" from the retry, but the user is charged twice.

Implementing Idempotency with Latch

To solve this, I developed a lightweight Python library called latch. It implements idempotency keys—a standard backend practice—specifically for the agent-tool-calling pattern. Instead of guessing if two calls are the same based on arguments (which is risky), it requires an explicit idempotency_key.

Here is the practical deployment of the @idempotent decorator:

from latch import idempotent

# Mock API for demonstration
class PaymentsAPI:
    def charge(self, order_id, amount):
        print(f"Processing charge for {order_id}...")
        return {"status": "success"}

payments_api = PaymentsAPI()

@idempotent()
def create_order(order_id: str, amount: float, idempotency_key: str) -> dict:
    return payments_api.charge(order_id, amount)

# First call: executes the function
create_order(order_id="A1", amount=42.0, idempotency_key="run-7-step-3")

# Second call with same key: returns cached result, no second API call
create_order(order_id="A1", amount=42.0, idempotency_key="run-7-step-3")

Expanding the AI Workflow

Idempotency is just the start. To build a production-ready LLM agent, you need a full set of guardrails to prevent "runaway" agents. I've expanded the library to include these specific patterns:

  • @idempotent: Stops duplicate side effects during retries.
  • @circuit_breaker: Prevents the agent from hammering a dependency that is already returning 500s.
  • @with_timeout: Ensures a hung external API call doesn't block the entire agent loop indefinitely.
  • @budget_guardrail: Stops a loop from burning through your API credits by capping total cost per session.
  • Saga Pattern (Saga class): Manages multi-step transactions (e.g., Charge → Book Flight → Book Hotel). If step 3 fails, it triggers compensating transactions to undo steps 1 and 2.

For anyone building a real-world AI workflow, treating tool calls as unreliable distributed events rather than local function calls is the only way to ensure data consistency.
PromptAIopensourcepython
More reusable prompt workflows are gathered in a practical ChatGPT prompt guide, with plenty of directly applicable cases.

All Replies (2)

C
Casey51 Novice 7/24/2026
I started passing a unique request ID from the agent side to handle this. Saves a lot of headache with Stripe API calls.
0 Reply
T
TaylorDreamer Intermediate 7/24/2026
What happens if the idempotency key expires before the retry actually hits? Do you have a TTL limit on those records?
0 Reply

Write a Reply

Markdown supported