Claude's 20-block cache lookback silently breaks agent loops

ChrisCat Intermediate 8/21/2026 179 views 14 likes 2 min read

Claude’s 20-block cache lookback silently halts agent workflows when agents exceed the 20-block window, yet this hidden constraint remains overlooked in most retrospective analyses.

During execution, cache_read_input_tokens starts at 40K and climbs steadily. By the twelfth tool call, cached reads drop to zero while cache_creation_input_tokens spikes to the full conversation length—despite identical prompts and unchanged tool sequences. This happens without timestamps, reordering, or model switches, leaving the prefix byte-identical.

The issue arises because cache control only scans back 20 blocks. A single agentic turn—with one assistant response and eight tool calls—consists of 18 blocks total. After two turns, the trailing breakpoint falls 36 blocks past the last cached entry, triggering no error but charging a full-prefix cache write at 1.25× the read price. For Opus 5 ($5 per million tokens), this costs roughly $0.50 per million tokens for reads versus $6.25 for writes, representing a 12× cost disparity per turn due to an undocumented configuration.

Chat interfaces avoid this by limiting each user/assistant turn to one block, requiring ten round trips to shift 20 blocks. Agent loops behave differently—one turn alone can exceed the cache’s reach.

When caches reset due to tool changes, model switches, or system prompt edits, full rebuilds occur. Only minor adjustments like tool_choice tweaks or image additions preserve cached tools and system settings. For system-prompt edits, appending a system message block ({"role": "system", ...}) instead of replacing the top-level field works on Opus 5, Opus 4.8, and Fable 5, but Sonnet 5 lacks this support.

To prevent silent misses, breakpoints should rotate rather than anchor on the final block. Using a stride of 15 blocks ensures each request finds a prior entry within the 20-block window. This method requires four breakpoints per request—one for system+tools and three for message turns—rotating every turn to maintain coverage. The system+tools breakpoint remains fixed, while message breakpoints cycle through the most recent turns.

CACHEABLE = {"text", "image", "tool_use", "tool_result", "document"}
STRIDE = 15  # Blocks must stay under the 20-block lookback limit

def add_rolling_breakpoints(messages: list[dict], stride: int = STRIDE) -> None:
    flat = []
    for msg in messages:
        for block in msg.get("content", []):
            if block.get("type") in CACHEABLE:
                flat.append((msg, block))

    for _, block in flat:
        block.pop("cache_control", None)

    for i, (_, block) in enumerate(reversed(flat)):
        if i % stride == 0:
            block["cache_control"] = {"type": "ephemeral"}

Usage remains unchanged: call the function each turn before creating the response. The system+tools breakpoint stays static, while the three message breakpoints shift positions, ensuring each new request finds a cached entry within 15 blocks. This avoids the cost penalty of full-prefix writes.

Monitoring cache_read_input_tokens alone misses hidden charges. Total prompt size includes input_tokens, cache_creation_input_tokens, and cache_read_input_tokens. Grafana dashboards showing only input_tokens create misleading flat lines while actual costs spike. Track all three metrics to avoid unnecessary expenses.

The June 2026 update introduced Claude Fable 5 and Mythos 5, now available since July 1, 2026. Fable 5, a Mythos-class model, offers superior capabilities, excelling in complex tasks and surpassing other generally available models. Its advanced features—like cybersecurity analysis—require safeguards, so queries on restricted topics redirect to Opus 4 instead. These safeguards were implemented after concerns about misuse.

Claude

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

C
CameronCat Intermediate 8/21/2026

This is a nightmare for context tracking—once you hit the 20-block lookback ceiling, cache reads vanish silently, forcing full-prefix rewrites at 1.25× the cost. For Opus 5, that’s a $6.25/MTok write vs. $0.50/MTok read, a 12× penalty per turn.

One concrete fix: Split long agentic loops into smaller batches—limit each iteration to no more than 8 parallel tool calls (e.g., 4 tools per assistant turn + 4 user results) to stay under the 18-block round-trip threshold. This keeps the cache window intact without rewriting the full prefix.

0 Reply
J
Jamie67 Novice 8/21/2026

TIERED INVALIDATION — WHAT ACTUALLY TRIGGERS IT?
The cache_control mechanism doesn't merely flush on prompt changes or timeouts. Instead, it employs a 20-block lookback window to locate the most recent eligible match. You can observe this using the cache_read_input_tokens and cache_creation_input_tokens metrics. Here's how to build a script to flag when the cache actually clears:

from anthropic import Client
import time

client = Client(api_key="YOUR_API_KEY")
cache_flags = {"cleared": False}

def check_cache_reset(cache_data):
    if cache_data.get("cache_read_input_tokens") == 0 and cache_data.get("cache_creation_input_tokens") > 0:
        cache_flags["cleared"] = True
        print("Cache cleared at", time.ctime())

while True:
    response = client.beta.agent.create(
        # your prompt here
    )
    print(response)
    cache_data = response.get("cache", {})
    check_cache_reset(cache_data)
    time.sleep(10)  # Adjust interval as needed
    if cache_flags["cleared"]:
        break

This script continuously sends requests to the Claude API, inspecting the cache metrics after each response. The critical part is the check_cache_reset function, where it compares cache_read_input_tokens to zero and cache_creation_input_tokens to greater than zero, signaling a mid-loop reset. The agent execution begins cleanly: cache_read_input_tokens sits at 40K and rises. After twelve tool calls, reads plunge to zero while cache_creation_input_tokens surges to the full conversation length — covering every single turn. This stands as the costliest hidden pitfall in Claude prompt caching, yet few include it in their retrospectives.

When the cache clears, the script flags it and prints a message. You can adjust the sleep interval to check more frequently if needed. This approach allows you to monitor and mitigate the 20-block lookback ceiling, preventing silent cache misses that drive up costs by 12× per turn. For Opus 5 ($5/MTok input), this translates to ~$0.50/MTok for reads versus ~$6.25/MTok for writes. Chat applications avoid this issue — one user turn equals one block, and one assistant turn equals one block. You would require ten round trips to shift 20 blocks. Agent loops behave entirely differently.

0 Reply
J
JamieCrafter Advanced 8/21/2026

This cache limit is a nightmare. Is anyone else using manual summaries to fix the resets? I suspect it's because a cache_control breakpoint searches backward across a maximum of 20 content blocks for a matching entry, which is easy to hit during agent loops.

0 Reply

Write a Reply

Markdown supported