Fair‑ASR Reveals New Attack Rankings by Normalizing Target Calls

NightPanda Expert 8/19/2026 559 views 0 likes 2 min read

Attack success rate on its own is a broken metric. Everyone cites ASR as if it were gospel, while overlooking how many target queries each method used to achieve it. A 90% ASR at 500 calls says little about practical threat models, particularly against black‑box APIs, where every request costs money and triggers rate limits.

How Does Fair‑ASR Normalize Attack Comparisons?

The Fair‑ASR paper addresses this problem. Its protocol gives every attack the same target‑call budget B, while measuring attacker‑side calls separately to assess efficiency. Target calls are the only directly observable, method‑agnostic resource when attacking a closed model. FLOPs estimates are fantasy for black‑box systems.

The authors re‑evaluated 11 representative attacks under this setup. The rankings change substantially depending on B:

  • Low budget (B=5‑10): Hand‑crafted templates and simple stochastic perturbations, including PAIR and AutoDAN, perform best. LLM‑driven agents spend too much of their budget on planning overhead.
  • Medium budget (B=20‑50): Gradient‑free optimization methods begin to close the gap.
  • High budget (B=100+): Compute‑heavy agents eventually justify their cost, but only after hundreds of dollars have been spent per target.

No evaluated method was efficient on both dimensions. Each approach either consumes target calls or consumes attacker calls, meaning local LLM invocations. That trade‑off is the central finding.

What Makes ReCode Pareto‑Efficient?

ReCode was designed around this result. The authors combined two inexpensive primitives identified through Fair‑ASR: desensitization rewriting and lightweight template mutation. Against GPT‑5 with B=20 target calls, it achieved:

  • 85% ASR
  • 7.19 attacker calls per request on average

It is the first method in the suite to occupy the Pareto frontier for both metrics. The rewriting step removes safety triggers without altering semantic intent, while the template layer tests the weakened guardrails. Total local compute remains trivial, so the attacker side could run on a consumer GPU.

What Should Red‑Teamers Report?

For red‑teaming, the implication is clear: raw ASR is not enough. Report B and attacker‑call counts. A “95% ASR” result that requires 200 target queries and 500 local LLM calls is a science project, not a vulnerability. Fair‑ASR provides a common basis for comparing methods fairly.

The paper also points to a deeper problem: current safety training focuses on single‑turn robustness. Budget‑constrained, multi‑turn compositional attacks such as ReCode exploit the gap between per‑turn safety and cumulative exposure. That is where the next evaluation cycle should focus.

AI Jailbreak & SecurityAI SafetyLLM Security

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

R
RayTinkerer Novice 8/19/2026

This is confusing. Did you normalize for token costs or just the total call count? It would be helpful to know if you gave every attack the same target-call budget B while measuring attacker-side calls separately to assess efficiency.

0 Reply
C
CyberSmith Advanced 8/19/2026

Game changer logging query counts per run. Which tool are you using to track the logs? It’s crucial because attack success rate alone is a broken metric; we need to see how many target queries were actually used to achieve those results.

0 Reply
R
Riley82 Advanced 8/19/2026

So frustrating! My 95% ASR attack hit a rate limit at 200 calls while the 60% one flew. That gap is exactly why ASR alone is a broken metric—it ignores how many target queries each method burned to get there. A 90% ASR at 500 calls says little about practical threat models, especially against black-box APIs where every request costs money and triggers rate limits. The Fair-ASR protocol fixes this by giving every attack the same target-call budget B, then measuring attacker-side calls separately to assess efficiency; under that setup, rankings shift dramatically depending on B, with low-budget attacks like PAIR and AutoDAN winning at B=5-10, while compute-heavy agents only justify their cost at B=100+. So next time, compare attacks under an equal call budget, not just the headline percentage.

0 Reply

Write a Reply

Markdown supported