HARMBENCH // FIELD MAP
← field map
MEASURING · 04 OF 04Zephyr 7B + R2D2 · GCG 69.5 → 5.5 · TAP 66.5 → 60.8 · 16 h on 8×A100

Defenses: R2D2 and the read-across caveat

HarmBench's own R2D2 defense, what it costs, and why an ASR number describes an attack, not a model.
TL;DR — HarmBench ships a defense too: R2D2, adversarial training against a pool of GCG test cases re-optimised as the model changes. Zephyr 7B's GCG ASR falls 69.5 → 5.5, past Llama 2 13B Chat's 30.0, for 1.34 points of MT-Bench. Read the same row sideways: 48.0 under PAIR, 54.3 under TAP-Transfer, 60.8 under TAP. Training against a gradient adversary bought robustness to gradient adversaries. An ASR number is a claim about the attack you ran, not about the model.
GCG ASR
69.5→5.5
12.6× lower; best model-level defense in the paper
TAP ASR
66.5→60.8
1.1× lower; same model, different column
MT-Bench
7.34→6.0
the capability bill, and it is not zero
Train cost
16 h
8×A100, 500 steps, 180 persistent test cases
Defense methods
1
of 33 targets, exactly one is a defense HarmBench implements

The prior three pages built the number: the grader, the pipeline, the grid. This page moves it down. HarmBench's authors built a defense with their own benchmark in the loop and printed its full row: strong in one place, weak in another, same line of the same table.

R2D2: adversarial training against an adversary that keeps moving

Most defended models are trained once: assemble harmful prompts, fine-tune to refuse, ship. But the attacker optimises after you stop, and GCG searches token by token against your specific weights. Putting the attacker inside the training loop is the textbook fix, but unaffordable for LLMs: GCG needs about twenty minutes per test case against a 7B model on an A100. Earlier attempts froze the adversary and got disappointing results.

Robust Refusal Dynamic Defense (R2D2) borrows vision's fix, persistent adversarial examples: a standing pool of 180 test cases. Each iteration samples 8, runs GCG 5 further steps to refresh them against the current weights, then takes a model step, so the attack is never solved from scratch nor fully stale. 20% resets every 50 updates. Total: 500 steps, roughly 16 hours on an 8×A100 node.

Base model Mistral 7B, instruction data UltraChat, harness the Zephyr 7B Beta SFT training code with R2D2 spliced into the trainer.

The two losses, and why one of them is not obvious

Each step minimises three terms. The SFT loss is ordinary UltraChat instruction tuning, present only to keep the model useful. The toward loss, −log fθ(trefusal | xi), makes a refusal likely; its refusal target is one hard-coded sentence for every test case. The third term is the interesting one:

L_away = -1 · log( 1 - f_θ(t_i | x_i) )

where ti is the harmful continuation GCG optimised toward and xi is the adversarial prompt: the exact negation of the GCG objective, pushing back down whatever GCG just made probable.

It is not redundant with the toward loss: "make refusal likely" and "make this specific completion unlikely" are different objectives, and a model can satisfy the first while failing the second. Raising a refusal string from 0.2 to 0.6 says almost nothing about a harmful continuation sitting at 0.05, and 0.05 is not safe: greedy decoding only needs the harmful branch to win token by token. The toward loss shapes the mode; the away loss attacks the specific tail. Train only the first and the model refuses confidently until a suffix flips it.

The implementation adds a detail the formula omits: the away term is floored per token, so once a token's log P falls below −5.0 (roughly 0.7% probability) its contribution is zeroed. Otherwise −log(1−p) keeps returning gradient on already-impossible tokens, and the unbounded away term swamps the SFT term and eats general ability.

The shipped defaults and the reported run disagree. The training script matches the paper on the per-step knobs (5 GCG steps per iteration, 8 cases refreshed per step, a reset every 50 steps, search width 512) but ships a pool of 8 test cases, not 180, and the repo README calls the released checkpoint (cais/zephyr_7b_r2d2) "a snapshot after 2000 steps" against the paper's M = 500. The multi-behavior path is commented out and unsupported.

The row, and the ratio column

Zephyr 7B against Zephyr 7B + R2D2 across every attack column of Table 6’s all-behaviors slice (the 400 text behaviors), sorted by how much R2D2 helped, because that ordering is the finding.

AttackFamilyZephyr 7B+ R2D2× lower
UATgradient62.30.0∞
GCG-Transfergradient61.10.0∞
GBDAgradient62.80.2314
PEZgradient62.52.921.6
GCG-Multigradient62.54.912.8
GCGgradient — the train-time adversary69.55.512.6
AutoPromptgradient60.55.511.0
ZeroShotstatic, LLM-written60.07.28.3
Human Jailbreaksfixed templates66.013.64.9
DirectRequestno attack at all65.814.24.6
AutoDANLLM-driven, genetic75.017.04.4
Stochastic Few-ShotLLM-driven, adaptive62.043.51.4
PAP-top5persuasion taxonomy32.924.31.4
TAP-TransferLLM-driven, tree search69.354.31.3
PAIRLLM-driven, iterative58.848.01.2
TAPLLM-driven, tree search66.560.81.1
Mean of the 16 columns62.318.93.3

∞ means the column reached exactly 0.0. The × column is arithmetic on the two printed columns, not reported by the paper; the last row is the plain mean of the sixteen cells. Table 11 gives 62.7 and 19.1 for the same models, a slightly different "Average ASR" slice: the same mismatch that makes Zephyr's GCG ASR 69.4 there and 69.5 here.

5.5 under GCG is the best any model-level defense in the paper achieves against its strongest attack, beating Llama 2 13B Chat's 30.0 and 70B Chat's 37.5, both products of full industrial RLHF. The denominator is the 400 text behaviors (200 standard, 100 contextual, 100 copyright); the 110 multimodal behaviors sit in Table 9. It includes the 100 copyright behaviors, MinHash-graded near the floor for every model and attack (see the copyright category), which pull both rows down about equally, so the comparison is fair but not a uniform compliance rate.

What it cost: Table 11

ModelMT-BenchGCG ASRAverage ASR
Zephyr 7B (SFT + DPO)7.3469.462.7
Mistral 7B Instruct v0.26.569.155.7
Koala 13B5.462.248.5
Zephyr 7B + R2D26.05.519.1

R2D2 lands at 6.0, 1.34 MT-Bench points below its own Zephyr base, below Mistral 7B's 6.5, above Koala 13B's 5.4. The paper calls this "does not necessarily harm performance"; it is a 1.34-point drop on a ten-point scale. The robustness gain is enormous and the capability cost is not zero: 43.6 points of average ASR traded for 1.34 of MT-Bench.

The paper states an asymmetry plainly. The main-table Zephyr 7B is Zephyr 7B Beta, SFT and DPO; R2D2 is built on the SFT-only variant, with no DPO stage. So the 7.34 → 6.0 gap is not purely the price of adversarial training, part is the missing preference-optimisation stage, and the same asymmetry sits under the 69.5 → 5.5 robustness claim. This is a preliminary demonstration, not a controlled ablation, and the paper says so.

Read the row across, not down

Ignore the absolute numbers and read the table by group. The top (UAT, GCG-Transfer, GBDA, PEZ, GCG-Multi, GCG, AutoPrompt) is the entire gradient family, methods that search token sequences using the model's own gradients; all seven collapse, five below 5%. The bottom (TAP, PAIR, TAP-Transfer, PAP-top5, Stochastic Few-Shot) is almost exactly the LLM-driven family, where another model refines a natural-language jailbreak from what the target did last. Those barely move: 12.6× in one column, 1.1× in another, same model and training run.

The split is not simply gradient versus not. DirectRequest, Human Jailbreaks and ZeroShot also fall far (4.6×, 4.9×, 8.3×) with no gradients. What they share is being non-adaptive: a fixed request, template, or blind prompt, which plain refusal training (the toward loss) handles for free. What survives is the intersection: attacks that adapt to this model in a space R2D2 never trained against. GCG adapts in token space, where the defense lives; PAIR and TAP adapt in semantic space, where it has nothing to say.

An ASR number is a claim about the attack you ran, not about the model. "5.5% ASR" and "60.8% ASR" are both true of the same weights, the same 510 behaviors, the same classifier, differing only in which attack produced the test cases. A robustness figure without its attacks named is a coincidence about which attacks the authors ran, which correlates with what they defended against. That is the practical content of the finding that no model is robust to all attacks and no attack breaks all models: a robustness claim needs an attack list or it means nothing.

The paper says this about its own method, in the section announcing its success:

"For some attacks, the improvement conferred by R2D2 is less pronounced. This is especially true for methods dissimilar to the train-time GCG adversary, including PAIR, TAP, and Stochastic Few-Shot. This suggests that incorporating multiple diverse attacks into adversarial training may be necessary to obtain generalizable robustness."— HarmBench, §6, Adversarial Training Results

A normal defense paper tables the attacks it handles and explains why the rest were out of scope. Here the weak columns print at the same width as the strong ones, possible only because the authors fixed the battery before building the defense.

One caution on the deltas: the grader agrees with humans about 93% of the time, so roughly 7% label noise sits under every cell. Far too little to threaten 69.5 → 5.5, but enough that 66.5 versus 60.8 is a difference you should not lean on. The read-across holds because the extremes differ by an order of magnitude.

The other "defenses" among the 33 targets

"33 target LLMs and defenses" invites a misreading. R2D2 is the only defense method HarmBench implements. Everything else is a shipped model its vendor safety-trained: a real category, since safety training is a model-level defense, but with no defense code to reproduce. Three groups from Table 6:

Read Claude's row twice. A near-zero row fits two very different models: one that refuses harmful requests while staying helpful, and one that refuses too much. HarmBench cannot tell them apart: no benign set, no false-refusal measurement, no capability check beyond the single MT-Bench score the authors ran on their own defended model. A model that refuses everything scores 0% ASR and tops the table. Claude's numbers do not by themselves mean "well-calibrated"; that is a claim this benchmark cannot make. See the classifier's failure modes.

The MT-Bench line is the only evidence that any defended model on the list is still useful. Without it, "average ASR 19.1" and "refuses everything" are indistinguishable from the grid.

Why system-level defenses are out of scope

Defenses split in two. Model-level defenses change the model: safety training, refusal mechanisms, system prompts, adversarial training. System-level defenses wrap it: input filters, output classifiers, prompt cleansing, perplexity checks, a screening model. HarmBench evaluates the first and excludes the second; §B.3 gives the reason:

"[W]hen assessing the robustness of a defense, it is vital to consider adaptive attacks, yet adaptive attacks for system-level defenses are highly specific to the individual defense. This makes it challenging to determine whether a defense is truly robust or simply hasn't been adequately tested (Athalye et al., 2018)."— HarmBench, §B.3

Athalye, Carlini and Wagner (2018) took nine defenses from one ICLR and broke seven; each had made gradients useless to the standard attack in a way that attack was not written to handle. The general form: a defense evaluated only against attacks that predate it is untested, its score unfalsifiable rather than good.

Apply that to a filter. An input filter rejecting high-perplexity text stops GCG dead, since a GCG suffix is gibberish by construction, and the battery would score it near-perfect. The adaptive attack is obvious once you name the defense: add a fluency penalty, or switch to AutoDAN, which already produces readable prompts. But that attack must be written for that filter, and an output classifier needs another. No fixed battery honestly evaluates an arbitrary system-level defense: the honest evaluation is defense-specific, so the generic battery measures its own ignorance, not the defense's strength.

Model-level defenses lack this property: the defense is in the weights, the battery's attacks already take the weights (or the API) as input, and there is no separate component to be naive about. So a fixed battery gives a meaningful if incomplete reading, and R2D2's row shows the incompleteness legibly, because the battery caught the generalisation gap and printed it. HarmBench supports evaluating system-level defenses; it just does not do so itself. Running the battery against your filter is a lower bound on how bad things are, not evidence they are good.

What to carry away

Always ask which attacks: one family is not a list. Adversarial training generalises less than you hope and further than you fear: R2D2 transferred cleanly across the whole gradient family, including attacks it never saw, and never crossed into a different search space. And publish the capability number: robustness is trivially achievable by refusing everything, and one line of Table 11 separates R2D2 from a model that scored well by being useless.

That closes the measuring arc, and the map. All 510 behaviors are listed in full: text, category, functional type, split, each linking to its harm page and its line in the source CSV.

Next: Sources, verified — every file, table and URL this map draws on, and the three places the repo and the paper disagree. Or return to the field map and start an arc you have not read; this page is the last of Measuring.