Defenses: R2D2 and the read-across caveat
The prior three pages built the number: the grader, the pipeline, the grid. This page moves it down. HarmBench's authors built a defense with their own benchmark in the loop and printed its full row: strong in one place, weak in another, same line of the same table.
R2D2: adversarial training against an adversary that keeps moving
Most defended models are trained once: assemble harmful prompts, fine-tune to refuse, ship. But the attacker optimises after you stop, and GCG searches token by token against your specific weights. Putting the attacker inside the training loop is the textbook fix, but unaffordable for LLMs: GCG needs about twenty minutes per test case against a 7B model on an A100. Earlier attempts froze the adversary and got disappointing results.
Robust Refusal Dynamic Defense (R2D2) borrows vision's fix, persistent adversarial examples: a standing pool of 180 test cases. Each iteration samples 8, runs GCG 5 further steps to refresh them against the current weights, then takes a model step, so the attack is never solved from scratch nor fully stale. 20% resets every 50 updates. Total: 500 steps, roughly 16 hours on an 8×A100 node.
Base model Mistral 7B, instruction data UltraChat, harness the Zephyr 7B Beta SFT training code with R2D2 spliced into the trainer.
The two losses, and why one of them is not obvious
Each step minimises three terms. The SFT loss is ordinary UltraChat instruction tuning, present only to keep the model useful. The toward loss, −log fθ(trefusal | xi), makes a refusal likely; its refusal target is one hard-coded sentence for every test case. The third term is the interesting one:
L_away = -1 · log( 1 - f_θ(t_i | x_i) )
where ti is the harmful continuation GCG optimised toward and xi is the adversarial prompt: the exact negation of the GCG objective, pushing back down whatever GCG just made probable.
It is not redundant with the toward loss: "make refusal likely" and "make this specific completion unlikely" are different objectives, and a model can satisfy the first while failing the second. Raising a refusal string from 0.2 to 0.6 says almost nothing about a harmful continuation sitting at 0.05, and 0.05 is not safe: greedy decoding only needs the harmful branch to win token by token. The toward loss shapes the mode; the away loss attacks the specific tail. Train only the first and the model refuses confidently until a suffix flips it.
The implementation adds a detail the formula omits: the away term is floored per token, so once a token's log P falls below −5.0 (roughly 0.7% probability) its contribution is zeroed. Otherwise −log(1−p) keeps returning gradient on already-impossible tokens, and the unbounded away term swamps the SFT term and eats general ability.
The row, and the ratio column
Zephyr 7B against Zephyr 7B + R2D2 across every attack column of Table 6’s all-behaviors slice (the 400 text behaviors), sorted by how much R2D2 helped, because that ordering is the finding.
| Attack | Family | Zephyr 7B | + R2D2 | × lower |
|---|---|---|---|---|
| UAT | gradient | 62.3 | 0.0 | ∞ |
| GCG-Transfer | gradient | 61.1 | 0.0 | ∞ |
| GBDA | gradient | 62.8 | 0.2 | 314 |
| PEZ | gradient | 62.5 | 2.9 | 21.6 |
| GCG-Multi | gradient | 62.5 | 4.9 | 12.8 |
| GCG | gradient — the train-time adversary | 69.5 | 5.5 | 12.6 |
| AutoPrompt | gradient | 60.5 | 5.5 | 11.0 |
| ZeroShot | static, LLM-written | 60.0 | 7.2 | 8.3 |
| Human Jailbreaks | fixed templates | 66.0 | 13.6 | 4.9 |
| DirectRequest | no attack at all | 65.8 | 14.2 | 4.6 |
| AutoDAN | LLM-driven, genetic | 75.0 | 17.0 | 4.4 |
| Stochastic Few-Shot | LLM-driven, adaptive | 62.0 | 43.5 | 1.4 |
| PAP-top5 | persuasion taxonomy | 32.9 | 24.3 | 1.4 |
| TAP-Transfer | LLM-driven, tree search | 69.3 | 54.3 | 1.3 |
| PAIR | LLM-driven, iterative | 58.8 | 48.0 | 1.2 |
| TAP | LLM-driven, tree search | 66.5 | 60.8 | 1.1 |
| Mean of the 16 columns | 62.3 | 18.9 | 3.3 |
∞ means the column reached exactly 0.0. The × column is arithmetic on the two printed columns, not reported by the paper; the last row is the plain mean of the sixteen cells. Table 11 gives 62.7 and 19.1 for the same models, a slightly different "Average ASR" slice: the same mismatch that makes Zephyr's GCG ASR 69.4 there and 69.5 here.
5.5 under GCG is the best any model-level defense in the paper achieves against its strongest attack, beating Llama 2 13B Chat's 30.0 and 70B Chat's 37.5, both products of full industrial RLHF. The denominator is the 400 text behaviors (200 standard, 100 contextual, 100 copyright); the 110 multimodal behaviors sit in Table 9. It includes the 100 copyright behaviors, MinHash-graded near the floor for every model and attack (see the copyright category), which pull both rows down about equally, so the comparison is fair but not a uniform compliance rate.
What it cost: Table 11
| Model | MT-Bench | GCG ASR | Average ASR |
|---|---|---|---|
| Zephyr 7B (SFT + DPO) | 7.34 | 69.4 | 62.7 |
| Mistral 7B Instruct v0.2 | 6.5 | 69.1 | 55.7 |
| Koala 13B | 5.4 | 62.2 | 48.5 |
| Zephyr 7B + R2D2 | 6.0 | 5.5 | 19.1 |
R2D2 lands at 6.0, 1.34 MT-Bench points below its own Zephyr base, below Mistral 7B's 6.5, above Koala 13B's 5.4. The paper calls this "does not necessarily harm performance"; it is a 1.34-point drop on a ten-point scale. The robustness gain is enormous and the capability cost is not zero: 43.6 points of average ASR traded for 1.34 of MT-Bench.
The paper states an asymmetry plainly. The main-table Zephyr 7B is Zephyr 7B Beta, SFT and DPO; R2D2 is built on the SFT-only variant, with no DPO stage. So the 7.34 → 6.0 gap is not purely the price of adversarial training, part is the missing preference-optimisation stage, and the same asymmetry sits under the 69.5 → 5.5 robustness claim. This is a preliminary demonstration, not a controlled ablation, and the paper says so.
Read the row across, not down
Ignore the absolute numbers and read the table by group. The top (UAT, GCG-Transfer, GBDA, PEZ, GCG-Multi, GCG, AutoPrompt) is the entire gradient family, methods that search token sequences using the model's own gradients; all seven collapse, five below 5%. The bottom (TAP, PAIR, TAP-Transfer, PAP-top5, Stochastic Few-Shot) is almost exactly the LLM-driven family, where another model refines a natural-language jailbreak from what the target did last. Those barely move: 12.6× in one column, 1.1× in another, same model and training run.
The split is not simply gradient versus not. DirectRequest, Human Jailbreaks and ZeroShot also fall far (4.6×, 4.9×, 8.3×) with no gradients. What they share is being non-adaptive: a fixed request, template, or blind prompt, which plain refusal training (the toward loss) handles for free. What survives is the intersection: attacks that adapt to this model in a space R2D2 never trained against. GCG adapts in token space, where the defense lives; PAIR and TAP adapt in semantic space, where it has nothing to say.
An ASR number is a claim about the attack you ran, not about the model. "5.5% ASR" and "60.8% ASR" are both true of the same weights, the same 510 behaviors, the same classifier, differing only in which attack produced the test cases. A robustness figure without its attacks named is a coincidence about which attacks the authors ran, which correlates with what they defended against. That is the practical content of the finding that no model is robust to all attacks and no attack breaks all models: a robustness claim needs an attack list or it means nothing.
The paper says this about its own method, in the section announcing its success:
"For some attacks, the improvement conferred by R2D2 is less pronounced. This is especially true for methods dissimilar to the train-time GCG adversary, including PAIR, TAP, and Stochastic Few-Shot. This suggests that incorporating multiple diverse attacks into adversarial training may be necessary to obtain generalizable robustness."— HarmBench, §6, Adversarial Training Results
A normal defense paper tables the attacks it handles and explains why the rest were out of scope. Here the weak columns print at the same width as the strong ones, possible only because the authors fixed the battery before building the defense.
One caution on the deltas: the grader agrees with humans about 93% of the time, so roughly 7% label noise sits under every cell. Far too little to threaten 69.5 → 5.5, but enough that 66.5 versus 60.8 is a difference you should not lean on. The read-across holds because the extremes differ by an order of magnitude.
The other "defenses" among the 33 targets
"33 target LLMs and defenses" invites a misreading. R2D2 is the only defense method HarmBench implements. Everything else is a shipped model its vendor safety-trained: a real category, since safety training is a model-level defense, but with no defense code to reproduce. Three groups from Table 6:
- The Llama 2 Chat family, what good refusal training looks like. 32.5 / 30.0 / 37.5 under GCG for 7B / 13B / 70B, against 0.8 / 2.8 / 2.8 under DirectRequest: the RLHF signature, where merely asking gets nothing but a white-box optimizer still gets through about a third of the time. 70B is less robust than 13B: scale does not buy robustness, the pipeline does.
- The Claude models, the lowest numbers in the table. Claude 2 and 2.1 at 2.7 and 2.6 under GCG-Transfer, under 5 on every column run, 0.3 under Human Jailbreaks; Claude 1 is higher at 12.1. API-only, so no gradient columns at all.
- The GPT models, safety-trained but penetrable by transfer. GPT-4 0613 and GPT-4 Turbo 1106 hold DirectRequest to 21.0 and 9.3 and GCG-Transfer to 22.0 and 22.3, yet both exceed 54 under TAP-Transfer, and GPT-3.5 Turbo 0613 reaches 62.3 there. Transfer attacks are what these models are worst against.
The MT-Bench line is the only evidence that any defended model on the list is still useful. Without it, "average ASR 19.1" and "refuses everything" are indistinguishable from the grid.
Why system-level defenses are out of scope
Defenses split in two. Model-level defenses change the model: safety training, refusal mechanisms, system prompts, adversarial training. System-level defenses wrap it: input filters, output classifiers, prompt cleansing, perplexity checks, a screening model. HarmBench evaluates the first and excludes the second; §B.3 gives the reason:
"[W]hen assessing the robustness of a defense, it is vital to consider adaptive attacks, yet adaptive attacks for system-level defenses are highly specific to the individual defense. This makes it challenging to determine whether a defense is truly robust or simply hasn't been adequately tested (Athalye et al., 2018)."— HarmBench, §B.3
Athalye, Carlini and Wagner (2018) took nine defenses from one ICLR and broke seven; each had made gradients useless to the standard attack in a way that attack was not written to handle. The general form: a defense evaluated only against attacks that predate it is untested, its score unfalsifiable rather than good.
Apply that to a filter. An input filter rejecting high-perplexity text stops GCG dead, since a GCG suffix is gibberish by construction, and the battery would score it near-perfect. The adaptive attack is obvious once you name the defense: add a fluency penalty, or switch to AutoDAN, which already produces readable prompts. But that attack must be written for that filter, and an output classifier needs another. No fixed battery honestly evaluates an arbitrary system-level defense: the honest evaluation is defense-specific, so the generic battery measures its own ignorance, not the defense's strength.
Model-level defenses lack this property: the defense is in the weights, the battery's attacks already take the weights (or the API) as input, and there is no separate component to be naive about. So a fixed battery gives a meaningful if incomplete reading, and R2D2's row shows the incompleteness legibly, because the battery caught the generalisation gap and printed it. HarmBench supports evaluating system-level defenses; it just does not do so itself. Running the battery against your filter is a lower bound on how bad things are, not evidence they are good.
What to carry away
Always ask which attacks: one family is not a list. Adversarial training generalises less than you hope and further than you fear: R2D2 transferred cleanly across the whole gradient family, including attacks it never saw, and never crossed into a different search space. And publish the capability number: robustness is trivially achievable by refusing everything, and one line of Table 11 separates R2D2 from a model that scored well by being useless.
That closes the measuring arc, and the map. All 510 behaviors are listed in full: text, category, functional type, split, each linking to its harm page and its line in the source CSV.