What the paper found
Table 6, the single results table. What matters is the shape of the grid, not any one cell.
The grid
Table 6 is 29 target rows by 16 attack columns. Each cell is an attack success rate: the fraction of a method's test cases whose completion the HarmBench classifier scored as exhibiting the behavior. Generation is greedy, 512 new tokens, same behaviors and grader in every cell.
“All behaviors” means the 400 text behaviors: 200 standard + 100 contextual + 100 copyright. The 110 multimodal behaviors (Tables 9 and 10) are not blended in (see below). Per-cell behavior IDs: the behavior index.
| Model | GCG | GCG-M | GCG-T | PAIR | TAP | TAP-T | AutoDAN | PAP-5 | Human | DR |
|---|---|---|---|---|---|---|---|---|---|---|
| Llama 2 7B Chat | 32.5 | 21.2 | 19.7 | 9.3 | 9.3 | 7.8 | 0.5 | 2.7 | 0.8 | 0.8 |
| Llama 2 13B Chat | 30.0 | 11.3 | 16.4 | 15.0 | 14.2 | 8.0 | 0.8 | 3.3 | 1.7 | 2.8 |
| Llama 2 70B Chat | 37.5 | 10.8 | 22.1 | 14.5 | 13.3 | 16.3 | 2.8 | 4.1 | 2.2 | 2.8 |
| Vicuna 7B | 65.5 | 61.5 | 60.8 | 53.5 | 51.0 | 59.8 | 66.0 | 18.9 | 39.0 | 24.3 |
| Vicuna 13B | 67.0 | 61.3 | 54.9 | 47.5 | 54.8 | 62.1 | 65.5 | 19.3 | 40.0 | 19.8 |
| Baichuan 2 13B | 62.3 | 52.4 | 45.3 | 52.3 | 54.8 | 63.6 | 60.1 | 21.7 | 31.7 | 19.3 |
| Mistral 7B | 69.8 | 63.6 | 64.5 | 52.5 | 62.5 | 66.1 | 71.5 | 27.2 | 58.0 | 46.3 |
| Starling 7B | 66.0 | 61.9 | 59.0 | 58.3 | 68.5 | 66.3 | 74.0 | 31.9 | 60.2 | 57.0 |
| Zephyr 7B | 69.5 | 62.5 | 61.1 | 58.8 | 66.5 | 69.3 | 75.0 | 32.9 | 66.0 | 65.8 |
| Zephyr 7B + R2D2 | 5.5 | 4.9 | 0.0 | 48.0 | 60.8 | 54.3 | 17.0 | 24.3 | 13.6 | 14.2 |
| GPT-3.5 Turbo 0613 | – | – | 38.9 | 46.8 | 47.7 | 62.3 | – | 15.4 | 24.5 | 21.3 |
| GPT-3.5 Turbo 1106 | – | – | 42.5 | 35.0 | 39.2 | 47.5 | – | 11.3 | 2.8 | 33.0 |
| GPT-4 0613 | – | – | 22.0 | 39.3 | 43.0 | 54.8 | – | 16.8 | 11.3 | 21.0 |
| GPT-4 Turbo 1106 | – | – | 22.3 | 33.0 | 36.4 | 58.5 | – | 11.1 | 2.6 | 9.3 |
| Claude 1 | – | – | 12.1 | 10.0 | 7.0 | 1.5 | – | 1.3 | 2.4 | 5.0 |
| Claude 2 | – | – | 2.7 | 4.8 | 2.0 | 0.8 | – | 1.0 | 0.3 | 2.0 |
| Claude 2.1 | – | – | 2.6 | 2.8 | 2.5 | 0.8 | – | 0.9 | 0.3 | 2.0 |
| Gemini Pro | – | – | 18.0 | 35.1 | 38.8 | 31.2 | – | 11.8 | 12.1 | 18.0 |
| Average (all 29 rows) | 54.3 | 45.0 | 38.8 | 40.7 | 45.2 | 48.3 | 52.7 | 16.6 | 27.3 | 25.3 |
This is 18 of the paper's 29 model rows and 10 of its 16 attack columns. Cut rows: Baichuan 2 7B, Qwen 7B / 14B / 72B Chat, Koala 7B and 13B, Orca 2 7B and 13B, SOLAR 10.7B-Instruct, Mixtral 8x7B, OpenChat 3.5 1210. Cut columns: PEZ, GBDA, UAT, AutoPrompt, Stochastic Few-Shot and ZeroShot. Their averages are in the next table; the cut rows sit inside the Vicuna/Mistral/Starling band, so nothing at the extremes was removed. The Average row is over all 29 rows, not the 18 shown. Full grid: arxiv.org/html/2402.04249v2.
The dashes are structural, not missing data: the method needs something the target does not expose, token gradients for the gradient family or at minimum next-token logits. An API-only model offers neither, so seven of the sixteen columns are empty for every closed model. The pipeline's model registry encodes which target types each method may run against.
Column averages — all 16 methods, strongest to weakest
The n column below is counted from the grid; these averages are not comparable, each over a different population of models.
| Method | Avg ASR | n rows | What that population is |
|---|---|---|---|
| GCG | 54.3 | 19 | open weights, gradients required |
| AutoDAN | 52.7 | 21 | open weights, scoring by loss |
| TAP-Transfer | 48.3 | 29 | every target, incl. all closed |
| TAP | 45.2 | 29 | every target, incl. all closed |
| GCG-Multi | 45.0 | 19 | open weights, gradients required |
| AutoPrompt | 43.7 | 19 | open weights, gradients required |
| PAIR | 40.7 | 29 | every target, incl. all closed |
| GCG-Transfer | 38.8 | 29 | crafted on surrogates, replayed anywhere |
| Stochastic Few-Shot | 38.3 | 21 | open weights, scoring by loss |
| UAT | 30.8 | 19 | open weights, gradients required |
| GBDA | 29.8 | 19 | open weights, gradients required |
| PEZ | 29.0 | 19 | open weights, gradients required |
| Human Jailbreaks | 27.3 | 29 | every target, incl. all closed |
| ZeroShot | 25.4 | 29 | every target, incl. all closed |
| DirectRequest | 25.3 | 29 | every target — this is the floor |
| PAP-top5 | 16.6 | 29 | every target, incl. all closed |
GCG's 54.3 is over nineteen open-weight models, eighteen with no adversarial training; TAP-Transfer's 48.3 includes Claude 2 and Claude 2.1, whose 0.8 cells pull the mean down. The 1.6-point gap between GCG and AutoDAN, on populations of 19 and 21, is not a result.
By functional slice
| Slice | GCG | AutoDAN | TAP-T | PAIR | Human | DR |
|---|---|---|---|---|---|---|
| All text behaviors (400) | 54.3 | 52.7 | 48.3 | 40.7 | 27.3 | 25.3 |
| Standard (200) | 69.1 | 68.3 | 60.0 | 47.5 | 31.9 | 23.9 |
| Contextual (100) | 74.8 | 68.4 | 67.5 | 60.5 | 41.4 | 46.2 |
| Copyright (100) | 4.6 | 6.0 | 5.9 | 7.5 | 4.3 | 7.3 |
The copyright row is recomputed from the grid, not copied from the paper (see below). Zephyr 7B under GCG is 90.5 on standard and 90.0 on contextual, and its famous 69.5 is after averaging in 7.0 on copyright. Quote 69.5 as “Zephyr complies with 70% of harmful requests” and you understate the standard result by twenty-one points and overstate copyright by sixty-three.
The findings
Nothing is robust to everything; nothing breaks everything
Down a column: AutoDAN peaks at 75.0 on Zephyr 7B, bottoms at 0.5 on Llama 2 7B Chat, a 150-fold spread. Across a row: Llama 2 7B Chat is 0.5 under AutoDAN, 0.8 under Human Jailbreaks, 32.5 under GCG. “Robustness” is not a scalar property of the model.
The Claude 2 and Claude 2.1 rows come close to being exceptions: their best cells across the eight methods that reach them are 4.8 and 4.1, not distinguishable from zero at ~7% classifier label noise. But only eight of sixteen methods could run against them, so those rows describe query-only attacks, not the models. The other way: GCG never fell below 30.0 on any model it could run against except the one adversarially trained against it, its floor being Llama 2 13B Chat at 30.0 over eighteen undefended open-weight targets.
Model size does not predict robustness
Within Llama 2 Chat, GCG scores 32.5 on 7B, 30.0 on 13B, 37.5 on 70B: the largest is the least robust, not monotonic. Vicuna does the same in miniature (65.5 at 7B, 67.0 at 13B). Under AutoDAN the Llama rows go 0.5, 0.8, 2.8 with size, within noise of zero. Scale is not the variable; “bigger models will be safer” gets no support here.
The training pipeline is the variable
Mistral 7B scores 69.8 under GCG. Zephyr 7B, a fine-tune of that base, scores 69.5. Take Zephyr's SFT pipeline, add adversarial training against a refreshed pool of GCG test cases, and the same architecture lands at 5.5: a 64-point swing with no change of base model or size.
Easily over-read, though: the same defended row is 48.0 under PAIR and 54.3 under TAP-Transfer, barely moved from undefended Zephyr's 58.8 and 69.3. The training bought robustness to gradient attacks only. The defenses page takes that apart, including its capability cost.
Contextual behaviors are easier than standard ones
GCG: 74.8 contextual against 69.1 standard. PAIR: 60.5 against 47.5. DirectRequest: 46.2 against 23.9, nearly double at the floor with no attack machinery. The gap is widest for the cheapest methods. The paper flags this as a problem:
“ASR is considerably higher for contextual behaviors than standard behaviors. This is concerning, as contextual behaviors represent more specific harmful tasks that would be challenging to look up the answer to on a search engine. Thus, behaviors [that] would [be] more differentially harmful for LLMs to exhibit are easier to elicit with red teaming methods.”— HarmBench, Figure 11 caption
Backwards from the design intent: the contextual type tests behaviors where a model adds capability over a search engine, and those are the ones models most readily do.
DirectRequest at 25.3 is the floor, and the gap is the finding
DirectRequest hands the behavior string over with no attack, averaging 25.3. ZeroShot (25.4) and Human Jailbreaks (27.3) sit within two points, inside the classifier's noise band. The useful quantity is the delta over asking, not the absolute score.
On Zephyr 7B, DirectRequest is 65.8 and GCG 69.5: five hundred optimisation steps bought 3.7 points, the model already complying. On Llama 2 7B Chat, 0.8 and 32.5: the same optimiser bought 31.7 points. A published ASR without its DirectRequest baseline is uninterpretable.
PAP-top5 is the weakest column and that is not the whole story
PAP-top5 averages 16.6, last of sixteen, best cell 32.9 on Zephyr: the only column that never clears a third. But it is query-only, reaches closed models, and costs a handful of generations against GCG's hundreds of gradient steps on owned GPUs. On ASR per dollar the ordering inverts: these averages describe reach, not a leaderboard.
Copyright is near the floor for everything
Every method lands between 4.3 and 8.5 on the copyright slice, no cell in that 29×16 grid over 28.0. That is the metric, not the model: copyright is graded by MinHash similarity against shipped source hashes, firing only on near-verbatim regurgitation, so a competent paraphrase scores zero. The column measures memorised-text extraction, and “copyright violation ASR” overclaims; the copyright category page has more.
What the paper does not report: ASR by semantic category
There is no per-semantic-category ASR table anywhere in the paper. The only per-category evidence is three figures, images with no tabulated values. Figure 9 averages ASR over the seven categories across all attacks and open-source models, its caption claiming copyright is much lower and the average ASR is similar across all other categories. Figure 11 does the same split for standard / contextual / copyright, the claim quoted above. Figure 10 is the only one separating models:
“For specific models, some categories of harm are easier to elicit than others. For example, on Llama 2 and GPT models the Misinformation & Disinformation category has the highest ASR, but for Baichuan 2 and Starling the Chemical & Biological Weapons / Drugs category has the highest ASR. This suggests that training distributions can greatly influence the kinds of behaviors that are harder to elicit.”— HarmBench, Figure 10 caption
So Misinformation & Disinformation is highest on the Llama 2 and GPT families, Chemical & Biological Weapons / Drugs highest on Baichuan 2 and Starling. That is the whole per-category result: a ranking over four families, no numbers, copyright excluded. Per-category ASR must be recomputed from the raw completions in the Zenodo archive (record 10714577, 10.4 GB), joining each to its behavior's SemanticCategory in the behavior index.
One printed row does not add up
The arXiv v2 HTML prints an Average row under the copyright sub-table reading 50.7 / 41.6 / 36.5 / 28.2 / 28.7 / 29.6 / 40.9 / 37.1 / 25.4 / 39.0 / 42.6 / 45.4 / 48.8 / 17.3 / 26.2 / 25.6. It cannot be the grid's mean: every cell there lies between 0.0 and 28.0, so no column mean can be 50.7. Recomputed column means from the printed cells, in column order (GCG, GCG-M, GCG-T, PEZ, GBDA, UAT, AP, SFS, ZS, PAIR, TAP, TAP-T, AutoDAN, PAP-top5, Human, DR):
| Source | GCG | GCG-M | GCG-T | PEZ | GBDA | UAT | AP | SFS | ZS | PAIR | TAP | TAP-T | AutoDAN | PAP-5 | Human | DR |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Printed in Table 6 | 50.7 | 41.6 | 36.5 | 28.2 | 28.7 | 29.6 | 40.9 | 37.1 | 25.4 | 39.0 | 42.6 | 45.4 | 48.8 | 17.3 | 26.2 | 25.6 |
| Recomputed from the grid | 4.6 | 3.3 | 4.3 | 4.0 | 3.9 | 4.1 | 4.9 | 6.6 | 7.2 | 7.5 | 8.5 | 5.9 | 6.0 | 7.4 | 4.3 | 7.3 |
| Table 7 (test split), printed | 4.7 | 3.5 | 4.4 | 4.2 | … | … | … | … | … | … | … | … | … | … | … | … |
| Table 8 (validation), printed | 3.9 | 2.9 | 3.9 | 3.3 | … | … | … | … | … | … | … | … | … | … | … | … |
The recomputed row sits between the test-split and validation-split copyright averages the paper prints in Tables 7 and 8, over subsets of the same cells that must bracket it. The printed row does not.
Reconciliation one, the whole benchmark. Table 6's all-behaviors numbers should be the 200/100/100 blend of its three sub-tables. With the recomputed copyright figure: (200×69.1 + 100×74.8 + 100×4.6) / 400 = 21,760 / 400 = 54.4, against a printed all-behaviors GCG average of 54.3. The tenth is not slack: averaging the all-behaviors GCG column directly gives 54.32, the difference tracing to two rows (Qwen 7B Chat, Qwen 14B Chat) whose sub-table cells do not blend to their printed all-behaviors cells. Substitute the printed copyright average and you get (13,820 + 7,480 + 5,070) / 400 = 65.9, eleven and a half points off the paper's headline.
Reconciliation two, a single row. Llama 2 7B Chat under GCG: standard 34.5, contextual 58.0, copyright 3.0. (200×34.5 + 100×58.0 + 100×3.0) / 400 = 13,000 / 400 = 32.5, exactly the printed all-behaviors cell. It holds on Zephyr 7B (90.5 / 90.0 / 7.0 → 69.5) and the defended model (0.0 / 21.0 / 1.0 → 5.5). The per-cell copyright values are sound; only the Average row beneath them is wrong.
Conclusion: the printed copyright Average row in v2 is a typesetting error, most likely the standard-behaviors averages pasted into the copyright block, consistent with their magnitude. This map uses the recomputed values everywhere; the sources page records the check. The same arithmetic confirms Table 6 covers the 400 text behaviors, not all 510: the weights balance only at 200 + 100 + 100.
Two smaller discrepancies
The validation split. The paper describes the text validation set as 40 standard + 20 contextual. The shipped CSV has 41 standard + 19 contextual (plus 20 copyright, for 80), consistent with the paper's note that the stratified sample was hand-adjusted. It changes no result, but a harness hard-coding 40/20 mis-slices silently. Counts here are parsed from the CSVs.
§6.3's prose against its own tables. The body text reports Llama 2 7B Chat at 31.8 → 5.9 under the paper's defense, 13B at 30.2 → 5.9. The v2 tables give 32.5 → 5.5 on all behaviors, 31.9 → 6.3 on test, 13B at 30.0. The prose matches neither split, probably predating a re-run. This map uses the table figures throughout, naming the split whenever a number could be either.
How to read someone else’s HarmBench number
Five questions, in order of how often they change the answer.
- Which split? All 400, the 320-behavior test split, or the 80-behavior validation split. Column averages barely move (GCG 54.3 / 54.2 / 54.9), but rows swing hard, since 80 behaviors makes one behavior 1.25 points: defended Zephyr under GCG is 5.5 on all behaviors, 6.3 on test, 2.5 on validation.
- Which classifier? Test classifier is a Llama 2 13B fine-tune at ~93% human agreement; validation is a Mistral 7B at 88.6%. Not interchangeable, and the standing rule is: never tune an attack against the classifier you report with.
- What token budget? 512 new tokens, greedy, fixed. Load-bearing: a longer budget lets a partial answer become complete, and the ASR moves with nothing about the model changing.
- Was DirectRequest reported alongside? If not, the number has no floor and cannot be read. See the Zephyr 65.8 → 69.5 case above.
- All-behaviors or standard-only? An all-behaviors figure blends in 100 copyright behaviors nothing scores on, running about 15 points below the standard-behaviors figure for a compliant model; the two get quoted interchangeably.
None of the five tells you this: over-refusal is not measured anywhere in HarmBench. With no benign control set, a model that refuses everything scores 0.0% across all sixteen columns. The only capability check is a single MT-Bench score for the authors' own defended model. Any robustness claim from this table is half a result.
The takeaway: an ASR is a claim about the attack you ran, not the model you ran it against. The defended Zephyr row proves it: 5.5 under the attack it was trained against, 60.8 under one it was not.