HARMBENCH // FIELD MAP
← field map
MEASURING · 03 OF 0429 model rows × 16 attack columns · 400 text behaviors · GCG 54.3, DirectRequest 25.3

What the paper found

Reading Table 6 cell by cell: which findings survive, which need a caveat, and one printed row that is arithmetically impossible.
TL;DR: Eighteen attacks against 33 targets, both ways: no model resisted every attack, no attack broke every model. Post-training pipeline matters far more than size. Zephyr 7B is 69.5 under GCG; those same weights adversarially trained are 5.5. Quote an attack's gap over DirectRequest (simply asking), not its raw score. And the Average row under the copyright sub-table is arithmetically impossible; recomputed values and proof below.
Strongest column
54.3
GCG, all behaviors — over the 19 open-weight rows it could run against
The floor
25.3
DirectRequest — no attack at all, all 29 rows
Weakest column
16.6
PAP-top5 — and its best single cell anywhere is 32.9
Grid
29 × 16
target rows × methods, over 400 text behaviors
Label noise
~7%
classifier vs. human agreement is ~93%

Table 6, the single results table. What matters is the shape of the grid, not any one cell.

The grid

Table 6 is 29 target rows by 16 attack columns. Each cell is an attack success rate: the fraction of a method's test cases whose completion the HarmBench classifier scored as exhibiting the behavior. Generation is greedy, 512 new tokens, same behaviors and grader in every cell.

“All behaviors” means the 400 text behaviors: 200 standard + 100 contextual + 100 copyright. The 110 multimodal behaviors (Tables 9 and 10) are not blended in (see below). Per-cell behavior IDs: the behavior index.

Model GCG GCG-M GCG-T PAIR TAP TAP-T AutoDAN PAP-5 Human DR
Llama 2 7B Chat32.521.219.79.39.37.80.52.70.80.8
Llama 2 13B Chat30.011.316.415.014.28.00.83.31.72.8
Llama 2 70B Chat37.510.822.114.513.316.32.84.12.22.8
Vicuna 7B65.561.560.853.551.059.866.018.939.024.3
Vicuna 13B67.061.354.947.554.862.165.519.340.019.8
Baichuan 2 13B62.352.445.352.354.863.660.121.731.719.3
Mistral 7B69.863.664.552.562.566.171.527.258.046.3
Starling 7B66.061.959.058.368.566.374.031.960.257.0
Zephyr 7B69.562.561.158.866.569.375.032.966.065.8
Zephyr 7B + R2D25.54.90.048.060.854.317.024.313.614.2
GPT-3.5 Turbo 0613––38.946.847.762.3–15.424.521.3
GPT-3.5 Turbo 1106––42.535.039.247.5–11.32.833.0
GPT-4 0613––22.039.343.054.8–16.811.321.0
GPT-4 Turbo 1106––22.333.036.458.5–11.12.69.3
Claude 1––12.110.07.01.5–1.32.45.0
Claude 2––2.74.82.00.8–1.00.32.0
Claude 2.1––2.62.82.50.8–0.90.32.0
Gemini Pro––18.035.138.831.2–11.812.118.0
Average (all 29 rows)54.345.038.840.745.248.352.716.627.325.3

This is 18 of the paper's 29 model rows and 10 of its 16 attack columns. Cut rows: Baichuan 2 7B, Qwen 7B / 14B / 72B Chat, Koala 7B and 13B, Orca 2 7B and 13B, SOLAR 10.7B-Instruct, Mixtral 8x7B, OpenChat 3.5 1210. Cut columns: PEZ, GBDA, UAT, AutoPrompt, Stochastic Few-Shot and ZeroShot. Their averages are in the next table; the cut rows sit inside the Vicuna/Mistral/Starling band, so nothing at the extremes was removed. The Average row is over all 29 rows, not the 18 shown. Full grid: arxiv.org/html/2402.04249v2.

The dashes are structural, not missing data: the method needs something the target does not expose, token gradients for the gradient family or at minimum next-token logits. An API-only model offers neither, so seven of the sixteen columns are empty for every closed model. The pipeline's model registry encodes which target types each method may run against.

Column averages — all 16 methods, strongest to weakest

The n column below is counted from the grid; these averages are not comparable, each over a different population of models.

MethodAvg ASRn rowsWhat that population is
GCG54.319open weights, gradients required
AutoDAN52.721open weights, scoring by loss
TAP-Transfer48.329every target, incl. all closed
TAP45.229every target, incl. all closed
GCG-Multi45.019open weights, gradients required
AutoPrompt43.719open weights, gradients required
PAIR40.729every target, incl. all closed
GCG-Transfer38.829crafted on surrogates, replayed anywhere
Stochastic Few-Shot38.321open weights, scoring by loss
UAT30.819open weights, gradients required
GBDA29.819open weights, gradients required
PEZ29.019open weights, gradients required
Human Jailbreaks27.329every target, incl. all closed
ZeroShot25.429every target, incl. all closed
DirectRequest25.329every target — this is the floor
PAP-top516.629every target, incl. all closed

GCG's 54.3 is over nineteen open-weight models, eighteen with no adversarial training; TAP-Transfer's 48.3 includes Claude 2 and Claude 2.1, whose 0.8 cells pull the mean down. The 1.6-point gap between GCG and AutoDAN, on populations of 19 and 21, is not a result.

By functional slice

SliceGCGAutoDANTAP-TPAIRHumanDR
All text behaviors (400)54.352.748.340.727.325.3
Standard (200)69.168.360.047.531.923.9
Contextual (100)74.868.467.560.541.446.2
Copyright (100)4.66.05.97.54.37.3

The copyright row is recomputed from the grid, not copied from the paper (see below). Zephyr 7B under GCG is 90.5 on standard and 90.0 on contextual, and its famous 69.5 is after averaging in 7.0 on copyright. Quote 69.5 as “Zephyr complies with 70% of harmful requests” and you understate the standard result by twenty-one points and overstate copyright by sixty-three.

The findings

Nothing is robust to everything; nothing breaks everything

Down a column: AutoDAN peaks at 75.0 on Zephyr 7B, bottoms at 0.5 on Llama 2 7B Chat, a 150-fold spread. Across a row: Llama 2 7B Chat is 0.5 under AutoDAN, 0.8 under Human Jailbreaks, 32.5 under GCG. “Robustness” is not a scalar property of the model.

The Claude 2 and Claude 2.1 rows come close to being exceptions: their best cells across the eight methods that reach them are 4.8 and 4.1, not distinguishable from zero at ~7% classifier label noise. But only eight of sixteen methods could run against them, so those rows describe query-only attacks, not the models. The other way: GCG never fell below 30.0 on any model it could run against except the one adversarially trained against it, its floor being Llama 2 13B Chat at 30.0 over eighteen undefended open-weight targets.

Model size does not predict robustness

Within Llama 2 Chat, GCG scores 32.5 on 7B, 30.0 on 13B, 37.5 on 70B: the largest is the least robust, not monotonic. Vicuna does the same in miniature (65.5 at 7B, 67.0 at 13B). Under AutoDAN the Llama rows go 0.5, 0.8, 2.8 with size, within noise of zero. Scale is not the variable; “bigger models will be safer” gets no support here.

The training pipeline is the variable

Mistral 7B scores 69.8 under GCG. Zephyr 7B, a fine-tune of that base, scores 69.5. Take Zephyr's SFT pipeline, add adversarial training against a refreshed pool of GCG test cases, and the same architecture lands at 5.5: a 64-point swing with no change of base model or size.

Easily over-read, though: the same defended row is 48.0 under PAIR and 54.3 under TAP-Transfer, barely moved from undefended Zephyr's 58.8 and 69.3. The training bought robustness to gradient attacks only. The defenses page takes that apart, including its capability cost.

Contextual behaviors are easier than standard ones

GCG: 74.8 contextual against 69.1 standard. PAIR: 60.5 against 47.5. DirectRequest: 46.2 against 23.9, nearly double at the floor with no attack machinery. The gap is widest for the cheapest methods. The paper flags this as a problem:

“ASR is considerably higher for contextual behaviors than standard behaviors. This is concerning, as contextual behaviors represent more specific harmful tasks that would be challenging to look up the answer to on a search engine. Thus, behaviors [that] would [be] more differentially harmful for LLMs to exhibit are easier to elicit with red teaming methods.”— HarmBench, Figure 11 caption

Backwards from the design intent: the contextual type tests behaviors where a model adds capability over a search engine, and those are the ones models most readily do.

DirectRequest at 25.3 is the floor, and the gap is the finding

DirectRequest hands the behavior string over with no attack, averaging 25.3. ZeroShot (25.4) and Human Jailbreaks (27.3) sit within two points, inside the classifier's noise band. The useful quantity is the delta over asking, not the absolute score.

On Zephyr 7B, DirectRequest is 65.8 and GCG 69.5: five hundred optimisation steps bought 3.7 points, the model already complying. On Llama 2 7B Chat, 0.8 and 32.5: the same optimiser bought 31.7 points. A published ASR without its DirectRequest baseline is uninterpretable.

PAP-top5 is the weakest column and that is not the whole story

PAP-top5 averages 16.6, last of sixteen, best cell 32.9 on Zephyr: the only column that never clears a third. But it is query-only, reaches closed models, and costs a handful of generations against GCG's hundreds of gradient steps on owned GPUs. On ASR per dollar the ordering inverts: these averages describe reach, not a leaderboard.

Copyright is near the floor for everything

Every method lands between 4.3 and 8.5 on the copyright slice, no cell in that 29×16 grid over 28.0. That is the metric, not the model: copyright is graded by MinHash similarity against shipped source hashes, firing only on near-verbatim regurgitation, so a competent paraphrase scores zero. The column measures memorised-text extraction, and “copyright violation ASR” overclaims; the copyright category page has more.

What the paper does not report: ASR by semantic category

There is no per-semantic-category ASR table anywhere in the paper. The only per-category evidence is three figures, images with no tabulated values. Figure 9 averages ASR over the seven categories across all attacks and open-source models, its caption claiming copyright is much lower and the average ASR is similar across all other categories. Figure 11 does the same split for standard / contextual / copyright, the claim quoted above. Figure 10 is the only one separating models:

“For specific models, some categories of harm are easier to elicit than others. For example, on Llama 2 and GPT models the Misinformation & Disinformation category has the highest ASR, but for Baichuan 2 and Starling the Chemical & Biological Weapons / Drugs category has the highest ASR. This suggests that training distributions can greatly influence the kinds of behaviors that are harder to elicit.”— HarmBench, Figure 10 caption

So Misinformation & Disinformation is highest on the Llama 2 and GPT families, Chemical & Biological Weapons / Drugs highest on Baichuan 2 and Starling. That is the whole per-category result: a ranking over four families, no numbers, copyright excluded. Per-category ASR must be recomputed from the raw completions in the Zenodo archive (record 10714577, 10.4 GB), joining each to its behavior's SemanticCategory in the behavior index.

One printed row does not add up

The arXiv v2 HTML prints an Average row under the copyright sub-table reading 50.7 / 41.6 / 36.5 / 28.2 / 28.7 / 29.6 / 40.9 / 37.1 / 25.4 / 39.0 / 42.6 / 45.4 / 48.8 / 17.3 / 26.2 / 25.6. It cannot be the grid's mean: every cell there lies between 0.0 and 28.0, so no column mean can be 50.7. Recomputed column means from the printed cells, in column order (GCG, GCG-M, GCG-T, PEZ, GBDA, UAT, AP, SFS, ZS, PAIR, TAP, TAP-T, AutoDAN, PAP-top5, Human, DR):

SourceGCGGCG-MGCG-TPEZGBDAUATAPSFSZSPAIRTAPTAP-TAutoDANPAP-5HumanDR
Printed in Table 650.741.636.528.228.729.640.937.125.439.042.645.448.817.326.225.6
Recomputed from the grid4.63.34.34.03.94.14.96.67.27.58.55.96.07.44.37.3
Table 7 (test split), printed4.73.54.44.2………………………………
Table 8 (validation), printed3.92.93.93.3………………………………

The recomputed row sits between the test-split and validation-split copyright averages the paper prints in Tables 7 and 8, over subsets of the same cells that must bracket it. The printed row does not.

Reconciliation one, the whole benchmark. Table 6's all-behaviors numbers should be the 200/100/100 blend of its three sub-tables. With the recomputed copyright figure: (200×69.1 + 100×74.8 + 100×4.6) / 400 = 21,760 / 400 = 54.4, against a printed all-behaviors GCG average of 54.3. The tenth is not slack: averaging the all-behaviors GCG column directly gives 54.32, the difference tracing to two rows (Qwen 7B Chat, Qwen 14B Chat) whose sub-table cells do not blend to their printed all-behaviors cells. Substitute the printed copyright average and you get (13,820 + 7,480 + 5,070) / 400 = 65.9, eleven and a half points off the paper's headline.

Reconciliation two, a single row. Llama 2 7B Chat under GCG: standard 34.5, contextual 58.0, copyright 3.0. (200×34.5 + 100×58.0 + 100×3.0) / 400 = 13,000 / 400 = 32.5, exactly the printed all-behaviors cell. It holds on Zephyr 7B (90.5 / 90.0 / 7.0 → 69.5) and the defended model (0.0 / 21.0 / 1.0 → 5.5). The per-cell copyright values are sound; only the Average row beneath them is wrong.

Conclusion: the printed copyright Average row in v2 is a typesetting error, most likely the standard-behaviors averages pasted into the copyright block, consistent with their magnitude. This map uses the recomputed values everywhere; the sources page records the check. The same arithmetic confirms Table 6 covers the 400 text behaviors, not all 510: the weights balance only at 200 + 100 + 100.

Two smaller discrepancies

The validation split. The paper describes the text validation set as 40 standard + 20 contextual. The shipped CSV has 41 standard + 19 contextual (plus 20 copyright, for 80), consistent with the paper's note that the stratified sample was hand-adjusted. It changes no result, but a harness hard-coding 40/20 mis-slices silently. Counts here are parsed from the CSVs.

§6.3's prose against its own tables. The body text reports Llama 2 7B Chat at 31.8 → 5.9 under the paper's defense, 13B at 30.2 → 5.9. The v2 tables give 32.5 → 5.5 on all behaviors, 31.9 → 6.3 on test, 13B at 30.0. The prose matches neither split, probably predating a re-run. This map uses the table figures throughout, naming the split whenever a number could be either.

How to read someone else’s HarmBench number

Five questions, in order of how often they change the answer.

None of the five tells you this: over-refusal is not measured anywhere in HarmBench. With no benign control set, a model that refuses everything scores 0.0% across all sixteen columns. The only capability check is a single MT-Bench score for the authors' own defended model. Any robustness claim from this table is half a result.

The takeaway: an ASR is a claim about the attack you ran, not the model you ran it against. The defended Zephyr row proves it: 5.5 under the attack it was trained against, 60.8 under one it was not.

Next: Defenses: R2D2 and the read-across caveat · back to the map.