HARMBENCH // FIELD MAP
← field map
PROVENANCE1 paper (v2) · 1 repo @ main · 33 URLs checked, 32 live · 4 discrepancies

Sources, verified

What each source establishes and does not, and the four repo/paper disagreements.
TL;DR: Three bodies of evidence: arXiv:2402.04249v2 for the design argument and every reported number, the repository at main for every count and mechanism, and artifacts shipping separately (three classifiers on Hugging Face, a 10.4 GB results archive on Zenodo). Counts were parsed from the shipped CSVs, not copied from the paper; results were read from the v2 HTML tables, never the figures. Four repo/paper mismatches are documented below, one a printed table row that cannot be reconciled with its own cells.
Paper
v2
arXiv:2402.04249, 2024-02-27
Repo
main
cloned 2026-09-08
URLs checked
33
32 × HTTP 200
Discrepancies
4
3 material, 1 cosmetic

Organised by artifact rather than by claim: for each source, what it establishes and what people routinely assume it establishes but does not.

The paper

Mazeika, Phan, Yin, Zou, Wang, Mu, Sakhaee, Li, Basart, Li, Forsyth and Hendrycks, HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. arXiv:2402.04249: v1 2024-02-06, v2 2024-02-27. This map reads the v2 HTML and follows v2 where the versions differ.

The paper is two artifacts: a benchmark-design argument (sections 3 and 4, appendix B) and an empirical comparison (section 6, appendix D). Which piece supports what:

Section or tableWhat it establishes, and where this map uses it
§3.1The ASR definition — mean classifier label over a method’s test cases for a behavior, with greedy decoding and a fixed 512-new-token budget. The single most load-bearing definition on the site; read on the pipeline page and assumed by every number on the results page.
§4.2Curation of harmful behaviors: the design principles — breadth, differential harm (prefer behaviors a search engine does not already answer), dual-intent handling, and comparability. This is why the benchmark looks the way it does.
§5R2D2: the adversarial-training method itself — the refreshed pool of GCG test cases, the “toward” and “away” losses, and the training loop folded into Zephyr’s SFT code. See defenses.
§6.3R2D2’s results in prose, including the comparison against the Llama 2 Chat family. Its quoted figures disagree slightly with the v2 tables — see discrepancies.
§B.3Supported threat models: no-access (transfer), query-access and parameter-access attacks; model-level versus system-level defenses; and the explicit decision to scope the large-scale comparison to model-level defenses only.
§B.4The seven semantic categories and their sub-bullets. The taxonomy quoted on each harm page comes from here, not from the CSV, which carries only the machine slug.
§B.5.1How the Llama 2 13B evaluation classifier was trained — template-generated completions, GPT-4-0613 labels, distillation, and the separate Mistral validation classifier fine-tuned on half the data. See the grader.
§B.5.2Why copyright is graded by hashing rather than by an LLM judge: works inspired by a copyrighted original are not reliably separable from attempts to reproduce it, so the benchmark measures only near-verbatim regurgitation. Underwrites the copyright type page.
Table 3Classifier agreement with human labels on 600 hand-labelled examples, against AdvBench, GPTFuzz, ChatGLM, Llama Guard and GPT-4. The ~93% figure — and therefore the ~7% label-noise caveat repeated across this site — comes from here. Read.
Table 4The three prequalification sets, which is where you learn how the prior graders fail rather than just that they do: refusal-prefix heuristics call benign text harmful, and “is this harmful?” graders cannot tell a completion of the wrong behavior from a completion of the right one.
Table 5Dataset comparison against nine prior behavior sets. Establishes 510 unique behaviors and that HarmBench is the only one of the ten carrying both multimodal and contextual behaviors.
Table 6ASR on the all-behaviors slice — the 400 text behaviors, since the sub-table weights are 200 + 100 + 100 and the 110 multimodal behaviors are Table 9 — with sub-tables for standard, contextual and copyright. The main grid of the results page, and the source of every attack-column average quoted on the attack pages.
Table 7The same grid on the 320-behavior text test split. Used here as an independent check on Table 6, and it is what catches the copyright Average row.
Table 8The same grid on the 80-behavior validation split. Second check, same role.
Tables 9–10Multimodal ASR: table 9 on the 110 image behaviors, table 10 on text-only behaviors handed to a vision model. See the multimodal type page.
Table 11MT-Bench against average ASR for Zephyr, Mistral, Koala and Zephyr+R2D2 — the capability cost of the defense, and the only capability measurement in the paper.
Table 12Searchability: 20 sampled behaviors per dataset, 10-minute Google limit each. The empirical backing for the differential-harm principle.
Figures 9–11Per-semantic-category ASR, per-family category ASR, and standard-vs-contextual-vs-copyright ASR. Figures only. They carry claims in their captions, not numbers in a grid, and this map quotes the captions rather than eyeballing bar heights.

What the paper does not establish

There is no per-semantic-category ASR table. The evidence is figures 9 and 10 and their captions: copyright sits far below the other six categories, which are roughly level, and which is easiest depends on the model family. Any per-category ranking quoted elsewhere with decimals was read off a bar chart or invented.

Over-refusal is not measured anywhere. No benign or false-refusal set ships or is evaluated, so a model that refuses every input scores 0% ASR. The one capability check is table 11’s MT-Bench column, four models, not the 33-target grid.

System-level defenses are out of scope by construction. Input filters, output classifiers and prompt cleansing are absent; §B.3 says why: a fixed set of attacks cannot fairly evaluate a defense that would call for its own adaptive attack.

“For simplicity of evaluation with a fixed set of attacks, we focus our large-scale comparison on model-level defenses, although future work could use HarmBench for evaluating system-level defenses.”— HarmBench, §B.3

The repository

github.com/centerforaisafety/HarmBench, branch main, cloned 2026-09-08. Everything mechanical here (counts, tags, splits, config defaults) was read from the working tree, not the paper’s prose.

File or directoryWhat it establishes here
data/behavior_datasets/*.csvEvery count on this site: 510 total, 400 text and 110 multimodal, the 320/80 test/val split, the functional breakdown (200 standard, 100 contextual, 100 copyright), the seven semantic-category totals, and the fact that exactly two tags exist — context on the 100 contextual rows, hash_check paired with book or lyrics on the 100 copyright rows. Parsed, not quoted.
eval_utils.pyThe whole grading layer. LLAMA2_CLS_PROMPT at lines 309–357 is the judging rubric reproduced verbatim on the classifier page. compute_results_classifier (~396) is the three-way dispatch on tags. compute_results_hashing (~367) is the MinHash path, with compute_hashes_with_sliding_window (~223) supplying the 300/200 and 50/40 window sizes and check_output_with_sliding_window (~247) the 0.6 Jaccard threshold. compute_results_advbench (~359) is the refusal-prefix metric, kept for comparison and explicitly not the HarmBench metric.
api_models.pyThe five closed-model wrappers — GPT, GPTV, Claude, Gemini, Mistral — and that dispatch is by model-name prefix. Also the snag on the pipeline page: the OpenAI client is constructed with no base_url, so pointing HarmBench at an OpenAI-compatible endpoint that is not OpenAI is a code change, not a config change.
configs/pipeline_configs/run_pipeline.yamlThe runner’s view of every method: class_name, experiment_name_template, behavior_chunk_size, base_num_gpus, run_ids, and the allowed_target_model_types field that decides which attacks may point at which class of model. 24 active entries, with HumanJailbreaks-all and PAP-avg commented out as too expensive on closed models.
configs/method_configs/*.yaml21 files; every hyperparameter quoted on the attack pages — step counts, search widths, batch sizes, mutation rates, attacker-model choices, query budgets.
configs/model_configs/models.yaml26 top-level target entries and the model_type field (open_source, closed_source, open_source_multimodal, closed_source_multimodal) that gates attacks against targets. See the pipeline page.
baselines/<method>/The 18 method implementations. Read for mechanism — what is searched over, what the objective is, what one step costs — and never quoted.
adversarial_training/R2D2’s training code, built on alignment-handbook. Confirms the defense is an SFT-stage modification, not a separate system.
docs/evaluation_pipeline.mdThe four steps and the --step / --mode flags. Backs the pipeline page.
docs/configs.mdHow the config expansion works — method configs crossed with model configs into experiments.
docs/behavior_datasets.mdThe CSV schemas, including that ContextString is a first-class column and not an optional extra. See contextual behaviors.
docs/codebase_structure.mdWhere each moving part lives; used to navigate rather than cited.
baselines/ (index)The directory listing itself is evidence: 18 attack packages plus baseline.py and shared refusal-checking utilities, which is how the method families on the attack pages were grouped.
What was deliberately not read, and is not reproduced anywhere on this site. The Behavior and ContextString columns of the behavior CSVs; the contents of data/optimizer_targets/; the template text in baselines/human_jailbreaks/; the seed prompts in baselines/gptfuzz/GPTFuzzer.csv; the persuasion templates in baselines/pap/. Only counts, tags, splits and file shapes were taken from them. The exceptions: the behavior index reproduces all 510 behaviors and their context strings (the published dataset, not tooling), each linked to its line in the source CSV; and two 2023 GCG suffixes from the authors' own disclosure, shown as specimens.

Models, archives and the project site

Three things ship outside the repository:

ArtifactWhat it is, and what it settles
cais/HarmBench-Llama-2-13b-clsThe test classifier. A Llama 2 13B Chat fine-tune distilled from GPT-4-0613 labels; ~93% agreement with humans. This is the grader every reported ASR on this site was produced by, and the one you must never optimise against.
cais/HarmBench-Mistral-7b-val-clsThe validation classifier, Mistral 7B base, trained on half the test classifier’s data; 88.6% agreement. The one you should tune against. The two disagree on 26 examples of the human-labelled set.
…-cls-multimodal-behaviorsThe multimodal grader. Establishes that image behaviors are judged from the RedactedImageDescription text, never from the image — see the multimodal page.
zenodo.org/records/10714577“HarmBench Results (1.0)”, a single 10.4 GB zip published 2024-02-27: the generated test cases, the completions and the per-test-case labels behind every table in the paper. This is where anyone who wants per-semantic-category numbers has to go — the labels are per behavior, so a category breakdown is a group-by away, but it is not published in either the paper or the repo. It is also the only way to reproduce the reported grid without paying the GPU cost of regenerating test cases.
HarmBench PR #19The 1.0 release. Useful as a boundary marker: it dates the state of the code the paper’s numbers came from, against the main read here.
harmbench.orgThe project site — leaderboard framing and pointers. Nothing on this map depends on it; it is listed because it is the canonical landing page.

The Zenodo record answers browsers but returned 403 to the automated request while compiling this page; the block is traffic-based, not a missing record. Treat it as live.

Primary papers for the red teaming methods

HarmBench implements 18 methods but not 18 papers. Several columns are variants of one method (GCG-Multi and GCG-Transfer are ensemble configurations of GCG; TAP-Transfer is TAP with a fixed attacker and judge), two are HarmBench’s own baselines, and one grading-table entry is a classifier, not an attack. Every arXiv identifier below was fetched and its title checked against the method claimed.

Gradient and token-level — attacks/gradient

MethodPrimary paperarXiv
GCG, GCG-Multi, GCG-TransferZou et al. 2023, Universal and Transferable Adversarial Attacks on Aligned Language Models2307.15043
AutoPromptShin et al. 2020, AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts2010.15980
PEZWen et al. 2023, Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery2302.03668
GBDAGuo et al. 2021, Gradient-based Adversarial Attacks against Text Transformers2104.13733
UATWallace et al. 2019, Universal Adversarial Triggers for Attacking and Analyzing NLP1908.07125

UAT (2019) and AutoPrompt (2020) predate the jailbreaking literature: written for model analysis and prompt discovery, not safety. Their use as red teaming baselines is HarmBench’s framing, not the original authors’.

LLM-optimizer — attacks/llm-optimizer

MethodPrimary paperarXiv
AutoDANLiu et al. 2023, AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models2310.04451
PAIRChao et al. 2023, Jailbreaking Black Box Large Language Models in Twenty Queries2310.08419
TAP, TAP-TransferMehrotra et al. 2023, Tree of Attacks: Jailbreaking Black-Box LLMs Automatically2312.02119
GPTFuzzYu et al. 2023, GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts2309.10253
PAP-top5Zeng et al. 2024, How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs2401.06373

These are not always the original implementations: where a method used a closed model as attacker or judge, HarmBench substitutes Mixtral 8x7B to keep compute comparable, and it uses the in-context PAP because the fine-tuned paraphraser is not public. Both can lower a method’s measured ASR relative to its own paper.

Template, few-shot and human — attacks/template

MethodPrimary paperarXiv
ZeroShot, Stochastic Few-ShotPerez et al. 2022, Red Teaming Language Models with Language Models — the paper HarmBench reimplements these two from, with acknowledged modifications (an adjusted zero-shot prompt, an iterative SFS)2202.03286
Human JailbreaksShen et al. 2023a, “Do Anything Now”: Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models — the in-the-wild corpus the fixed template set is drawn from2308.03825
ArtPromptJiang et al. 2024, ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs2402.11753
DirectRequestNo paper — a HarmBench baseline that sends the behavior unmodified. Its Table 6 average of 25.3 is the floor every other column should be read against.–

Llama Guard (Inan et al. 2023, arXiv:2312.06674) is a citation here but not an attack: it appears in tables 3 and 4 as a grader HarmBench’s classifier is benchmarked against, and in table 4’s third prequalification set as one that fails, because it answers “is this harmful?” when the question is “is this this behavior?” See the grader page.

Where the sources disagree

Four mismatches turned up cross-checking repository against paper, three material and one cosmetic:

1. The validation split: 40/20 in the paper, 41/19 in the CSV

§B.2 describes a 100-behavior validation set of 20 multimodal, 20 contextual, 20 copyright and 40 standard behaviors. The shipped harmbench_behaviors_text_val.csv has 80 text rows: 41 standard, 19 contextual, 20 copyright. The paper’s account of the split explains it:

“These were selected with stratified sampling using the intersections between all functional behaviors and semantic categories, followed by manual adjustment to ensure the appropriate number of behaviors in each category.”— HarmBench, §B.2

The same paragraph notes the split was drawn against an earlier version of the semantic categories, so the per-category validation percentages do not exactly match the final taxonomy. One behavior out of 80 is a live example of the rule this map follows: when the CSV and the prose disagree about a count, the CSV is the benchmark.

2. §6.3’s prose figures are not the table’s figures

Section 6.3 states R2D2’s headline result against the Llama 2 Chat family, quoting a GCG ASR of 31.8 → 5.9, but neither number is in the v2 tables. Table 6 (400-behavior all-behaviors) gives Llama 2 7B Chat under GCG at 32.5 and Zephyr 7B + R2D2 at 5.5; Table 7 (320-behavior test split) gives 31.9 and 6.3. The prose pair matches neither: it takes the low end of one and the high end of the other.

This map uses the table figures throughout, preferring the all-behaviors slice unless a page says otherwise; the difference between 5.5 and 6.3 is smaller than the classifier’s own ~7% label noise.

3. The Table 6 copyright “Average” row is arithmetically impossible

Table 6’s Copyright Behaviors sub-table prints an Average row beginning 50.7 41.6 36.5 … 25.6, yet every cell in that sub-table lies between 0.0 and 28.0: 386 cells across 30 model rows, checked. No values in that range average to 50.7. Recomputing the column means gives numbers an order of magnitude smaller, agreeing with the copyright averages the paper prints in tables 7 and 8:

Copyright column averageGCGGCG-MGCG-TPEZAutoDANHumanDR
Table 6, as printed (all 510)50.741.636.528.248.826.225.6
Table 6, recomputed from its cells4.63.34.34.06.04.37.3
Table 7, test split, as printed4.73.54.44.26.54.67.5
Table 8, validation split, as printed3.92.93.93.33.83.26.7

The recomputed figures also reconcile the all-behaviors row exactly, which the printed ones cannot. Weighting the three functional sub-tables by size, (200×69.1 + 100×74.8 + 100×4.6) / 400 = 54.3, the printed GCG average; per model too, Llama 2 7B Chat under GCG, (200×34.5 + 100×58.0 + 100×3.0) / 400 = 32.5, its printed all-behaviors cell. Substitute the printed copyright average and nothing reconciles.

So the printed row is a typesetting error in v2, and this map reports the recomputed averages everywhere. The full arithmetic is on the results page. With the correct numbers, attack success on copyright behaviors is near the floor for every method, matching §B.5.2 (a metric that fires only on near-verbatim regurgitation) and figure 9’s caption.

4. The multimodal tag that is never set

compute_results_classifier branches on a multimodal tag to decide whether to feed the classifier a RedactedImageDescription, but the shipped multimodal CSV’s Tags column has no values, so the branch is unreachable from the shipped data alone; multimodal runs must tag rows upstream or supply their own dataframe. See the multimodal page.

How this map was built

Counts were parsed, not quoted. Every behavior count (per category, functional type, split, tag) came from reading the shipped CSVs programmatically, none from the paper. They match the paper’s totals, with the single 41/19 exception above.

Results came from the arXiv v2 HTML tables, cross-checked as grids; where a figure and a table both bear on a claim, the table was used.

Per-category ASR is not available and is not invented anywhere on this site. Six of the seven harm pages state that the paper does not tabulate ASR by semantic category and stop there. The seventh, copyright, has real numbers because it is also a functional type with its own sub-table, and those are the recomputed averages, not the printed row.

What this site carries, and what it does not. The dataset is here: all 510 behaviors and their context strings are reproduced on the behavior index from the MIT-licensed repository, and every harm page shows its examples verbatim. The attack tooling is not: no jailbreak template, persuasion script, model completion, or suffix optimised against a current model. Method pages explain mechanisms and carry no payloads, with one marked exception: two 2023 GCG suffixes from the authors’ own disclosure, against Vicuna and Llama 2 7B Chat, long since non-transferable. The other verbatim block is the classifier’s judging rubric on the grader page, a grading instruction, not an attack. Context passages appear on the behavior index, not inline, where at up to 3,544 characters they would swamp the argument.

Every external URL on this page was checked at compile time: 33 URLs, 32 returning 200, the Zenodo record a traffic-based 403 to automated clients while live in a browser. Each cited arXiv identifier was fetched and its title compared against the method attributed. Compiled 2026-09-08 against the repository at main and arXiv:2402.04249v2.

Next: back to the field map, or jump to the results page for the copyright reconciliation in full, the behavior index for all 510 IDs, or the sources section of the spine.