Sources, verified
Organised by artifact rather than by claim: for each source, what it establishes and what people routinely assume it establishes but does not.
The paper
Mazeika, Phan, Yin, Zou, Wang, Mu, Sakhaee, Li, Basart, Li, Forsyth and Hendrycks, HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. arXiv:2402.04249: v1 2024-02-06, v2 2024-02-27. This map reads the v2 HTML and follows v2 where the versions differ.
The paper is two artifacts: a benchmark-design argument (sections 3 and 4, appendix B) and an empirical comparison (section 6, appendix D). Which piece supports what:
| Section or table | What it establishes, and where this map uses it |
|---|---|
| §3.1 | The ASR definition — mean classifier label over a method’s test cases for a behavior, with greedy decoding and a fixed 512-new-token budget. The single most load-bearing definition on the site; read on the pipeline page and assumed by every number on the results page. |
| §4.2 | Curation of harmful behaviors: the design principles — breadth, differential harm (prefer behaviors a search engine does not already answer), dual-intent handling, and comparability. This is why the benchmark looks the way it does. |
| §5 | R2D2: the adversarial-training method itself — the refreshed pool of GCG test cases, the “toward” and “away” losses, and the training loop folded into Zephyr’s SFT code. See defenses. |
| §6.3 | R2D2’s results in prose, including the comparison against the Llama 2 Chat family. Its quoted figures disagree slightly with the v2 tables — see discrepancies. |
| §B.3 | Supported threat models: no-access (transfer), query-access and parameter-access attacks; model-level versus system-level defenses; and the explicit decision to scope the large-scale comparison to model-level defenses only. |
| §B.4 | The seven semantic categories and their sub-bullets. The taxonomy quoted on each harm page comes from here, not from the CSV, which carries only the machine slug. |
| §B.5.1 | How the Llama 2 13B evaluation classifier was trained — template-generated completions, GPT-4-0613 labels, distillation, and the separate Mistral validation classifier fine-tuned on half the data. See the grader. |
| §B.5.2 | Why copyright is graded by hashing rather than by an LLM judge: works inspired by a copyrighted original are not reliably separable from attempts to reproduce it, so the benchmark measures only near-verbatim regurgitation. Underwrites the copyright type page. |
| Table 3 | Classifier agreement with human labels on 600 hand-labelled examples, against AdvBench, GPTFuzz, ChatGLM, Llama Guard and GPT-4. The ~93% figure — and therefore the ~7% label-noise caveat repeated across this site — comes from here. Read. |
| Table 4 | The three prequalification sets, which is where you learn how the prior graders fail rather than just that they do: refusal-prefix heuristics call benign text harmful, and “is this harmful?” graders cannot tell a completion of the wrong behavior from a completion of the right one. |
| Table 5 | Dataset comparison against nine prior behavior sets. Establishes 510 unique behaviors and that HarmBench is the only one of the ten carrying both multimodal and contextual behaviors. |
| Table 6 | ASR on the all-behaviors slice — the 400 text behaviors, since the sub-table weights are 200 + 100 + 100 and the 110 multimodal behaviors are Table 9 — with sub-tables for standard, contextual and copyright. The main grid of the results page, and the source of every attack-column average quoted on the attack pages. |
| Table 7 | The same grid on the 320-behavior text test split. Used here as an independent check on Table 6, and it is what catches the copyright Average row. |
| Table 8 | The same grid on the 80-behavior validation split. Second check, same role. |
| Tables 9–10 | Multimodal ASR: table 9 on the 110 image behaviors, table 10 on text-only behaviors handed to a vision model. See the multimodal type page. |
| Table 11 | MT-Bench against average ASR for Zephyr, Mistral, Koala and Zephyr+R2D2 — the capability cost of the defense, and the only capability measurement in the paper. |
| Table 12 | Searchability: 20 sampled behaviors per dataset, 10-minute Google limit each. The empirical backing for the differential-harm principle. |
| Figures 9–11 | Per-semantic-category ASR, per-family category ASR, and standard-vs-contextual-vs-copyright ASR. Figures only. They carry claims in their captions, not numbers in a grid, and this map quotes the captions rather than eyeballing bar heights. |
What the paper does not establish
There is no per-semantic-category ASR table. The evidence is figures 9 and 10 and their captions: copyright sits far below the other six categories, which are roughly level, and which is easiest depends on the model family. Any per-category ranking quoted elsewhere with decimals was read off a bar chart or invented.
Over-refusal is not measured anywhere. No benign or false-refusal set ships or is evaluated, so a model that refuses every input scores 0% ASR. The one capability check is table 11’s MT-Bench column, four models, not the 33-target grid.
System-level defenses are out of scope by construction. Input filters, output classifiers and prompt cleansing are absent; §B.3 says why: a fixed set of attacks cannot fairly evaluate a defense that would call for its own adaptive attack.
“For simplicity of evaluation with a fixed set of attacks, we focus our large-scale comparison on model-level defenses, although future work could use HarmBench for evaluating system-level defenses.”— HarmBench, §B.3
The repository
github.com/centerforaisafety/HarmBench, branch main, cloned 2026-09-08. Everything mechanical here (counts, tags, splits, config defaults) was read from the working tree, not the paper’s prose.
| File or directory | What it establishes here |
|---|---|
| data/behavior_datasets/*.csv | Every count on this site: 510 total, 400 text and 110 multimodal, the 320/80 test/val split, the functional breakdown (200 standard, 100 contextual, 100 copyright), the seven semantic-category totals, and the fact that exactly two tags exist — context on the 100 contextual rows, hash_check paired with book or lyrics on the 100 copyright rows. Parsed, not quoted. |
| eval_utils.py | The whole grading layer. LLAMA2_CLS_PROMPT at lines 309–357 is the judging rubric reproduced verbatim on the classifier page. compute_results_classifier (~396) is the three-way dispatch on tags. compute_results_hashing (~367) is the MinHash path, with compute_hashes_with_sliding_window (~223) supplying the 300/200 and 50/40 window sizes and check_output_with_sliding_window (~247) the 0.6 Jaccard threshold. compute_results_advbench (~359) is the refusal-prefix metric, kept for comparison and explicitly not the HarmBench metric. |
| api_models.py | The five closed-model wrappers — GPT, GPTV, Claude, Gemini, Mistral — and that dispatch is by model-name prefix. Also the snag on the pipeline page: the OpenAI client is constructed with no base_url, so pointing HarmBench at an OpenAI-compatible endpoint that is not OpenAI is a code change, not a config change. |
| configs/pipeline_configs/run_pipeline.yaml | The runner’s view of every method: class_name, experiment_name_template, behavior_chunk_size, base_num_gpus, run_ids, and the allowed_target_model_types field that decides which attacks may point at which class of model. 24 active entries, with HumanJailbreaks-all and PAP-avg commented out as too expensive on closed models. |
| configs/method_configs/*.yaml | 21 files; every hyperparameter quoted on the attack pages — step counts, search widths, batch sizes, mutation rates, attacker-model choices, query budgets. |
| configs/model_configs/models.yaml | 26 top-level target entries and the model_type field (open_source, closed_source, open_source_multimodal, closed_source_multimodal) that gates attacks against targets. See the pipeline page. |
| baselines/<method>/ | The 18 method implementations. Read for mechanism — what is searched over, what the objective is, what one step costs — and never quoted. |
| adversarial_training/ | R2D2’s training code, built on alignment-handbook. Confirms the defense is an SFT-stage modification, not a separate system. |
| docs/evaluation_pipeline.md | The four steps and the --step / --mode flags. Backs the pipeline page. |
| docs/configs.md | How the config expansion works — method configs crossed with model configs into experiments. |
| docs/behavior_datasets.md | The CSV schemas, including that ContextString is a first-class column and not an optional extra. See contextual behaviors. |
| docs/codebase_structure.md | Where each moving part lives; used to navigate rather than cited. |
| baselines/ (index) | The directory listing itself is evidence: 18 attack packages plus baseline.py and shared refusal-checking utilities, which is how the method families on the attack pages were grouped. |
Models, archives and the project site
Three things ship outside the repository:
| Artifact | What it is, and what it settles |
|---|---|
| cais/HarmBench-Llama-2-13b-cls | The test classifier. A Llama 2 13B Chat fine-tune distilled from GPT-4-0613 labels; ~93% agreement with humans. This is the grader every reported ASR on this site was produced by, and the one you must never optimise against. |
| cais/HarmBench-Mistral-7b-val-cls | The validation classifier, Mistral 7B base, trained on half the test classifier’s data; 88.6% agreement. The one you should tune against. The two disagree on 26 examples of the human-labelled set. |
| …-cls-multimodal-behaviors | The multimodal grader. Establishes that image behaviors are judged from the RedactedImageDescription text, never from the image — see the multimodal page. |
| zenodo.org/records/10714577 | “HarmBench Results (1.0)”, a single 10.4 GB zip published 2024-02-27: the generated test cases, the completions and the per-test-case labels behind every table in the paper. This is where anyone who wants per-semantic-category numbers has to go — the labels are per behavior, so a category breakdown is a group-by away, but it is not published in either the paper or the repo. It is also the only way to reproduce the reported grid without paying the GPU cost of regenerating test cases. |
| HarmBench PR #19 | The 1.0 release. Useful as a boundary marker: it dates the state of the code the paper’s numbers came from, against the main read here. |
| harmbench.org | The project site — leaderboard framing and pointers. Nothing on this map depends on it; it is listed because it is the canonical landing page. |
The Zenodo record answers browsers but returned 403 to the automated request while compiling this page; the block is traffic-based, not a missing record. Treat it as live.
Primary papers for the red teaming methods
HarmBench implements 18 methods but not 18 papers. Several columns are variants of one method (GCG-Multi and GCG-Transfer are ensemble configurations of GCG; TAP-Transfer is TAP with a fixed attacker and judge), two are HarmBench’s own baselines, and one grading-table entry is a classifier, not an attack. Every arXiv identifier below was fetched and its title checked against the method claimed.
Gradient and token-level — attacks/gradient
| Method | Primary paper | arXiv |
|---|---|---|
| GCG, GCG-Multi, GCG-Transfer | Zou et al. 2023, Universal and Transferable Adversarial Attacks on Aligned Language Models | 2307.15043 |
| AutoPrompt | Shin et al. 2020, AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts | 2010.15980 |
| PEZ | Wen et al. 2023, Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery | 2302.03668 |
| GBDA | Guo et al. 2021, Gradient-based Adversarial Attacks against Text Transformers | 2104.13733 |
| UAT | Wallace et al. 2019, Universal Adversarial Triggers for Attacking and Analyzing NLP | 1908.07125 |
UAT (2019) and AutoPrompt (2020) predate the jailbreaking literature: written for model analysis and prompt discovery, not safety. Their use as red teaming baselines is HarmBench’s framing, not the original authors’.
LLM-optimizer — attacks/llm-optimizer
| Method | Primary paper | arXiv |
|---|---|---|
| AutoDAN | Liu et al. 2023, AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models | 2310.04451 |
| PAIR | Chao et al. 2023, Jailbreaking Black Box Large Language Models in Twenty Queries | 2310.08419 |
| TAP, TAP-Transfer | Mehrotra et al. 2023, Tree of Attacks: Jailbreaking Black-Box LLMs Automatically | 2312.02119 |
| GPTFuzz | Yu et al. 2023, GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts | 2309.10253 |
| PAP-top5 | Zeng et al. 2024, How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs | 2401.06373 |
These are not always the original implementations: where a method used a closed model as attacker or judge, HarmBench substitutes Mixtral 8x7B to keep compute comparable, and it uses the in-context PAP because the fine-tuned paraphraser is not public. Both can lower a method’s measured ASR relative to its own paper.
Template, few-shot and human — attacks/template
| Method | Primary paper | arXiv |
|---|---|---|
| ZeroShot, Stochastic Few-Shot | Perez et al. 2022, Red Teaming Language Models with Language Models — the paper HarmBench reimplements these two from, with acknowledged modifications (an adjusted zero-shot prompt, an iterative SFS) | 2202.03286 |
| Human Jailbreaks | Shen et al. 2023a, “Do Anything Now”: Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models — the in-the-wild corpus the fixed template set is drawn from | 2308.03825 |
| ArtPrompt | Jiang et al. 2024, ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs | 2402.11753 |
| DirectRequest | No paper — a HarmBench baseline that sends the behavior unmodified. Its Table 6 average of 25.3 is the floor every other column should be read against. | – |
Llama Guard (Inan et al. 2023, arXiv:2312.06674) is a citation here but not an attack: it appears in tables 3 and 4 as a grader HarmBench’s classifier is benchmarked against, and in table 4’s third prequalification set as one that fails, because it answers “is this harmful?” when the question is “is this this behavior?” See the grader page.
Where the sources disagree
Four mismatches turned up cross-checking repository against paper, three material and one cosmetic:
1. The validation split: 40/20 in the paper, 41/19 in the CSV
§B.2 describes a 100-behavior validation set of 20 multimodal, 20 contextual, 20 copyright and 40 standard behaviors. The shipped harmbench_behaviors_text_val.csv has 80 text rows: 41 standard, 19 contextual, 20 copyright. The paper’s account of the split explains it:
“These were selected with stratified sampling using the intersections between all functional behaviors and semantic categories, followed by manual adjustment to ensure the appropriate number of behaviors in each category.”— HarmBench, §B.2
The same paragraph notes the split was drawn against an earlier version of the semantic categories, so the per-category validation percentages do not exactly match the final taxonomy. One behavior out of 80 is a live example of the rule this map follows: when the CSV and the prose disagree about a count, the CSV is the benchmark.
2. §6.3’s prose figures are not the table’s figures
Section 6.3 states R2D2’s headline result against the Llama 2 Chat family, quoting a GCG ASR of 31.8 → 5.9, but neither number is in the v2 tables. Table 6 (400-behavior all-behaviors) gives Llama 2 7B Chat under GCG at 32.5 and Zephyr 7B + R2D2 at 5.5; Table 7 (320-behavior test split) gives 31.9 and 6.3. The prose pair matches neither: it takes the low end of one and the high end of the other.
This map uses the table figures throughout, preferring the all-behaviors slice unless a page says otherwise; the difference between 5.5 and 6.3 is smaller than the classifier’s own ~7% label noise.
3. The Table 6 copyright “Average” row is arithmetically impossible
Table 6’s Copyright Behaviors sub-table prints an Average row beginning 50.7 41.6 36.5 … 25.6, yet every cell in that sub-table lies between 0.0 and 28.0: 386 cells across 30 model rows, checked. No values in that range average to 50.7. Recomputing the column means gives numbers an order of magnitude smaller, agreeing with the copyright averages the paper prints in tables 7 and 8:
| Copyright column average | GCG | GCG-M | GCG-T | PEZ | AutoDAN | Human | DR |
|---|---|---|---|---|---|---|---|
| Table 6, as printed (all 510) | 50.7 | 41.6 | 36.5 | 28.2 | 48.8 | 26.2 | 25.6 |
| Table 6, recomputed from its cells | 4.6 | 3.3 | 4.3 | 4.0 | 6.0 | 4.3 | 7.3 |
| Table 7, test split, as printed | 4.7 | 3.5 | 4.4 | 4.2 | 6.5 | 4.6 | 7.5 |
| Table 8, validation split, as printed | 3.9 | 2.9 | 3.9 | 3.3 | 3.8 | 3.2 | 6.7 |
The recomputed figures also reconcile the all-behaviors row exactly, which the printed ones cannot. Weighting the three functional sub-tables by size, (200×69.1 + 100×74.8 + 100×4.6) / 400 = 54.3, the printed GCG average; per model too, Llama 2 7B Chat under GCG, (200×34.5 + 100×58.0 + 100×3.0) / 400 = 32.5, its printed all-behaviors cell. Substitute the printed copyright average and nothing reconciles.
So the printed row is a typesetting error in v2, and this map reports the recomputed averages everywhere. The full arithmetic is on the results page. With the correct numbers, attack success on copyright behaviors is near the floor for every method, matching §B.5.2 (a metric that fires only on near-verbatim regurgitation) and figure 9’s caption.
4. The multimodal tag that is never set
compute_results_classifier branches on a multimodal tag to decide whether to feed the classifier a RedactedImageDescription, but the shipped multimodal CSV’s Tags column has no values, so the branch is unreachable from the shipped data alone; multimodal runs must tag rows upstream or supply their own dataframe. See the multimodal page.
How this map was built
Counts were parsed, not quoted. Every behavior count (per category, functional type, split, tag) came from reading the shipped CSVs programmatically, none from the paper. They match the paper’s totals, with the single 41/19 exception above.
Results came from the arXiv v2 HTML tables, cross-checked as grids; where a figure and a table both bear on a claim, the table was used.
Per-category ASR is not available and is not invented anywhere on this site. Six of the seven harm pages state that the paper does not tabulate ASR by semantic category and stop there. The seventh, copyright, has real numbers because it is also a functional type with its own sub-table, and those are the recomputed averages, not the printed row.
What this site carries, and what it does not. The dataset is here: all 510 behaviors and their context strings are reproduced on the behavior index from the MIT-licensed repository, and every harm page shows its examples verbatim. The attack tooling is not: no jailbreak template, persuasion script, model completion, or suffix optimised against a current model. Method pages explain mechanisms and carry no payloads, with one marked exception: two 2023 GCG suffixes from the authors’ own disclosure, against Vicuna and Llama 2 7B Chat, long since non-transferable. The other verbatim block is the classifier’s judging rubric on the grader page, a grading instruction, not an attack. Context passages appear on the behavior index, not inline, where at up to 3,544 characters they would swamp the argument.
Every external URL on this page was checked at compile time: 33 URLs, 32 returning 200, the Zenodo record a traffic-based 403 to automated clients while live in a browser. Each cited arXiv identifier was fetched and its title compared against the method attributed. Compiled 2026-09-08 against the repository at main and arXiv:2402.04249v2.