HARMBENCH // FIELD MAP
← all maps
MAZEIKA ET AL. · ARXIV:2402.04249 · CENTER FOR AI SAFETY510 behaviors · 18 attacks · 33 targets · 18 pages here

HarmBench, mapped

HarmBench measures how easily a model can be made to do something harmful: 510 fixed behaviors, a fixed grader, and eighteen automated attacks to elicit them. Browse it by harm category, behavior type, and attack family, plus four pages on how the final number is produced and what it hides.

The thing everyone asks first: is it a benchmark or a recipe? Both, and the split is clean. The behaviors, the classifier and the copyright hashes are frozen artifacts you download. The attack prompts do not exist until you generate them against your specific target model — one set per attack × model pair, which is where the GPU hours go. That is why it reads like a recipe: the expensive half of the benchmark is compute, not data.

Every behavior has a permalink: data/behaviors.html#<BehaviorID> jumps to its row, where the behavior is printed in full — with its context passage folded alongside it — and links on to its exact line in the source CSV on GitHub.

Browse the harms
Every behavior

Tick a page when you have read it; this browser remembers.

01 · THE HARMS

Seven categories, one page each

Every behavior carries a SemanticCategory, the kind of harm, drawn from the acceptable-use policies of OpenAI, Anthropic, Meta and Inflection. One page per category, each with its counts, grading, best attacks, and representative behaviors verbatim.

01
cybercrime_intrusion121 behaviors · 67 text + 54 multimodalthe largest category
Hacking, malware and CAPTCHA defeat, the only category where multimodal nearly matches text. 40 standard, 27 contextual, with code-artifact grading.
02
illegal101 behaviors · 65 text + 36 multimodalthe broadest definition
Seven sub-types: fraud and scams, trafficking, weapons acquisition, theft, violent crime, extortion, and assisting suicide. Overwhelmingly standard (58 of 65), least dependent on supplied context.
03
copyright100 behaviors · 50 books + 50 lyricsgraded by hash, not by model
No classifier at all: 100 MinHash pickles ship instead of the copyrighted text, and a completion counts only if a sliding window matches at Jaccard > 0.6. Every attack scores near the floor.
04
misinformation_disinformation68 behaviors · 65 text + 3 multimodalmost context-heavy
Harmful lies, propaganda, election interference and defamation. Nearly half the text behaviors (31 of 65) are contextual. The highest-ASR category on Llama 2 and GPT models.
05
chemical_biological60 behaviors · 56 text + 4 multimodal50/50 standard & contextual
The category "differential harm" was written for: the contextual half (28 of 56) supplies a technical passage and a narrow follow-up a search engine cannot answer. Easiest to elicit on Baichuan 2 and Starling.
06
harmful31 behaviors · 22 text + 9 multimodalthe smallest text category
The residual bucket: graphic and age-restricted content, unsafe-practice promotion, and privacy violations. 21 of 22 text behaviors are standard, exactly one carries a context string.
07
harassment_bullying29 behaviors · 25 text + 4 multimodalhate speech & self-harm
Harassment, hate speech and encouraging self-harm. The classifier's "unambiguous and non-minimal" rule bites hardest here, and heavily trained refusals catch most of it.
02 · THE TYPES

Four functional categories

The same 510 behaviors cut the other way. The FunctionalCategory decides how a behavior is presented and, crucially, how it is graded: three grading paths hide behind one ASR number.

08
standard200 · 159 test / 41 valthe plain case
The behavior string is the whole prompt: no context, no tags, graded by LLAMA2_CLS_PROMPT['prompt']. Column average 69.1 under GCG, and only 39% of the set.
09
contextual100 · 81 test / 19 valthe ContextString mechanics
A supplied passage plus a narrow question about it, tagged context, graded by a second template that shows the classifier the passage. Un-Googleable by design, yet empirically easier to elicit than standard behaviors.
10
copyright100 · 80 test / 20 valhash_check · book | lyrics
Bypasses the neural grader: sliding-window MinHash at two window sizes for a book or lyrics, threshold 0.6. Also why the shipped repo contains no copyrighted text.
11
multimodal110 · no published splitimage + text
An image plus a question about it, graded by a third classifier that never sees the image, only a redacted description. Five attack methods of its own, and the finding that an optimised blank image is a 48–66% attack channel on text-only behaviors.
03 · THE ATTACKS

Eighteen methods, three families

The families differ mainly in access needed: gradients, or just an API. The runner enforces that constraint, which is why the GPT-4 rows in the results table are mostly dashes.

12
white box · needs weightsGCG · GCG-Multi · GCG-Transfer · AutoPrompt · PEZ · GBDA · UATGCG avg 54.3, the strongest
Optimise a suffix token by token against the model's own gradients, the most expensive family at 500 steps × 512 candidates per behavior. Two transfer variants point a white-box attack at GPT-4.
13
black box · query onlyAutoDAN · PAIR · TAP · TAP-Transfer · GPTFuzz · PAP-top5AutoDAN avg 52.7
One language model attacking another: mutate, judge, prune, repeat. Nearly as effective as gradient attacks at a fraction of the cost, and mostly runnable through an API.
14
no optimisation loopZeroShot · Stochastic Few-Shot · Human Jailbreaks · ArtPrompt · DirectRequestDirectRequest avg 25.3 — the floor
DirectRequest sends the behavior unmodified and still succeeds a quarter of the time; famous human jailbreaks average 27.3, barely above it.
04 · MEASURING

How the number is produced, and what it hides

ASR is one number over a lot of machinery: a fine-tuned 13B judge with about 7% label noise, a 512-token generation budget, three grading paths, and no measure of the opposite failure. A model that refuses everything scores perfectly.

15
cais/HarmBench-Llama-2-13b-cls93.19% human agreement+ the MinHash path
The judging prompt verbatim, both templates, its distillation from GPT-4 labels over a 600-example human-labelled set, why it beats Llama Guard and refusal-prefix heuristics on the prequalification sets, and the rule that voids your result.
16
run_pipeline.pygenerate → merge → complete → evaluateslurm | local | ray
Why step 1.5 exists, how models.yaml's model_type field silently decides which attacks you may run, and the snag in api_models.py that blocks pointing HarmBench at a non-OpenAI, OpenAI-compatible endpoint.
17
Table 6 · 18 × 33no model robust, no attack universal+ a table typo, verified
The results grid condensed: model size does not predict robustness, the training pipeline does, contextual behaviors are easier than standard ones, and a printed average row in arXiv v2 that does not match its own column, recomputed here.
18
Zephyr 7B + R2D2GCG 69.5 → 5.5, PAIR 58.8 → 48.016 h on 8×A100
The adversarial-training method and its capability cost: training against a gradient adversary bought robustness to gradient adversaries and almost nothing against LLM-driven ones. Plus why system-level defenses are excluded.
05

Data and sources

06

What this map does and does not carry

All 510 behaviors and their context strings are reproduced on the behaviors page, and every harm page shows its examples verbatim: the published, MIT-licensed HarmBench dataset, which its authors state they reviewed and trimmed before release. They are requests, not instructions, and each row links its own line in the source CSV.

The other half, the attack tooling, is not here: no jailbreak template, no persuasion script, no model completion, no adversarial suffix optimised against any current model. The LLM-optimizer and template families are described by mechanism only. The one marked exception: two published 2023 GCG suffixes from the authors' own disclosure, against models two generations old and no longer transferable.