HARMBENCH // FIELD MAP
← field map
THE ATTACKS · 02 OF 036 methods · AutoDAN 52.7 · TAP-T 48.3 · TAP 45.2 · PAIR 40.7 · PAP-top5 16.6

LLM-optimizer attacks

One model writes the prompts, another grades them: the family that made closed-model evaluation possible at a fraction of GCG's compute.
TL;DR — Six methods replace the gradient with a generator and a judge: propose, query the target, score, keep the best, mutate, repeat. Top entry AutoDAN at 52.7 is 1.6 points off GCG's 54.3 with no backward pass, a gap smaller than the grader's label noise. Four can point at an API, so the GPT, Claude and Gemini rows of Table 6 have numbers at all. Caveats: HarmBench swapped every closed-source attacker and judge for Mixtral 8x7B, and GPTFuzz ships with no Table 6 column, so do not quote it an ASR.
Gradient & token-level Templates, few-shot & humans Table 6, read What the grader misses
methods here
6
5 with published averages, 1 without
best average
52.7
AutoDAN — 2nd of 16 columns
vs GPT-4 Turbo
58.5
TAP-Transfer, best cell in that row
API-runnable
4 / 6
AutoDAN and GPTFuzz need the weights
weakest column
16.6
PAP-top5 — zero target queries

The gradient family asks what the strongest attack is if you own the weights; this family asks what you can build with a chat endpoint and a budget: almost as much, for almost nothing.

The shape they share, and the access they actually need

Four of the six run the same loop: an attacker model proposes prompts for one behavior, each goes to the target, a judge scores the reply, and high scorers are kept, mutated or extended until one crosses a threshold or the budget runs out. The survivor is the one test case. Nothing differentiates through the target, so it can sit behind an HTTP call.

Check allowed_target_model_types in run_pipeline.yaml before assuming any is query-only: PAIR, TAP, TAP-Transfer and PAP-top5 carry [open_source, closed_source], while AutoDAN and GPTFuzz carry [open_source] only, for code reasons: AutoDAN's fitness is the target's cross-entropy loss (a logits read), GPTFuzz needs a local vLLM target with a RoBERTa judge. Gradient-free is not black-box; model_type in models.yaml enforces it. All six take behavior_chunk_size: 5; five want base_num_gpus: 2, GPTFuzz one. Underneath every number sits one substitution:

"Some methods use closed-source LLMs in their original implementations, most often as an evaluator to guide optimization or as an attacker LLM to optimize test cases. We replace these with Mixtral 8x7B to reduce the costs of running our large-scale comparison. While this can reduce their effectiveness, it has two substantial benefits: (1) reducing evaluation costs, and (2) enabling compute comparisons."— HarmBench, §C.1

Each average is thus a lower bound, the ordering the durable result; every judge steering these searches is weaker than the Llama 2 13B classifier that grades the outcome.

Not shown here. AutoDAN and GPTFuzz seed from corpora of in-the-wild jailbreak templates shipped in the repository, and PAP carries forty persuasion templates with worked examples. This page gives their size and role, not their content.

AutoDAN — Generating Stealthy Jailbreak Prompts on Aligned LLMs

needs target loss (open source only)class AutoDAN100 steps × 64 candidatesavg ASR 52.7

A genetic algorithm over whole prompt prefixes. Population: 64 strings from the first 64 of a 128-item corpus of handcrafted persona jailbreaks (a ~415 KB Python list). Each generation scores all 64 (prefix + behavior) by the target's cross-entropy loss on a fixed continuation; low loss wins. num_elites: 0.1 (six) survive; 58 come from softmax roulette over the negated losses, recombined at crossover: 0.5 on sentence boundaries with num_points: 5 cuts; only mutation: 0.01 (about 58 per run) reach mistralai/Mistral-7B-Instruct-v0.2 under vLLM to paraphrase (a GPT mutator is commented out). So AutoDAN is overwhelmingly sentence-level recombination of existing human jailbreaks, a model touched about once in a hundred:

"A semi-automated method that initializes test cases from handcrafted jailbreak prompts. These are then evolved using a hierarchical genetic algorithm to elicit specific behaviors from the target LLM."— HarmBench, §C.1

Every eval_steps: 5 generations (25 for Llama 2) it checks a 36-entry refusal-substring list of bare tokens like never, matched anywhere, so the emitted test case is the loss-argmin, not a verified success. model_short_name and developer_name exist because the seed templates address the assistant by name. Result 52.7 (second column): 68.3 standard, 68.4 contextual, 6.0 copyright, near-zero on Llama 2 (0.5 / 0.8 / 2.8). Liu et al., 2023, arXiv:2310.04451 (verified).

PAIR — Jailbreaking Black Box Large Language Models in Twenty Queries

chat API onlyclass PAIR20 streams × 3 rounds ≤ 60 queriesavg ASR 40.7

An attacker model, told the behavior and desired opening, produces a prompt plus a self-critique; it goes to the target; the reply returns with a judge score; the attacker revises. HarmBench runs n_streams: 20 conversations independently, steps: 3 deep (PAIR's defaults): worst case 60 target queries per behavior. Attacker and judge are both Mixtral-8x7B-Instruct-v0.1 on two GPUs, with attack_max_n_tokens: 500, target_max_n_tokens: 150, judge_max_n_tokens: 5 and cutoff_score: 10; the 1–10 judge score is read by a bracket regex that falls back to 1 when unreadable.

The attacker sees only 150 tokens of the reply while step 3 grades 512, a quarter of the grader's evidence. For contextual behaviors PAIR queries the target without the ContextString, prepending it only to the final test case, where TAP passes it through every query: PAIR's 60.5 contextual trails TAP's 61.8 and TAP-Transfer's 67.5. Overall 40.7 / 47.5 standard / 60.5 contextual / 7.5 copyright. Chao et al., 2023, arXiv:2310.08419 (verified).

TAP — Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

chat API onlyclass TAPdepth 10 × branch 4, width 10avg ASR 45.2

PAIR with branching and a pre-filter. One root (n_streams: 1) expands branching_factor: 4 ways per level for depth: 10 levels, frontier held to width: 10. Before the target, an on-topic judge prunes rewrites that no longer ask what the behavior asked; a second prune, on harmfulness score, follows the target round.

Worst case: level 0 yields 4 candidates, level 1 16, levels 2–9 40 each (340 attacker generations, 340 on-topic judge calls), but only 4 + 10 + 8×10 = 94 target queries, the frontier capped. Against PAIR's 60 target queries and 120 auxiliary calls, TAP buys +4.5 points at ~1.6× the target traffic and six times the attacker compute. Since phase 1 already truncates to width and both judges floor at 1 on an unparseable response, the phase-2 prune never removes a candidate in the default config.

45.2 all behaviors; 55.3 standard, 61.8 contextual, 8.5 copyright, the highest copyright cell of the sixteen columns, on a very low bar (see the copyright category). Mehrotra et al., 2023, arXiv:2312.02119 (verified).

TAP-Transfer — TAP run once against a surrogate, then replayed

no target access at replay timeclass TAP, experiment gpt-4-0613_judge_gpt-4-1106-preview_targetone attack run, 33 targetsavg ASR 48.3

Not a new algorithm: the same TAP class pinned to one experiment block, naming gpt-4-0613 as judge and gpt-4-1106-preview as target, attacker still Mixtral. TAP runs to completion against GPT-4 Turbo as a surrogate; the test cases are replayed unmodified at every other model, one forward pass each.

On the closed rows it is the best attack in the table: 54.8 on GPT-4-0613, 58.5 on GPT-4-Turbo-1106, 62.3 on GPT-3.5-Turbo-0613, against direct TAP's 43.0, 36.4 and 47.7. Replaying a prompt developed against a different model beats attacking that model directly, on exactly the models where adaptation should matter most. Exceptions: on Gemini Pro transfer loses (31.2 against 38.8), on all three Claude models it collapses to 0.8–1.5, below direct TAP and far below GCG-Transfer's 12.1 on Claude 1, and on Zephyr + R2D2 direct TAP wins 60.8 to 54.3.

Pinning a GPT judge routes it through the GPTJudge path, whose on_topic_score reads the numeric bracket regex while the on-topic rubric asks for [[YES]] or [[NO]]: no match, so every candidate takes the fallback score of 1 and the phase-1 topicality prune degenerates into keeping width candidates by index (the open-source judge path parses yes/no correctly, unaffected). TAP-Transfer scored 48.3 with its signature pruning disabled. Same paper as TAP.

GPTFuzz — GPTFUZZER: Red Teaming LLMs with Auto-Generated Jailbreak Prompts

local vLLM target + local judge (open source only)class GPTFuzzmax 1,000 queries / 100 iterationsno Table 6 column

GPTFuzz has code in baselines/gptfuzz/, a config file and a live row in run_pipeline.yaml, yet is not one of Table 6's sixteen columns nor one of the sixteen text-only methods in §6: there is no published HarmBench ASR for it. The only numbers HarmBench attaches to the name belong to GPTFuzz the classifier: 75.42% human agreement in Table 3 and 65.8 average on the prequalification sets in Table 4 (35.2 on completions harmful for the wrong behavior). A GPTFuzz attack ASR quoted as a HarmBench result did not come from this paper.

Mutation-based fuzzing with a bandit over seeds. The seed file is 77 in-the-wild jailbreak templates (id and text columns, 260 to 5,498 characters, median around 1,756), each with a slot for the behavior. Per iteration it picks a seed by Monte-Carlo tree search over past reward, applies one of five gpt-3.5-turbo mutation operators at temperature 0 (crossover, expand, generate-similar, rephrase, shorten), queries the target for max_new_tokens: 512 and scores the reply. energy: 1, max_jailbreak: 1, max_query: 1000 and max_iteration: 100 cap it.

Two config-versus-code gaps. seed_selection_strategy: round_robin is stored but never read: the driver hardcodes the MCTS policy. And the internal judge hubert233/GPTFuzz is the same RoBERTa Table 4 scores at 35.2 on wrong-behavior completions, so the fuzzer can stop early on a completion harmful but not this behavior, which the HarmBench classifier grades as a miss. Yu et al., 2023, arXiv:2309.10253 (verified).

PAP-top5 — How Johnny Can Persuade LLMs to Jailbreak Them

no target access at allclass PAP, experiment top_55 generations, 0 target queriesavg ASR 16.6

PAP starts from a taxonomy of forty persuasion techniques (name, definition, worked example). top_k_persuasion_taxonomy: 5 takes the first five in the PAP paper's Figure 7 order: Logical Appeal, Authority Endorsement, Misrepresentation, Evidence-based Persuasion, Expert Endorsement. Mixtral 8x7B rewrites the request once per technique, and for contextual behaviors the ContextString is prepended to each.

The whole method: no judge, no target query, no iteration, five generations and done. A language model writes its prompts rather than searching; structurally closer to ZeroShot than PAIR. PAP emits five test cases per behavior and all five are graded, where every other method emits one curated best, so 16.6 is the mean of five unguided rewrites, not one optimised prompt. HarmBench ran the in-context variant, not PAP's fine-tuned paraphraser (not public).

The column: 13.9 standard, 31.3 contextual (a technical passage to be persuasive about more than doubles it), 7.4 copyright. On permissive models PAP is worse than simply asking (32.9 against DirectRequest's 65.8 on Zephyr 7B); on the hardest it flips: 2.7 / 3.3 / 4.1 across the three Llama 2 Chat sizes against DirectRequest's 0.8 / 2.8 / 2.8. Yet 16.6 comes from five generations, zero target queries and fluent language with no adversarial suffix for a perplexity filter, where GCG's 54.3 needs a 500-step search and leaves a visible artefact. PAP-avg (the full_40 experiment) is commented out of run_pipeline.yaml as too expensive on closed models, so the full taxonomy has no number. Zeng et al., 2024, arXiv:2401.06373 (verified).

The family side by side

Query counts are worst case per behavior from the shipped configs, counting only calls to the model under attack; attacker and judge calls are separate and, for TAP, far larger.

Methodclass_nameAccess neededAttacker / judgeTarget queriesTest casesTable 6 avgAPI target?
AutoDANAutoDANtarget loss / logitsMistral 7B v0.2 mutator~6,400 scored forwards152.7no
PAIRPAIRchat onlyMixtral 8x7B, both roles60140.7yes
TAPTAPchat onlyMixtral 8x7B, both roles94145.2yes
TAP-TransferTAPnone at replayMixtral attacker, GPT-4 judge + surrogate0148.3yes
GPTFuzzGPTFuzzlocal vLLM + local judgegpt-3.5-turbo mutator, RoBERTa judge1,0001–no
PAP-top5PAPnoneMixtral 8x7B rewriter, no judge0516.6yes
For referenceGCG / DirectRequestgradients / none—500 × 512154.3 / 25.3no / yes

What the family's numbers say

A cost result, not capability. GCG's 54.3 takes a 500-step, 512-candidate search per behavior; AutoDAN's 52.7 takes 100 generations and ~58 paraphraser calls, TAP's 45.2 takes 94 chat turns, PAP's 16.6 five generations. The gradient family's whole edge over the best gradient-free method is 1.6 points, inside the ~7% label noise of the classifier's 93.19% agreement: GCG and AutoDAN are indistinguishable. This family is the closed-model evaluation: the only non-dash cells down the GPT, Claude and Gemini rows of Table 6 are these, plus GCG-Transfer (22.0, 22.3 on the two GPT-4 endpoints vs TAP-Transfer's 54.8, 58.5) and the template baselines. The Claude column stays honest: every method 0.8 to 10.0, the transfer variant worst.

Each method steers on a proxy (a Mixtral rating capped at five output tokens, a RoBERTa classifier blind to the behavior a completion serves, a refusal-substring list) while the real grader is a stronger model run afterwards on a 512-token completion PAIR and TAP saw at most 150 tokens of. Against the floor: DirectRequest averages 25.3 unchanged, Human Jailbreaks 27.3, so PAP-top5's 16.6 is below both; PAIR's 40.7 is ~15 points of lift, TAP's 45.2 ~20. IDs are in the behavior index, the splits in the results read.

Next: Templates, few-shot and humans — the baselines that tell you whether any of this was worth it · back to the map.