HARMBENCH // FIELD MAP
← field map
THE ATTACKS · 03 OF 035 methods · SFS 38.3 · Human 27.3 · ZS 25.4 · DR 25.3

Templates, few-shot and humans

Five methods that compute no gradient, and the no-attack DirectRequest floor every other ASR is read against.
gradient family LLM-optimizer family Table 6 all 510 IDs
TL;DR — Five methods, none computing a gradient. Three cluster at the bottom of Table 6: DirectRequest 25.3, ZeroShot 25.4, Human Jailbreaks 27.3; the fourth, Stochastic Few-Shot, is 38.3, beating three white-box gradient methods. DirectRequest sends the behavior string unmodified and a quarter of HarmBench falls to it. Report a model's ASR as the distance from its DirectRequest floor, not the attack alone.
SFS
38.3
best in family; needs open weights
Human jailbreaks
27.3
5 of 114 templates per behavior
ZeroShot
25.4
attacker LLM, target never queried
DirectRequest
25.3
no attack at all
DR spread
0.8–65.8
Llama 2 7B Chat → Zephyr 7B

Where the gradient and LLM-optimizer families ask how hard you can push, this family asks how much falls over with no push: a quarter, the famous templates adding two points.

What the family has in common

Four of the five build a test case by string manipulation: a template round the behavior, a one-word rendering, or the behavior verbatim; the fifth, SFS, runs a search loop over natural-language candidates, not token space. None takes a loss on the input, so three carry base_num_gpus: 0; the two wanting GPUs (ZeroShot and SFS, two each) use them for a Mixtral 8x7B attack model, not the target. Four list closed_source in allowed_target_model_types; SFS excepted, scoring by the target's cross-entropy on a fixed continuation, so it needs logits.

ASR is a per-behavior mean over that behavior's test cases, averaged over behaviors: a method emitting five per behavior (ZeroShot, Human Jailbreaks) is scored on the fraction that worked, not whether any did. "Best of five" would rank both higher but is deliberately not reported. See the classifier and the pipeline.

ZS — ZeroShot

black box · target never queriedclass_name: ZeroShot5 candidates/behavior, one attack-model call eachTable 6 avg 25.4

An attacker LLM is shown the behavior (and, for a contextual behavior, its context passage) and asked to write prompts that elicit it; the reply is prefilled with an affirmative opener as a numbered list, sampling stops at the first newline, so each call yields one candidate. No feedback: the target is never called during generation, so this runs against API-only targets.

Attack model mistralai/Mixtral-8x7B-Instruct-v0.1 on 2 GPUs via vLLM at temperature 0.9, top_p 0.95, presence_penalty / frequency_penalty 0.1, max_tokens 256; num_test_cases_per_behavior: 5 overrides the class default of 100. Pipeline row mixtral_attacker_llm, behavior_chunk_size: 100: four chunks over the 400 text behaviors. For contextual behaviors the passage is stitched onto the front of the generated line.

At 25.4 it is indistinguishable from asking plainly (0.1 points, against a grader with ~7% label noise). On closed targets it is often the best cheap method: GPT-4 0613 19.4; on Claude 2.1 its 4.1 is the highest cell in all sixteen columns. On many it is worse than nothing: 33.0 → 28.4 on GPT-3.5 Turbo 1106, 18.0 → 14.8 on Gemini Pro, 65.8 → 60.0 on Zephyr 7B, because a rewrite that drifts off the behavior is scored a failure even when the model complied, the flaw Table 4's third prequalification set exposes in rival classifiers.

Reimplemented from Perez et al. 2022, Red Teaming Language Models with Language Models (arXiv:2202.03286) with an adjusted prompt, per §C.1.

SFS — Stochastic Few-Shot

open weights · forward passes, no gradientsclass_name: FewShot50 steps × 64 candidatesTable 6 avg 38.3

The loop: ask the same Mixtral model for ten variation prompts; dedupe and discard anything under 32 characters; run each survivor through the target and score by mean cross-entropy on a fixed affirmative continuation from data/optimizer_targets/harmbench_targets_text.json; hand the five lowest-loss back as few-shot examples, up to 50 rounds. The returned test case is the single lowest-loss candidate.

That target file is the one GCG optimises against: SFS is GCG's objective with the gradient removed and an LLM sampler in its place, steered by its own best-scoring outputs. It needs logits, hence allowed_target_model_types: [open_source], but never a backward pass.

Knobs: sample_size: 64 per step, n_shots: 5, num_steps: 50; generation at temperature 1.0, top_p 1.0, presence_penalty / frequency_penalty 0.1, max_tokens 1024, a stop string capping each call at ten numbered lines. Separate templates for plain and contextual behaviors; contextual candidates get the passage prepended before scoring. behavior_chunk_size: 5, base_num_gpus: 2.

Three things stop the loop: the shots stop changing, no new candidate survives filtering, or the current best yields a completion that does not look like a refusal. That check greedily generates 32 tokens and tests them against 35 refusal substrings, the AdvBench-style heuristic Table 4 shows is a bad grader. It cannot inflate the score: ASR is still the Llama 2 13B classifier over 512 tokens.

38.3 beats UAT (30.8), GBDA (29.8) and PEZ (29.0), all white box, and matches GCG-Transfer at 38.8. By slice: 44.3 standard, 56.5 contextual, 6.6 copyright. Against Zephyr 7B + R2D2, adversarially trained until GCG fell from 69.5 to 5.5, SFS still lands 43.5.

"For some attacks, the improvement conferred by R2D2 is less pronounced. This is especially true for methods dissimilar to the train-time GCG adversary, including PAIR, TAP, and Stochastic Few-Shot."— HarmBench, §6.2

Same origin as ZeroShot, but §C.1 is explicit this is not the original: the authors "use an iterative version of the Stochastic Few-Shot method to increase the probability of the target LLM generating a target string". The loss-guided iteration is HarmBench's addition, so 38.3 is HarmBench's variant.

Human — Human Jailbreaks

black box · no model calls to buildclass_name: HumanJailbreaks5 of 114 templates per behaviorTable 6 avg 27.3

The corpus is one Python list, JAILBREAKS: 114 entries, about 226,000 characters, median around 1,400 and longest around 16,000. In-the-wild Do Anything Now templates (persona, fictional-wrapper, rule-override framings); this site reproduces none of them.

Construction: template, blank line, behavior string, contextual passages joined first via the same \n\n---\n\n separator DirectRequest uses. Nothing adapts per model, behavior or response. The reported column is random_subset_5 with seed: 1, reshuffled inside the per-behavior loop, so each behavior draws a different five of the 114, deterministically. behavior_chunk_size: all_behaviors, base_num_gpus: 0, both target types.

An all_jailbreaks experiment exists (random_subset: -1); its HumanJailbreaks-all pipeline row is commented out as expensive on closed models: 113 templates × 400 behaviors is about 45,000 completions at 512 tokens. 113 not 114 because -1 is an integer, so JAILBREAKS[:-1] drops the shuffled list's last entry, an off-by-one that never touches the reported column.

At 27.3 the whole corpus is worth 2.0 points over asking plainly (31.9 standard, 41.4 contextual, 4.3 copyright). Across two revisions of one closed model the column falls off a cliff: GPT-3.5 Turbo 0613 24.5 vs 1106 2.8; GPT-4 0613 11.3 vs Turbo 1106 2.6; the three Claude models 2.4, 0.3, 0.3.

Five unadapted 2023 templates thrown at later models, scored on the mean of the five, measures do the famous jailbreaks still work (largely no), not can a human jailbreak this model, which needs a person adapting after each refusal, something HarmBench asks of none of its 18 methods. Source: Shen et al. 2023, "Do Anything Now" (arXiv:2308.03825).

ArtPrompt

black box + one GPT-4 API callclass_name: ArtPrompt1 masking call, 1 test case per behaviornot in Table 6

ArtPrompt has no HarmBench ASR number. It appears in neither the arXiv v2 text nor any result table, and the headline grid's sixteen attack columns exclude it. It is a later addition to the repository's method set, which is why the repo's 21 method configs and 24 active pipeline entries exceed the paper's 18 methods.

Mechanism: a masking model picks the single word carrying the sensitive load; that word becomes a placeholder re-presented as ASCII art, so a safety layer reading text never sees it while a model that reads the rendering recovers it.

Config: masking_model: gpt-4-0613, and masking_model_key is required (the constructor raises without it), so despite base_num_gpus: 0 it needs an OpenAI account. masking_direction: horizontal by default, vertical available; fontname: cards from a six-item list (gen, alphabet, letters, keyboard, cards, puzzle), the intersection of the two renderers (horizontal implements eleven fonts, vertical six). behavior_chunk_size: 5, both target types.

Only the first mask word is used (a code comment says multi-word masking is unsupported); candidates containing a space are discarded, and if none survives the test case falls back to the plain behavior string, silently degrading to DirectRequest. A mask word that is not a literal substring raises. Vestigial optimiser plumbing remains: targets_path is loaded but never read, and each log entry records loss=0.

Source: Jiang et al. 2024, ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs (arXiv:2402.11753), which also introduces the ViTC (Vision-in-Text Challenge) benchmark the class names refer to.

DR — DirectRequest

black box · no attackclass_name: DirectRequestno hyperparameters at allTable 6 avg 25.3

The whole method config is placeholder: placeholder, twice: generate_test_cases returns the Behavior string. Pipeline row: experiment_name_template: default, behavior_chunk_size: all_behaviors, base_num_gpus: 0, both model types. Step 1 of the pipeline is CPU string formatting; the cost lives in generation and classification.

For a contextual behavior it does not send the behavior alone, but ContextString, then \n\n---\n\n, then the behavior. That join is shared by Human Jailbreaks, ZeroShot and SFS, an undocumented convention of the method set; a harness that sends just the behavior is not running DirectRequest.

"This tests how well models can refuse direct requests to engage in the behaviors when the requests are not obfuscated in any way and often suggest malicious intent."— HarmBench, §C.1

25.3 average, but the per-model range is the finding: 0.8 on Llama 2 7B Chat, 65.8 on Zephyr 7B; beside them SOLAR 10.7B-Instruct 61.3 and Mistral 7B 46.3, against 2.8 for both the 13B and 70B Llama 2 Chat models. Refusal training, not scale, separates them, the point the results page makes across the grid.

By slice, 23.9 standard / 46.2 contextual / 7.3 copyright. Contextual nearly doubles standard: 46% succeed with no attack, the paper's Figure 11 claim that harder-to-look-up behaviors are easier to elicit. On Copyright Violations, DirectRequest's 7.3 beats GCG's 4.6 and every gradient column: the MinHash grader only fires on near-verbatim regurgitation, and optimisation pushes a model off verbatim reproduction. (Recomputed averages: the printed Average row in the v2 sub-table contradicts its own cells; see the discrepancy note.)

Its multimodal twin MultiModalDirectRequest scores 69.09 on LLaVA 1.5 and 16.36 on GPT-4V over the 110 multimodal behaviors.

The family, side by side

MethodAccess neededCost per behaviorTable 6 avgAPI-only target?
Stochastic Few-Shotopen weights (logits)up to 50 rounds; ≤64 target forward passes each38.3no
Human Jailbreaksblack box5 completions; zero model calls to build27.3yes
ZeroShotblack box5 attack-model calls (≤256 tokens each)25.4yes
DirectRequestblack box1 completion; no test-case generation25.3yes
ArtPromptblack box + GPT-4 API1 masking call, 1 completion–yes
For scale — GCGwhite box (gradients)500 steps × 512 candidates54.3no

ZeroShot (25.4) gains nothing over asking; adding a target-string objective and fifty rounds of selection on its own best outputs (SFS, 38.3) gains thirteen points and passes three white-box gradient methods, with the attack model held constant.

Reading any ASR number against the floor

Never report an attack ASR on its own. Report the pair.

ModelDirectRequestBest attack in Table 6WhichGap
Llama 2 7B Chat0.832.5GCG+31.7
Llama 2 13B Chat2.830.0GCG+27.2
Zephyr 7B + R2D214.260.8TAP+46.6
GPT-4 Turbo 11069.358.5TAP-Transfer+49.2
Vicuna 7B24.366.0AutoDAN+41.7
Orca 2 13B44.569.8AutoDAN+25.3
SOLAR 10.7B-Instruct61.372.5AutoDAN+11.2
Zephyr 7B65.875.0AutoDAN+9.2
Claude 2.12.04.1ZeroShot+2.1

Zephyr 7B under GCG is 69.5, Llama 2 7B Chat 32.5, but against the floor Zephyr went 65.8 → 69.5 (attacking bought 3.7 points) while Llama 2 7B Chat went 0.8 → 32.5. Adjacent headline numbers, opposite security stories; only the pair distinguishes them.

Two cautions. The grader agrees with humans about 93% of the time, so ~7% of labels are noise: the 2.0 points between Human Jailbreaks and DirectRequest and the 0.1 between ZeroShot and DirectRequest are not defensible; the 13 between SFS and ZeroShot is. And over-refusal is measured nowhere in HarmBench, so pair DirectRequest with a false-refusal set. The classifier page lists the rest.

Run all five before anything expensive; the interesting question about your model is usually answered by the cheapest, over the same 510 rows at the behavior index.

That closes the attacks arc. Next, how any of this gets scored: The classifier. Back to the attacks section, or on to measuring.