Running it: the four steps
Part of the benchmark is a download, part a computation. The 510 behaviors, the Llama 2 13B classifier and the 100 MinHash pickles that grade copyright behaviors are downloads. The test cases are not: a test case is what an attack produced against a specific model, one set per (method × target model) pair.
What ships, and what costs you GPU hours
Three ways to get a number, spanning four orders of magnitude: download the precomputed test cases (the 10.4 GB HarmBench Results 1.0 archive on Zenodo, record 10714577, shipped with the 1.0 release) and re-run steps 2 and 3; run only the cheap attacks (no optimisation, minutes); or run the gradient attacks (GPU-months).
Completions and results are keyed by target model, test cases by experiment, so any test-case set points at any model, which is how GCG-Transfer and TAP-Transfer exist as columns in Table 6.
The four steps
Four scripts at the repo root, chained by JSON files: each reads the previous step's file and writes its own. Diagnose a bad run by which file is missing or short.
Step 1 — generate_test_cases.py
The attack runs here: the script loads a red-teaming class from baselines/, initialises it from configs/method_configs/<Method>_config.yaml, hands it a slice of the behavior CSV, and saves the result. Everything expensive happens in that call: GCG's 500 optimisation steps, PAIR's attacker–target conversation, GPTFuzz's mutation loop.
The slice comes from --behavior_start_idx/--behavior_end_idx, or --behavior_ids_subset to fill gaps after a failed job. Single-behavior methods write {save_dir}/test_cases_individual_behaviors/{behavior_id}/test_cases.json, beside a logs.json and a method_config.json recording the config, the pinned transformers/vllm/ray/fastchat versions, and any API token masked to its last four characters.
Step 1.5 — merge_test_cases.py
The base-class contract requires only that after save_test_cases, running merge_test_cases yields a test_cases.json mapping behavior IDs to lists of test cases. Non-parallel methods write that file directly and inherit a merge that is literally pass. Parallel methods scatter output across per-behavior directories, one behavior or chunk per job, and the merge stitches it back.
Three classes implement a merge. RedTeamingMethod is the no-op. SingleBehaviorRedTeamingMethod (GCG, PEZ, GBDA, UAT, AutoPrompt) unions the per-behavior directories, asserting no ID appears twice. EnsembleGCG optimises one adversarial string jointly against every behavior, hence behavior_chunk_size: all_behaviors; its second axis is run_ids, where GCG-Multi and GCG-Transfer declare run_ids: [0,1,2,3,4] and the merge concatenates test_cases_{run_id}.json, so each behavior gets five test cases and an ASR averaged over five optimisations.
Step 2 — generate_completions.py
Send each merged test case to the target model, save at completions/{model_name}.json. Decoding is greedy and deterministic (do_sample=False on the HuggingFace path, temperature=0 on the vLLM and API paths), and the budget is max_new_tokens=512 for every experiment in the paper.
Pass anything but 512 and run_pipeline.py redirects completions and results into a {N}_tokens/ subdirectory, so non-standard budgets never mix with the published grid. Since the metric is whether the model produced the harmful artifact, one needing 700 tokens scores as a failure at 512.
Step 3 — evaluate_completions.py
The grading pass reads the behavior CSV (tags and context strings) and the completions, then dispatches per behavior: the hash_check tag routes to MinHash, everything else to the Llama 2 13B classifier under vLLM at temperature=0.0, max_tokens=1, a single-token label. Output is results/{model_name}.json, one 0/1 label per test case.
Step 3 re-clips every completion to 512 tokens with the classifier's own tokenizer (truncation_side="right"). ASR is a mean of means (rate per behavior, then averaged), which diverges from a flat mean only at GCG-Multi's five test cases per behavior. A behavior ID in the completions but absent from the CSV you passed is skipped with a printed warning, which is how a 400-behavior completions file gets graded against the 80-behavior validation split.
Chaining it: run_pipeline.py
scripts/run_pipeline.py reads run_pipeline.yaml, expands each method into per-model experiments, sizes each job's GPU request, and submits the chain:
python ./scripts/run_pipeline.py \
--methods DirectRequest,PAIR,TAP \
--models your_model \
--behaviors_path ./data/behavior_datasets/harmbench_behaviors_text_val.csv \
--step all --mode local_parallel \
--cls_path cais/HarmBench-Llama-2-13b-cls
--step takes all, 1, 1.5, 2, 3 or 2_and_3, the last being the Zenodo workhorse: skip generation, score an existing test-case set against a new model. --mode takes three values:
- slurm: the intended path. Each chunk becomes an sbatch job, steps wired with --dependency=afterok: on the previous step's job IDs. Hard-coded wall times: 24 hours for step 1, ten minutes for the merge, five hours each for completions and evaluation.
- local: sequential on this machine. Steps 1.5, 2 and 3 run under capture_output=True with no live output, so a long silent step 2 is not necessarily hung.
- local_parallel: Ray across one box's GPUs, marked experimental. It registers a custom visible_gpus resource sized to torch.cuda.device_count() and sets CUDA_VISIBLE_DEVICES per subprocess, the right mode for a single 8×A100 machine.
Results are keyed by the class name, not the pipeline entry name: everything saves under {base_save_dir}/{class_name}/{experiment_name}/, so GCG-Transfer output appears under results/EnsembleGCG/llama2_7b_vicuna_7b_llama2_13b_vicuna_13b_multibehavior_1000steps/ and no GCG-Transfer directory exists on disk.
Adding a target model — and the field that gates everything
Adding your own model means one entry in configs/model_configs/models.yaml. For anything transformers-compatible that is all it is: a model block passed to load_model_and_tokenizer (a model_name_or_path plus overrides such as dtype, trust_remote_code, a chat_template name, explicit pad/eos tokens), a num_gpus, and model_type. An API model carries the model string and a token instead of num_gpus. The file has 35 entries: 21 open_source, 10 closed_source, 3 open_source_multimodal, 1 closed_source_multimodal, with two warts not to copy: mistral-medium misspells its block as mode:, and gpt-4-vision-preview is typed closed_source while the separate gpt4v entry is closed_source_multimodal.
model_type takes one of four values (open_source, closed_source, open_source_multimodal, closed_source_multimodal), and every run_pipeline.yaml entry declares an allowed_target_model_types list. A model whose type is not in the list is dropped with a bare continue, no warning or error. Ask for --methods GCG --models gpt-4-0613 and the command exits cleanly having submitted nothing.
The GPT-4 and Claude rows in Table 6 are mostly dashes because GCG and its white-box relatives need gradients through the target and the config refuses to schedule the job. (Filter out every requested model in local_parallel mode and the skip surfaces as a bare NoneType traceback in step 2.)
A step-1 job's GPU request is base_num_gpus plus the target's num_gpus, but only when the entry's experiment_name_template contains <model_name>; fixed-name transfer entries get base_num_gpus alone, so GCG-Transfer asks for 4 (its four attack models). Job count is ceil(num_behaviors / behavior_chunk_size) × len(run_ids), from 1 to 320. Both read off the table below, verbatim from run_pipeline.yaml.
| Pipeline entry | class_name | experiment_name_template | chunk | base GPUs | run_ids | allowed_target_model_types |
|---|---|---|---|---|---|---|
| GCG | GCG | <model_name> | 1 | 0 | – | open_source |
| GCG_custom_targets | GCG | <model_name>_custom_targets | 5 | 0 | – | open_source |
| GCG-Multi | EnsembleGCG | <model_name> | all | 0 | 0–4 | open_source |
| GCG-Transfer | EnsembleGCG | llama2_7b_vicuna_7b_llama2_13b_vicuna_13b_multibehavior_1000steps | all | 4 | 0–4 | open_source, closed_source |
| AutoPrompt | AutoPrompt | <model_name> | 1 | 0 | – | open_source |
| PEZ | PEZ | <model_name> | 5 | 0 | – | open_source |
| GBDA | GBDA | <model_name> | 5 | 0 | – | open_source |
| UAT | UAT | <model_name> | 5 | 0 | – | open_source |
| AutoDAN | AutoDAN | <model_name> | 5 | 2 | – | open_source |
| PAIR | PAIR | <model_name> | 5 | 2 | – | open_source, closed_source |
| TAP | TAP | <model_name> | 5 | 2 | – | open_source, closed_source |
| TAP-Transfer | TAP | gpt-4-0613_judge_gpt-4-1106-preview_target | 5 | 2 | – | open_source, closed_source |
| GPTFuzz | GPTFuzz | <model_name> | 5 | 1 | – | open_source |
| PAP-top5 | PAP | top_5 | 5 | 2 | – | open_source, closed_source |
| FewShot (SFS) | FewShot | <model_name> | 5 | 2 | – | open_source |
| ZeroShot | ZeroShot | mixtral_attacker_llm | 100 | 2 | – | open_source, closed_source |
| HumanJailbreaks | HumanJailbreaks | random_subset_5 | all | 0 | – | open_source, closed_source |
| ArtPrompt | ArtPrompt | <model_name> | 5 | 0 | – | open_source, closed_source |
| DirectRequest | DirectRequest | default | all | 0 | – | open_source, closed_source |
| MultiModalPGD | MultiModalPGD | <model_name> | 5 | 1 | – | open_source_multimodal |
| MultiModalPGDPatch | MultiModalPGDPatch | <model_name> | 5 | 1 | – | open_source_multimodal |
| MultiModalPGDBlankImage | MultiModalPGD | <model_name> | 5 | 1 | – | open_source_multimodal |
| MultiModalRenderText | MultiModalRenderText | <model_name> | 5 | 1 | – | open_source_multimodal, closed_source_multimodal |
| MultiModalDirectRequest | MultiModalDirectRequest | default | 5 | 1 | – | open_source_multimodal, closed_source_multimodal |
| 24 active entries | two more are present but commented out | |||||
The five MultiModal* rows are the only ones accepting the multimodal types, enforcing the multimodal functional type: a vision model in models.yaml is invisible to the other nineteen methods. MultiModalPGDBlankImage is MultiModalPGD under a different experiment config, not a distinct class.
Two entries sit commented out as “expensive to run on closed-source models”: HumanJailbreaks-all (the full jailbreak library, not the random_subset_5 the paper reports) and PAP-avg (the full 40-strategy persuasion taxonomy, not the top 5), each roughly an order of magnitude more API calls per behavior. So the table's Human and PAP-top5 columns are the cheap configurations: Human Jailbreaks at 27.3 and PAP-top5 at 16.6.
The API-only snag: base_url
A target behind an HTTP API routes through api_models.py, which defines five classes (GPT, GPTV, Claude, Gemini, Mistral) and one dispatch function that picks between them by pattern-matching on the model name:
def api_models_map(model_name_or_path=None, token=None, **kwargs):
if 'gpt-' in model_name_or_path:
if 'vision' in model_name_or_path:
return GPTV(model_name_or_path, token)
else:
return GPT(model_name_or_path, token)
elif 'claude-' in model_name_or_path:
return Claude(model_name_or_path, token)
elif 'gemini-' in model_name_or_path:
return Gemini(model_name_or_path, token)
elif re.match(r'mistral-(tiny|small|medium|large)$', model_name_or_path):
return Mistral(model_name_or_path, token)
return None
No provider field, no endpoint field: a gpt-something name gets the OpenAI client, constructed like this:
class GPT():
API_TIMEOUT = 60
def __init__(self, model_name, api_key):
self.model_name = model_name
self.client = OpenAI(api_key=api_key, timeout=self.API_TIMEOUT)
No base_url: the constructor takes only a key and a timeout, so it only ever talks to api.openai.com. Much inference is served over the OpenAI wire format by something that is not OpenAI (a self-hosted vLLM or SGLang server, a vendor endpoint, a router), which HarmBench cannot reach by configuration. Add a class to api_models.py, or thread a base_url through the constructor and surface it in the model config: a code change, not a YAML edit.
api_models_map returns None when nothing matches, and its only caller, load_generation_function in generate_completions.py, falls through to the local-weights branch. So a closed_source entry whose name misses all four patterns hands your model string to a HuggingFace loader and fails with a missing-repository error: a hub 404 while wiring up a new endpoint means checking the dispatch function first.
The config docs say new APIs are incorporated in load_generation_function, api_models.py and eval_utils.py; it is the most common reason a run does not start.
The practical shape of a run
Cheapest: DirectRequest has no hyperparameters, sends the behavior as-is, and produces one test-case set shared by every model (experiment_name_template: default, behavior_chunk_size: all_behaviors, zero GPUs). HumanJailbreaks and ZeroShot are nearly as cheap.
Priciest: GCG is 500 optimisation steps at a search width of 512 candidates, one behavior per job (behavior_chunk_size: 1), per target: on the 320-behavior test split, 320 step-1 jobs per model, each a quarter-million forward passes. GCG-Multi multiplies that by five run IDs.
So the first move is not --methods all: run the 80-behavior validation split with DirectRequest end to end, and confirm a plausible ASR before committing GPU time.
python ./scripts/run_pipeline.py \
--methods DirectRequest \
--models your_model \
--behaviors_path ./data/behavior_datasets/harmbench_behaviors_text_val.csv \
--step all --mode local_parallel
harmbench_behaviors_text_val.csv is the right file: the 80 held-out behaviors, spanning all seven semantic categories and all three text functional types, the split to iterate against so the test classifier is never optimised against. It runs in minutes, yet a dropped ContextString column or wrong chat template shows up at once as an implausible number.
That exercises all four steps at once, each failing differently and cheap to fix alone. Then add one black-box attack, PAIR or TAP (both accept closed-source targets, two GPUs for the attacker), before anything gradient-based touches the cluster.
Read the number against the grid: DirectRequest averages 25.3 across the paper's 33 targets, but 65.8 on Zephyr 7B and 0.8 on Llama 2 7B Chat, so a first run near zero is a safety-trained model, not a broken pipeline. HarmBench ships no benign set, so a model that refuses everything scores 0% and looks flawless.