HARMBENCH // FIELD MAP
← field map
MEASURING · 02 OF 044 steps · 3 modes · 24 pipeline entries · 512 tokens, fixed

Running it: the four steps

What you type to run HarmBench, and where the runner silently refuses what you asked.
TL;DR. HarmBench is a dataset you download plus a pipeline you run. Behaviors, classifiers and copyright hashes ship; the test cases do not, so you pay in GPU hours, one set per (attack × target model) pair. Four scripts, chained by run_pipeline.py: generate_test_cases.py, an optional merge_test_cases.py, generate_completions.py, evaluate_completions.py, run in slurm, local or local_parallel mode. model_type gates everything: a wrong-type model is skipped with a bare continue, no warning. An OpenAI-compatible endpoint means patching Python, not YAML.
pipeline scripts
4
one of them optional (step 1.5)
pipeline entries
24
+2 commented out as too expensive
models.yaml
35
21 open · 10 closed · 4 multimodal
max_new_tokens
512
deviate and results land elsewhere
precomputed archive
10.4 GB
Zenodo, if you skip step 1 entirely

Part of the benchmark is a download, part a computation. The 510 behaviors, the Llama 2 13B classifier and the 100 MinHash pickles that grade copyright behaviors are downloads. The test cases are not: a test case is what an attack produced against a specific model, one set per (method × target model) pair.

What ships, and what costs you GPU hours

Three ways to get a number, spanning four orders of magnitude: download the precomputed test cases (the 10.4 GB HarmBench Results 1.0 archive on Zenodo, record 10714577, shipped with the 1.0 release) and re-run steps 2 and 3; run only the cheap attacks (no optimisation, minutes); or run the gradient attacks (GPU-months).

Completions and results are keyed by target model, test cases by experiment, so any test-case set points at any model, which is how GCG-Transfer and TAP-Transfer exist as columns in Table 6.

The four steps

Four scripts at the repo root, chained by JSON files: each reads the previous step's file and writes its own. Diagnose a bad run by which file is missing or short.

Step 1 — generate_test_cases.py

The attack runs here: the script loads a red-teaming class from baselines/, initialises it from configs/method_configs/<Method>_config.yaml, hands it a slice of the behavior CSV, and saves the result. Everything expensive happens in that call: GCG's 500 optimisation steps, PAIR's attacker–target conversation, GPTFuzz's mutation loop.

The slice comes from --behavior_start_idx/--behavior_end_idx, or --behavior_ids_subset to fill gaps after a failed job. Single-behavior methods write {save_dir}/test_cases_individual_behaviors/{behavior_id}/test_cases.json, beside a logs.json and a method_config.json recording the config, the pinned transformers/vllm/ray/fastchat versions, and any API token masked to its last four characters.

Step 1.5 — merge_test_cases.py

The base-class contract requires only that after save_test_cases, running merge_test_cases yields a test_cases.json mapping behavior IDs to lists of test cases. Non-parallel methods write that file directly and inherit a merge that is literally pass. Parallel methods scatter output across per-behavior directories, one behavior or chunk per job, and the merge stitches it back.

Three classes implement a merge. RedTeamingMethod is the no-op. SingleBehaviorRedTeamingMethod (GCG, PEZ, GBDA, UAT, AutoPrompt) unions the per-behavior directories, asserting no ID appears twice. EnsembleGCG optimises one adversarial string jointly against every behavior, hence behavior_chunk_size: all_behaviors; its second axis is run_ids, where GCG-Multi and GCG-Transfer declare run_ids: [0,1,2,3,4] and the merge concatenates test_cases_{run_id}.json, so each behavior gets five test cases and an ASR averaged over five optimisations.

SingleBehaviorRedTeamingMethod.merge_test_cases silently skips any per-behavior file that does not exist, so a job that OOM'd or hit its 24-hour SLURM wall drops from the merged file and step 3 computes an ASR over whatever survived. Check the length of test_cases.json against your behavior count before trusting a number.

Step 2 — generate_completions.py

Send each merged test case to the target model, save at completions/{model_name}.json. Decoding is greedy and deterministic (do_sample=False on the HuggingFace path, temperature=0 on the vLLM and API paths), and the budget is max_new_tokens=512 for every experiment in the paper.

Pass anything but 512 and run_pipeline.py redirects completions and results into a {N}_tokens/ subdirectory, so non-standard budgets never mix with the published grid. Since the metric is whether the model produced the harmful artifact, one needing 700 tokens scores as a failure at 512.

Step 3 — evaluate_completions.py

The grading pass reads the behavior CSV (tags and context strings) and the completions, then dispatches per behavior: the hash_check tag routes to MinHash, everything else to the Llama 2 13B classifier under vLLM at temperature=0.0, max_tokens=1, a single-token label. Output is results/{model_name}.json, one 0/1 label per test case.

Step 3 re-clips every completion to 512 tokens with the classifier's own tokenizer (truncation_side="right"). ASR is a mean of means (rate per behavior, then averaged), which diverges from a flat mean only at GCG-Multi's five test cases per behavior. A behavior ID in the completions but absent from the CSV you passed is skipped with a printed warning, which is how a 400-behavior completions file gets graded against the 80-behavior validation split.

Chaining it: run_pipeline.py

scripts/run_pipeline.py reads run_pipeline.yaml, expands each method into per-model experiments, sizes each job's GPU request, and submits the chain:

python ./scripts/run_pipeline.py \
  --methods DirectRequest,PAIR,TAP \
  --models your_model \
  --behaviors_path ./data/behavior_datasets/harmbench_behaviors_text_val.csv \
  --step all --mode local_parallel \
  --cls_path cais/HarmBench-Llama-2-13b-cls

--step takes all, 1, 1.5, 2, 3 or 2_and_3, the last being the Zenodo workhorse: skip generation, score an existing test-case set against a new model. --mode takes three values:

Results are keyed by the class name, not the pipeline entry name: everything saves under {base_save_dir}/{class_name}/{experiment_name}/, so GCG-Transfer output appears under results/EnsembleGCG/llama2_7b_vicuna_7b_llama2_13b_vicuna_13b_multibehavior_1000steps/ and no GCG-Transfer directory exists on disk.

Adding a target model — and the field that gates everything

Adding your own model means one entry in configs/model_configs/models.yaml. For anything transformers-compatible that is all it is: a model block passed to load_model_and_tokenizer (a model_name_or_path plus overrides such as dtype, trust_remote_code, a chat_template name, explicit pad/eos tokens), a num_gpus, and model_type. An API model carries the model string and a token instead of num_gpus. The file has 35 entries: 21 open_source, 10 closed_source, 3 open_source_multimodal, 1 closed_source_multimodal, with two warts not to copy: mistral-medium misspells its block as mode:, and gpt-4-vision-preview is typed closed_source while the separate gpt4v entry is closed_source_multimodal.

model_type takes one of four values (open_source, closed_source, open_source_multimodal, closed_source_multimodal), and every run_pipeline.yaml entry declares an allowed_target_model_types list. A model whose type is not in the list is dropped with a bare continue, no warning or error. Ask for --methods GCG --models gpt-4-0613 and the command exits cleanly having submitted nothing.

The GPT-4 and Claude rows in Table 6 are mostly dashes because GCG and its white-box relatives need gradients through the target and the config refuses to schedule the job. (Filter out every requested model in local_parallel mode and the skip surfaces as a bare NoneType traceback in step 2.)

A step-1 job's GPU request is base_num_gpus plus the target's num_gpus, but only when the entry's experiment_name_template contains <model_name>; fixed-name transfer entries get base_num_gpus alone, so GCG-Transfer asks for 4 (its four attack models). Job count is ceil(num_behaviors / behavior_chunk_size) × len(run_ids), from 1 to 320. Both read off the table below, verbatim from run_pipeline.yaml.

Pipeline entryclass_nameexperiment_name_templatechunkbase GPUsrun_idsallowed_target_model_types
GCGGCG<model_name>10–open_source
GCG_custom_targetsGCG<model_name>_custom_targets50–open_source
GCG-MultiEnsembleGCG<model_name>all00–4open_source
GCG-TransferEnsembleGCGllama2_7b_vicuna_7b_llama2_13b_vicuna_13b_multibehavior_1000stepsall40–4open_source, closed_source
AutoPromptAutoPrompt<model_name>10–open_source
PEZPEZ<model_name>50–open_source
GBDAGBDA<model_name>50–open_source
UATUAT<model_name>50–open_source
AutoDANAutoDAN<model_name>52–open_source
PAIRPAIR<model_name>52–open_source, closed_source
TAPTAP<model_name>52–open_source, closed_source
TAP-TransferTAPgpt-4-0613_judge_gpt-4-1106-preview_target52–open_source, closed_source
GPTFuzzGPTFuzz<model_name>51–open_source
PAP-top5PAPtop_552–open_source, closed_source
FewShot (SFS)FewShot<model_name>52–open_source
ZeroShotZeroShotmixtral_attacker_llm1002–open_source, closed_source
HumanJailbreaksHumanJailbreaksrandom_subset_5all0–open_source, closed_source
ArtPromptArtPrompt<model_name>50–open_source, closed_source
DirectRequestDirectRequestdefaultall0–open_source, closed_source
MultiModalPGDMultiModalPGD<model_name>51–open_source_multimodal
MultiModalPGDPatchMultiModalPGDPatch<model_name>51–open_source_multimodal
MultiModalPGDBlankImageMultiModalPGD<model_name>51–open_source_multimodal
MultiModalRenderTextMultiModalRenderText<model_name>51–open_source_multimodal, closed_source_multimodal
MultiModalDirectRequestMultiModalDirectRequestdefault51–open_source_multimodal, closed_source_multimodal
24 active entriestwo more are present but commented out

The five MultiModal* rows are the only ones accepting the multimodal types, enforcing the multimodal functional type: a vision model in models.yaml is invisible to the other nineteen methods. MultiModalPGDBlankImage is MultiModalPGD under a different experiment config, not a distinct class.

Two entries sit commented out as “expensive to run on closed-source models”: HumanJailbreaks-all (the full jailbreak library, not the random_subset_5 the paper reports) and PAP-avg (the full 40-strategy persuasion taxonomy, not the top 5), each roughly an order of magnitude more API calls per behavior. So the table's Human and PAP-top5 columns are the cheap configurations: Human Jailbreaks at 27.3 and PAP-top5 at 16.6.

The API-only snag: base_url

A target behind an HTTP API routes through api_models.py, which defines five classes (GPT, GPTV, Claude, Gemini, Mistral) and one dispatch function that picks between them by pattern-matching on the model name:

def api_models_map(model_name_or_path=None, token=None, **kwargs):
    if 'gpt-' in model_name_or_path:
        if 'vision' in model_name_or_path:
            return GPTV(model_name_or_path, token)
        else:
            return GPT(model_name_or_path, token)
    elif 'claude-' in model_name_or_path:
        return Claude(model_name_or_path, token)
    elif 'gemini-' in model_name_or_path:
        return Gemini(model_name_or_path, token)
    elif re.match(r'mistral-(tiny|small|medium|large)$', model_name_or_path):
        return Mistral(model_name_or_path, token)
    return None

No provider field, no endpoint field: a gpt-something name gets the OpenAI client, constructed like this:

class GPT():
    API_TIMEOUT = 60

    def __init__(self, model_name, api_key):
        self.model_name = model_name
        self.client = OpenAI(api_key=api_key, timeout=self.API_TIMEOUT)

No base_url: the constructor takes only a key and a timeout, so it only ever talks to api.openai.com. Much inference is served over the OpenAI wire format by something that is not OpenAI (a self-hosted vLLM or SGLang server, a vendor endpoint, a router), which HarmBench cannot reach by configuration. Add a class to api_models.py, or thread a base_url through the constructor and surface it in the model config: a code change, not a YAML edit.

api_models_map returns None when nothing matches, and its only caller, load_generation_function in generate_completions.py, falls through to the local-weights branch. So a closed_source entry whose name misses all four patterns hands your model string to a HuggingFace loader and fails with a missing-repository error: a hub 404 while wiring up a new endpoint means checking the dispatch function first.

The config docs say new APIs are incorporated in load_generation_function, api_models.py and eval_utils.py; it is the most common reason a run does not start.

The practical shape of a run

Cheapest: DirectRequest has no hyperparameters, sends the behavior as-is, and produces one test-case set shared by every model (experiment_name_template: default, behavior_chunk_size: all_behaviors, zero GPUs). HumanJailbreaks and ZeroShot are nearly as cheap.

Priciest: GCG is 500 optimisation steps at a search width of 512 candidates, one behavior per job (behavior_chunk_size: 1), per target: on the 320-behavior test split, 320 step-1 jobs per model, each a quarter-million forward passes. GCG-Multi multiplies that by five run IDs.

So the first move is not --methods all: run the 80-behavior validation split with DirectRequest end to end, and confirm a plausible ASR before committing GPU time.

python ./scripts/run_pipeline.py \
  --methods DirectRequest \
  --models your_model \
  --behaviors_path ./data/behavior_datasets/harmbench_behaviors_text_val.csv \
  --step all --mode local_parallel

harmbench_behaviors_text_val.csv is the right file: the 80 held-out behaviors, spanning all seven semantic categories and all three text functional types, the split to iterate against so the test classifier is never optimised against. It runs in minutes, yet a dropped ContextString column or wrong chat template shows up at once as an implausible number.

That exercises all four steps at once, each failing differently and cheap to fix alone. Then add one black-box attack, PAIR or TAP (both accept closed-source targets, two GPUs for the attacker), before anything gradient-based touches the cluster.

Read the number against the grid: DirectRequest averages 25.3 across the paper's 33 targets, but 65.8 on Zephyr 7B and 0.8 on Llama 2 7B Chat, so a first run near zero is a safety-trained model, not a broken pipeline. HarmBench ships no benign set, so a model that refuses everything scores 0% and looks flawless.

Next: What the paper found · back to the map.