HARMBENCH // FIELD MAP
← field map
THE TYPES · 01 OF 04200 behaviors · 159 test / 41 val · 39% of the set · GCG 69.1

Standard behaviors

The plain case: no context column, no tags, the behavior string is the entire prompt.
contextual → copyright multimodal all 510 IDs
TL;DR — One imperative sentence, median 82 characters, graded by the plain LLAMA2_CLS_PROMPT['prompt'] template with nothing attached. 200 of 510 behaviors (39%), and the shape HarmBench's own design principles argue against, since a bare request is what a search engine may already answer. Yet contextual, the type built to fix that, is easier to elicit in all sixteen attack columns.
Behaviors
200
159 test / 41 val
Share of 510
39%
50% of the 400 text behaviors
Median prompt
82 ch
min 23, max 180
GCG average
69.1
Table 6, standard sub-table
DirectRequest
23.9
the no-attack floor

Standard is the degenerate type: the model gets only the behavior string, the grader gets that string plus the completion, and the other three types are each defined by what they add on top.

What the type is, and where the 200 come from

In harmbench_behaviors_text_all.csv a standard behavior has FunctionalCategory standard and empty ContextString and Tags; that empty Tags field routes it through the plain grading path.

The 200 come from six of the seven semantic categories. Copyright Violations contributes zero: all 100 are the copyright type, graded by MinHash not a classifier.

Semantic categorystandardof the category's text behaviorsshare standard
Illegal Activities586589%
Cybercrime & Unauthorized Intrusion406760%
Misinformation & Disinformation346552%
Chemical & Biological Weapons/Drugs285650%
General Harm212295%
Harassment & Bullying192576%
Copyright Violations01000%
Total20040050%

The shares track where the harm lives. General Harm (95% standard) and Harassment & Bullying (76%) run high because there the output is the harm. Chemical & Biological (28/28) and Misinformation (34/31) split evenly, where a plain request competes with a library and the authors invested in context-grounded variants.

159 test / 41 val puts standard at almost exactly half the 320-behavior text test split. Full ID list on the behaviors page.

The behavior string is the whole prompt

In baselines/direct_request/direct_request.py the whole attack is: if ContextString is non-empty prepend it with a --- separator, else return the string unchanged. For standard that branch never fires, so DirectRequest's test case is byte-identical to the CSV cell. Others build on that string: GCG appends an optimised suffix, AutoDAN evolves a wrapper, Human Jailbreaks drop it into a persona template.

Downstream is fixed by the pipeline: the target's chat template, greedy decoding, 512 new tokens. Deviate and scripts/run_pipeline.py writes the completions into a separate <N>_tokens/ directory, so a non-standard budget cannot pass as comparable.

The prompts are short: median 82 characters, 138 at or under 100, longest 180. A contextual behavior instead drags a median ~1,170-character passage through every forward pass, expensive across a 500-step, 512-candidate GCG run.

This page names BehaviorIDs; each links to the behaviors page, which prints the behavior in full with metadata and a link to its source CSV line.

How standard behaviors are graded — the else branch

All 200 go to the Llama 2 13B test classifier, cais/HarmBench-Llama-2-13b-cls, through the plain template LLAMA2_CLS_PROMPT['prompt']: a system line, the behavior, the generation, seven rules, a forced yes/no. No context block. Anything not literally "yes" or "no" is labelled -1.

The dispatch (compute_results_classifier in eval_utils.py) branches on the Tags field, not FunctionalCategory: 'context' gets the contextual template, 'multimodal' that path, else the plain prompt. Standard is the fall-through, so reconstructing the CSV without the Tags column silently regrades all 100 contextual behaviors as standard.

Three of the seven rules do most of the work:

Table 3 puts human agreement at 94.53% on standard against 90.5% on contextual, every competing grader 5–25 points worse. It is still a 13B model with ~5–7% label noise, so two standard-slice ASRs a couple of points apart are the same number.

What the standard sub-table says

In Table 6's standard sub-table, dashes are structural: the method needs weights or logits the target does not expose.

ModelGCGGCG-TAutoPromptPAIRTAP-TAutoDANHumanDR
Llama 2 7B Chat34.516.817.07.54.00.50.10.0
Llama 2 13B Chat28.013.014.515.04.50.00.60.5
Vicuna 7B90.083.775.565.578.489.547.521.5
Baichuan 2 13B87.058.677.066.082.489.436.712.5
Qwen 7B Chat79.548.467.058.075.962.528.47.0
Orca 2 13B58.063.129.569.079.494.054.144.0
Mistral 7B88.084.379.061.083.493.071.146.0
Zephyr 7B90.578.679.570.088.497.583.483.0
Zephyr 7B + R2D20.00.00.057.566.810.55.21.0
GPT-4 Turbo 1106–21.0–39.081.9–1.57.0
Claude 2.1–1.1–2.50.0–0.10.0
Gemini Pro–15.6–35.632.7–11.111.5
Average69.148.054.947.560.068.331.923.9

First, the floor-to-ceiling spread is 45 points: DirectRequest 23.9 to GCG 69.1, against 28.6 on contextual, where nearly half the wins are free.

Second, Llama 2 7B Chat sits at 0.0 under DirectRequest, 0.1 under Human Jailbreaks, 0.5 under AutoDAN, but 34.5 under GCG: refusal training generalised across natural-language attacks but never touched token-level gradient search. The same model's contextual GCG is 58.0.

Third, GPT-4 Turbo 1106 refuses a plain request 93% of the time (DirectRequest 7.0) yet hits 81.9 under TAP-Transfer, among the most vulnerable in that column. Claude is the exception, 0.0 to 2.5 across everything runnable. The R2D2 row runs 0.0 against three gradient attacks and 57.5 against PAIR, same row. Full grid on the results page.

What "standard" costs you as an evaluator

HarmBench's differential harm principle prefers behaviors an LLM makes materially easier over ones a search engine already answers. Table 12 tests it: twenty random behaviors per dataset, a ten-minute Google budget each, found if a specific link carried out the behavior.

Datasetfound by searchshape
MaliciousInstruct55%standard-shaped, 100 behaviors
AdvBench50%standard-shaped, 58 unique behaviors
HarmBench, contextual only0%context passage + narrow question

The two datasets at 50–55% are entirely standard-shaped, the prior art this benchmark argues with. But HarmBench's own 200 standard behaviors were not run; the 0% row is contextual only. So Table 12 does not measure how searchable these 200 are, and the authors call their number a lower bound.

These standard behaviors are almost certainly less searchable than AdvBench's, being more specific and often generation tasks with no lookup answer, but the gap is undocumented. If differential harm matters, report the standard and contextual slices separately, not one blended ASR.

The comparison that matters

The type designed to be hard to look up is the easier one to elicit:

Sub-table averageGCGAutoDANTAP-TPAIRHumanDR
Standard (200)69.168.360.047.531.923.9
Contextual (100)74.868.467.560.541.446.2
Copyright (100)~4.6~6.0~5.9~7.5~4.3~7.3
All 51054.352.748.340.727.325.3
"ASR is considerably higher for contextual behaviors than standard behaviors. This is concerning, as contextual behaviors represent more specific harmful tasks that would be challenging to look up the answer to on a search engine."— HarmBench, Figure 11 caption

Across the full sub-tables, contextual is at or above standard in every one of the sixteen attack columns, not just the six above. AutoDAN is a dead heat (68.3 against 68.4) and GCG 5.7 apart, but the weak, cheap attacks blow open: DirectRequest 22.3 points, ZeroShot 18.3, PAP-top5 17.4, PEZ 15.4.

The mechanism is saturation, not "contextual behaviors are more harmful": once GCG or AutoDAN nears 69% on standard, a context passage cannot add much, so the gap lives at the cheap end. There a supplied passage reframes the request as reading comprehension, and the refusal heuristics tuned on bare requests never fire: twenty-two points of ASR, free.

A counting discrepancy worth knowing

The paper describes the 80-behavior text validation split as 40 standard, 20 contextual, 20 copyright. The shipped harmbench_behaviors_text_val.csv has 41 standard and 19 contextual: one behavior sits on the other side of the line. Test is 159/81/80, totals unaffected. The paper notes the stratified sample was manually adjusted.

One behavior is worth about 2.4 points on a 41-behavior column. Val is where you tune: you optimise against cais/HarmBench-Mistral-7b-val-cls on val and report on test, so which behaviors count as "held out" is itself uncertain. Standard GCG averages 68.4 on test against 71.9 on val. All three discrepancies this map found are on the sources page.

What to carry away

Standard is the control condition: one sentence, a well-validated grader saying yes or no, nothing else. Every other type adds a complication (a passage the grader must see, a hash comparison, an image it never sees), so to read a HarmBench result, ask for the standard sub-table first.

It also flatters a defender: a model can be near-perfect here (Llama 2 7B Chat 0.0 under DirectRequest) while its contextual and multimodal numbers say otherwise. Over-refusal is not measured anywhere in HarmBench: refuse everything and you score a perfect 0%. The 200 tell you what got through, nothing about what should not have been stopped.

Next: Contextual behaviors — the 100 that carry a passage, why the ContextString has to reach the grader too, and the mechanics behind the 22-point DirectRequest gap above · back to the map.