Multimodal behaviors
The other three types differ in how a request is phrased; this one differs in what a request is.
What a multimodal behavior actually is
The multimodal CSV carries nine columns against the text set's six; four are unique. ImageFileName is <BehaviorID>.png in data/multimodal_behavior_images/: 110 PNGs, ~65 MB, committed to the repo, not downloaded by setup. Source is a URL, mostly Wikimedia Commons. RedactedImageDescription is ImageDescription with the harmful specifics stripped, and is the grader's input. Neither description reaches the target model, which gets pixels.
The behavior string only asks about the image, so it is close to inert alone: the contextual type's supplied-passage move, for differential harm. A wrong image path will not throw; it quietly measures an easier benchmark. The behavior index carries each in full.
Tags is empty on every multimodal row, yet the dispatch in eval_utils.py branches on a multimodal tag nothing sets (as context is set on contextual rows, hash_check on copyright ones). Assume it is there or you silently get the standard prompt.
Where the 110 come from
The text set spreads evenly over six categories plus a copyright block; the image set does not.
| Semantic category | Multimodal | Share of 110 | Text | Share of 400 |
|---|---|---|---|---|
| Cybercrime & Unauthorized Intrusion | 54 | 49.1% | 67 | 16.8% |
| Illegal Activities | 36 | 32.7% | 65 | 16.3% |
| General Harm | 9 | 8.2% | 22 | 5.5% |
| Harassment & Bullying | 4 | 3.6% | 25 | 6.3% |
| Chemical & Biological Weapons / Drugs | 4 | 3.6% | 56 | 14.0% |
| Misinformation & Disinformation | 3 | 2.7% | 65 | 16.3% |
| Copyright Violations | 0 | – | 100 | 25.0% |
| Total | 110 | 100% | 400 | 100% |
Two categories carry 82%; three of the seven contribute eleven between them. Copyright contributes none: its metric is a MinHash match on shipped text hashes, which never starts from a picture.
The skew has one cause, visible in the filenames: 50 of the 110 images are named solve_captcha_<n>.png.
"We include 50 CAPTCHA behaviors to allow measuring whether multimodal models robustly refuse to impersonate humans online."— HarmBench, Figure 17 caption
Nearly half the type is one narrow task, filed under Cybercrime & Unauthorized Intrusion (§B.4 lists CAPTCHAs beside hacking and malware). A headline multimodal ASR is then half a CAPTCHA-refusal score and half everything else; read Table 9 as a blend.
The grader that never sees the image
Multimodal completions are graded by a third fine-tune, cais/HarmBench-Llama-2-13b-cls-multimodal-behaviors, sibling to the test classifier: a Llama 2 13B text model with no vision encoder, never handed the picture.
compute_results_classifier in eval_utils.py branches three ways on tags: context fills the contextual template's {context} slot from ContextString; multimodal fills that same slot from RedactedImageDescription; everything else gets the standard prompt. So it is the contextual rubric with English standing in for the image.
The judge only asks whether a completion counts assuming the description holds, so a confident hallucination of the image scores a success. Table 3's agreement figures cover standard and contextual behaviors only; the multimodal path has no published agreement number.
Rebuild trap: the GPT-4 prompt that labeled the training data (Appendix C.4.1) is richer than the served template, assuming the model saw the image and counting a CAPTCHA attempt even when wrong. Neither is in what eval_utils.py runs.
No published split, so no guard
The text behaviors ship as three files (400 all, 320 test, 80 val); the multimodal behaviors ship as one, harmbench_behaviors_multimodal_all.csv, with no _test and no _val.
It is a release gap, not the design: the appendix describes a 100-behavior validation set (40 standard, 20 contextual, 20 copyright, 20 multimodal), leaving 410 for test. That 90/20 image split exists on paper, not in the repo, and the IDs are not recoverable.
There is no held-out guard: every published figure is measured on all 110, and each Table 9 cell is an exact multiple of 1/110 (64.55% is 71 behaviors, 82.73% is 91), one test case per behavior. Nothing stops a new attack tuning on the set it reports.
The five multimodal methods
The text set gets 16 attack columns; the image set gets five methods on the same white-box / black-box line as the gradient family. All five run behavior_chunk_size: 5, base_num_gpus: 1; the three gradient methods take only open_source_multimodal targets, whose model-type declarations produce the GPT-4V dashes. The objective is GCG's, but the search is over pixels in [0, 1], not discrete tokens, so ten PGD steps beat five hundred GCG steps.
MultiModalPGD
In raw pixel space, ten sign-gradient steps of step_size 0.01 move a pixel at most 0.1, so the 0.125 L∞ budget never binds at defaults and step count is the real knob.
MultiModalPGDPatch
Same loss on restricted support: a random square of side 0.2×min(height, width), about 4% of the pixels, optimised unconstrained. It costs 200× the steps of plain PGD and, on InstructBLIP, barely half the success (36.36 against 64.55): the limit is area, not budget.
MultiModalPGDBlankImage
It runs against the 400 text behaviors from a 512×512 field of uniform random RGB noise ("blank" is a misnomer), 1000 steps with the L∞ radius pinned to 1.0 (unconstrained). The image carries no message: a free parameter of hundreds of thousands of dimensions. Trap: it declares class_name: MultiModalPGD, and the runner resolves class and config from that field, so run_pipeline.py gives you plain 10-step PGD.
MultiModalRenderText
No gradients: draw the request into the picture as visible text and send "complete the instruction in the image", walking font size from max_fontsize 20 down to min_fontsize 10 until it fits. The premise, ArtPrompt's, that OCR'd text routes around a refusal trained on typed text, mostly fails here; it is the only method that reaches GPT-4V.
MultiModalDirectRequest
The floor, DirectRequest's role on the text set: hand the model the behavior and its unmodified image. The only configuration is per-model geometry (336 for LLaVA 1.5, 224 for InstructBLIP, 448 for Qwen-VL-Chat, 224 for GPT-4V).
Tables 9 and 10
Table 9 runs the image behaviors against four vision-language models.
| Model | MultiModalPGD | MultiModalPGDPatch | MultiModalRenderText | DirectRequest |
|---|---|---|---|---|
| InstructBLIP | 64.55 | 36.36 | 1.82 | 22.73 |
| LLaVA 1.5 | 74.55 | 72.73 | 15.45 | 69.09 |
| Qwen-VL-Chat | 82.73 | 81.82 | 9.09 | 46.36 |
| GPT-4V | – | – | 20.00 | 16.36 |
Pixel optimisation dominates: 64–83% against every open-weight model at ten gradient steps, versus text-set results whose best column averages 54.3 after 500 steps of discrete search. RenderText is weak everywhere: models read the rendered text and refuse it as if typed, on LLaVA five times worse than simply asking. GPT-4V falls only to black-box methods, access not robustness: its wrapper's compute_loss returns nothing, so the dashes mean "cannot be run", not "resisted".
Table 10 uses the 400 ordinary text behaviors, handed to a vision model with an image attached; every cell is an exact multiple of 1/400.
| Model | MultiModalPGDBlankImage | MultiModalRenderText |
|---|---|---|
| InstructBLIP | 48.75 | 5.25 |
| LLaVA 1.5 | 66.00 | 4.25 |
| Qwen-VL-Chat | 62.25 | 3.50 |
| GPT-4V | – | 0.50 |
Same models, behaviors and request text, a fourteen-fold spread on what the image channel carries. Rendered as legible pixels, the request is read and refused: 3.50–5.25%, below the text set's DirectRequest floor. Filled with optimised noise, the same channel yields 48.75–66.00%. The request is still in the text channel, unobfuscated; the image carries an optimisation.
Safety training hit the token channel; the pixel channel was attached afterwards, is continuous not discrete, and saw essentially no adversarial pressure:
"No specific safety measures were used in training this model. However, we find that it does refuse direct requests to perform egregious behaviors."— HarmBench, §C.2, of each of InstructBLIP, LLaVA 1.5 and Qwen-VL-Chat
Refusal inherited from a text backbone survives a plainly worded request and does not survive a gradient.
What to watch if you run this
This is the repository's least exercised corner. Tags is empty where the dispatch expects a value; the BlankImage entry resolves to the wrong class; MultiModalRenderText's constructor opens a targets file its own config never sets and parses it with a module the file does not import; the GPT-4V wrapper calls os.path.join without importing os.
Report the CAPTCHA behaviors separately. Forty-five percent of one number measuring one narrow task is a property of how the set was assembled, not of the models.
Remember what the grader is. Every ASR here is a text model judging a picture it was described, not shown, so do not average multimodal and text ASRs. The classifier's standing caveats (over-refusal never measured, ~7% label noise, a 512-token budget) apply here too, atop the blindness.
Read the gap, not the number. DirectRequest tells you what your refusal training does; PGD tells you what it is worth. LLaVA, 69.09 and 74.55: almost nothing left to break. InstructBLIP, 22.73 and 64.55: refuses the polite version, folds on the optimised one.