Best-of-N Jailbreaking

Sanmi Koyejo (Stanford University / Virtue AI) · Erik Jones (UC Berkeley) · Fazl Barez (University of Oxford) · Rylan Schaeffer (Stanford University) · John Hughes (Anthropic) · Sara Price (Anthropic) · Aengus Lynch (University College London, University of London) · Arushi Somani (Anthropic) · Henry Sleight (MATS Program) · Ethan Perez (Anthropic) · Mrinank Sharma (University of Oxford)
attack success ratesaudio language modelsaugmented promptsbest-of-n jailbreakingblack-box algorithmcircuit breakersclosed-source language modelsharmful response elicitationinput sensitivitymodality-specific augmentationsmultimodal exploitationoptimized prefix attackpower-law-like behaviorreasoning modelsstate-of-the-art defensesvision language models

We introduce Best-of-N (BoN) Jailbreaking, a simple black-box algorithm that jailbreaks frontier AI systems across modalities. BoN Jailbreaking works by repeatedly sampling variations of a prompt with a combination of augmentations