Best-of-N Jailbreaking
attack success ratesaudio language modelsaugmented promptsbest-of-n jailbreakingblack-box algorithmcircuit breakersclosed-source language modelsharmful response elicitationinput sensitivitymodality-specific augmentationsmultimodal exploitationoptimized prefix attackpower-law-like behaviorreasoning modelsstate-of-the-art defensesvision language models
We introduce Best-of-N (BoN) Jailbreaking, a simple black-box algorithm that jailbreaks frontier AI systems across modalities. BoN Jailbreaking works by repeatedly sampling variations of a prompt with a combination of augmentations