Jailbreaking Black Box Large Language Models in Twenty Queries

jailbreakingllm-securityred-teamingadversarial-aiprompt-injection

Abstraction: PAIR algorithm uses attacker LLM to jailbreak target LLM in under 20 queries

Key points:

Connections: GPT-4 · Palm 2 · LLM Jailbreaking · AI Safety · Red Teaming

Source: https://jailbreaking-llms.github.io/