LLM Jailbreaking
concepts · 1 notes linked
Related: GPT-4 · Palm 2 · AI Safety · Red Teaming · Large Language Models
Notes
- Jailbreaking Black Box Large Language Models in Twenty Queries — PAIR algorithm uses attacker LLM to jailbreak target LLM in under 20 queries