jailbreak attacks

Security vulnerabilities in AI models where adversarial inputs cause the model to produce unintended or harmful outputs. Understanding and mitigating these attacks is crucial for ensuring the safe deployment of AI systems.

15 papers