attack effectiveness
A measure of how successfully an adversarial attack compromises the integrity or performance of an AI model, highlighting vulnerabilities to malicious inputs.
- AngleRoCL: Angle-Robust Concept Learning for Physically View-Invariant Adversarial Patches
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
- Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment
- Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability
- Semantic Representation Attack against Aligned Large Language Models