attack success rates
Metrics that measure the effectiveness of adversarial attacks against AI systems, quantifying how often a model can be successfully deceived.
- AngleRoCL: Angle-Robust Concept Learning for Physically View-Invariant Adversarial Patches
- Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious Tools
- BAM-ICL: Causal Hijacking In-Context Learning with Budgeted Adversarial Manipulation
- Backdoor Cleaning without External Guidance in MLLM Fine-tuning
- BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization
- Best-of-N Jailbreaking
- Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs
- LoSplit: Loss-Guided Dynamic Split for Training-Time Defense Against Graph Backdoor Attacks
- Reinforcement Learning with Backtracking Feedback
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Safety Pretraining: Toward the Next Generation of Safe AI
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
- Semantic Representation Attack against Aligned Large Language Models