adversarial manipulation
Adversarial manipulation refers to the intentional introduction of perturbations to input data to mislead AI models. This highlights vulnerabilities in models and is a crucial area of study for enhancing robustness against attacks.
- Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious Tools
- BAM-ICL: Causal Hijacking In-Context Learning with Budgeted Adversarial Manipulation
- CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
- Fortifying Time Series: DTW-Certified Robust Anomaly Detection
- Predicting the Performance of Black-box Language Models with Follow-up Queries
- TRAP: Targeted Redirecting of Agentic Preferences