red-teaming
Red-teaming in AI involves deliberately testing and challenging AI systems to identify vulnerabilities, weaknesses, or ethical issues. This practice helps improve model robustness and safety in real-world applications.
- CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion Models
- CoP: Agentic Red-teaming for Large Language Models using Composition of Principles
- GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
- LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
- PurpCode: Reasoning for Safer Code Generation
- Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling
- The Right to Red-Team: Adversarial AI Literacy as a Civic Imperative in K-12 Education