Sanmi Koyejo
- Aligning Compound AI Systems via System-level DPO
- AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
- Best-of-N Jailbreaking
- Fantastic Bugs and Where to Find Them in AI Benchmarks
- KGGen: Extracting Knowledge Graphs from Plain Text with Language Models
- Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
- Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
- The Leaderboard Illusion
- Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness