Dawn Song
- A Sustainable AI Economy Needs Data Deals That Work for Generators
- BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
- OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
- OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
- SECODEPLT: A Unified Benchmark for Evaluating the Security Risks and Capabilities of Code GenAI
- Scalable Best-of-N Selection for Large Language Models via Self-Certainty
- VMDT: Decoding the Trustworthiness of Video Foundation Models
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations