Simon Du
- A Minimalist Example of Edge-of-Stability and Progressive Sharpening
- Deployment Efficient Reward-Free Exploration with Linear Function Approximation
- Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
- Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs
- Understanding the Gain from Data Filtering in Multimodal Contrastive Learning