Haitao Mi
- Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
- MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data Curation
- The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
- Thoughts Are All Over the Place: On the Underthinking of Long Reasoning Models
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
- Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
- UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression