Mengdi Wang
- CURE: Co-Evolving Coders and Unit Testers via Reinforcement Learning
- DISC: Dynamic Decomposition Improves LLM Inference Scaling
- Does Thinking More Always Help? Mirage of Test-Time Scaling in Reasoning Models
- MMaDA: Multimodal Large Diffusion Language Models
- ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs
- Securing the Language of Life: Inheritable Watermarks from DNA Language Models to Proteins
- Training-Free Guidance Beyond Differentiability: Scalable Path Steering with Tree Search in Diffusion and Flow Models