Wenxuan Wang
- End-to-End Vision Tokenizer Tuning
- MedChain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence
- Off-policy Reinforcement Learning with Model-based Exploration Augmentation
- Synthetic Series-Symbol Data Generation for Time Series Foundation Models
- Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
- Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
- Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training