NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Yibin Wang
3 papers
Huazhong University of Science and Technology
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning