reinforcement fine-tuning
A technique where a pre-trained model is subsequently refined through reinforcement learning to improve its performance on a specific task. This often combines the advantages of supervised learning with the adaptability of reinforcement learning.
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- Angles Don’t Lie: Unlocking Training‑Efficient RL Through the Model’s Own Signals
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning
- BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- Pre-Trained Policy Discriminators are General Reward Models
- Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
- SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
- TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
- UFT: Unifying Supervised and Reinforcement Fine-Tuning
- Understanding Data Influence in Reinforcement Finetuning
- VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning