reasoning trajectories
The pathways taken by an AI system as it applies logic and inference to process information, leading to conclusions or decisions based on a sequence of reasoning steps.
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
- Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
- ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- Multi-step Visual Reasoning with Visual Tokens Scaling and Verification
- R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPO
- SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning Models
- Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
- SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning
- VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning