reasoning benchmarks
Standard tests or datasets used to evaluate the reasoning capabilities of AI models, assessing their performance on tasks requiring logical deduction.
- Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
- BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
- Behavior Injection: Preparing Language Models for Reinforcement Learning
- C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for Reasoning
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
- Corrector Sampling in Language Models
- Dependency Matters: Enhancing LLM Reasoning with Explicit Knowledge Grounding
- Efficient Large Language Model Inference with Neural Block Linearization
- ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
- Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- Hyperbolic Fine-Tuning for Large Language Models
- KLASS: KL-Guided Fast Inference in Masked Diffusion Models
- Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
- Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
- MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- Predicting the Performance of Black-box Language Models with Follow-up Queries
- QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
- R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPO
- SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought
- SPMDM: Enhancing Masked Diffusion Models through Simplifing Sampling Path
- SeRL: Self-play Reinforcement Learning for Large Language Models with Limited Data
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
- The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity