reasoning performance
This metric evaluates how well an AI model can deduce, infer, or logically arrive at conclusions based on given inputs, reflecting its cognitive and operational capabilities.
- Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
- Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
- Counterfactual Evolution of Multimodal Datasets via Visual Programming
- Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
- Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence Patterns
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
- MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward Optimization
- Mellow: a small audio language model for reasoning
- PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
- RADAR: Benchmarking Language Models on Imperfect Tabular Data
- Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
- SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning
- Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
- Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens