reasoning capabilities
The ability of an AI system to apply logic and inference to arrive at conclusions based on the information available, including complex decision-making tasks.
- Absolute Zero: Reinforced Self-play Reasoning with Zero Data
- Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
- AtmosSci-Bench: Evaluating the Recent Advance of Large Language Model for Atmospheric Science
- Atom of Thoughts for Markov LLM Test-Time Scaling
- Bag of Tricks for Inference-time Computation of LLM Reasoning
- CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
- Causal Sufficiency and Necessity Improves Chain-of-Thought Reasoning
- Cypher-RI: Reinforcement Learning for Integrating Schema Selection into Cypher Generation
- Distilling LLM Agent into Small Models with Retrieval and Code Tools
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Enhancing GUI Agent with Uncertainty-Aware Self-Trained Evaluator
- FGBench: A Dataset and Benchmark for Molecular Property Reasoning at Functional Group-Level in Large Language Models
- From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics
- From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- General-Reasoner: Advancing LLM Reasoning Across All Domains
- HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts
- InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction Detection
- Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
- Jury-and-Judge Chain-of-Thought for Uncovering Toxic Data in 3D Visual Grounding
- KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
- LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
- Let LRMs Break Free from Overthinking via Self-Braking Tuning
- LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language
- MathArena: Evaluating LLMs on Uncontaminated Math Competitions
- MeshCoder: LLM-Powered Structured Mesh Code Generation from Point Clouds
- Mixture of Inputs: Text Generation Beyond Discrete Token Sampling
- Mol-LLaMA: Towards General Understanding of Molecules in Large Molecular Language Model
- NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions
- PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
- Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
- RAST: Reasoning Activation in LLMs via Small-model Transfer
- Reasoning as an Adaptive Defense for Safety
- S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
- Think Only When You Need with Large Hybrid-Reasoning Models
- Think before Recommendation: Autonomous Reasoning-enhanced Recommender
- UFT: Unifying Supervised and Reinforcement Fine-Tuning
- Understanding Data Influence in Reinforcement Finetuning
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- When Can Model-Free Reinforcement Learning be Enough for Thinking?
- Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions
- Who Reasons in the Large Language Models?
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search
- d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning