reasoning tasks
Diverse cognitive challenges that require understanding, interpreting, and making inferences about data, often evaluated in AI systems.
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
- Can Dependencies Induced by LLM-Agent Workflows Be Trusted?
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
- Composing Global Solutions to Reasoning Tasks via Algebraic Objects in Neural Nets
- ConfTuner: Training Large Language Models to Express Their Confidence Verbally
- Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
- DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
- Fast attention mechanisms: a tale of parallelism
- Generalizable Reasoning through Compositional Energy Minimization
- LILO: Learning to Reason at the Frontier of Learnability
- MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
- Matching Markets Meet LLMs: Algorithmic Reasoning with Ranked Preferences
- Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of Go
- PID-controlled Langevin Dynamics for Faster Sampling on Generative Models
- PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
- PlanU: Large Language Model Reasoning through Planning under Uncertainty
- Reasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning Models
- SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
- Scalable Best-of-N Selection for Large Language Models via Self-Certainty
- Scaling Physical Reasoning with the PHYSICS Dataset
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
- ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
- Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
- Weaver: Shrinking the Generation-Verification Gap by Scaling Compute for Verification