NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Amrit Singh Bedi
3 papers
University of Central Florida
Does Thinking More Always Help? Mirage of Test-Time Scaling in Reasoning Models
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
On the Sample Complexity Bounds of Bilevel Reinforcement Learning