NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
suboptimality gap
3 papers
Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback
Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs