NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
reward mechanism
3 papers
Learning to Clean: Reinforcement Learning for Noisy Label Correction
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models