NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Yixin Liu
3 papers
Yale University
On Evaluating LLM Alignment by Evaluating LLMs as Judges
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks