NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
evaluators
3 papers
AI Debate Aids Assessment of Controversial Claims
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
On Evaluating LLM Alignment by Evaluating LLMs as Judges