NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
dataset diversity
3 papers
Can LLMs Outshine Conventional Recommenders? A Comparative Evaluation
SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
THUNDER: Tile-level Histopathology image UNDERstanding benchmark