NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
evaluation methods
4 papers
Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
Torch-Uncertainty: Deep Learning Uncertainty Quantification