NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
evaluation performance
3 papers
AgentAuditor: Human-level Safety and Security Evaluation for LLM Agents
Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
Scaling Up Active Testing to Large Language Models