NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
Lei Hou
3 papers
Tsinghua University, Tsinghua University
AGENTIF: Benchmarking Large Language Models Instruction Following Ability in Agentic Scenarios
How do Transformers Learn Implicit Reasoning?
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons