NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
alignment techniques
3 papers
KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
Semantic Representation Attack against Aligned Large Language Models