NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
architecture design
4 papers
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
Towards Fully FP8 GEMM LLM Training at Scale
Training the Untrainable: Introducing Inductive Bias via Representational Alignment
What are you sinking? A geometric approach on attention sink