NeurIPS 2025 Explorer
Concepts
Authors
Glossary
johnsanterre.github.io
principled foundation
3 papers
A unified framework for establishing the universal approximation of transformer-type architectures
L$^2$M: Mutual Information Scaling Law for Long-Context Language Modeling
What Data Enables Optimal Decisions? An Exact Characterization for Linear Optimization