empirical findings
- More of the Same: Persistent Representational Harms Under Increased Representation
- Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers
- ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
- Uncertainty-Guided Exploration for Efficient AlphaZero Training
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding