induction heads
- From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
- Interpretable Next-token Prediction via the Generalized Induction Head
- Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning
- What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains