Week 7: Sequence Models & Attention
Course Generated Slides: Transformers
1. Sequence Models
- Recurrent Neural Networks (RNNs)
- Vanishing/exploding gradients
- Long Short-Term Memory (LSTM)
- Gated Recurrent Units (GRU)
- Applications
- Language modeling
- Sequence prediction
- Time series analysis
2. Attention Mechanisms
- Motivation for attention
- Limitations of RNNs
- Long-range dependencies
- Types of attention
- Bahdanau attention
- Luong attention
- Self-attention
- Visualization and interpretation
- Attention heatmaps
- Interpreting attention weights
3. Transformers
- Transformer architecture
- Multi-head attention
- Positional encoding
- Feed-forward layers
- Advantages over RNNs
- Parallelization
- Handling long sequences
- Applications
- Machine translation
- Text summarization
- Question answering
Required Reading
- Attention Is All You Need - The original Transformer paper that revolutionized sequence modeling and laid the foundation for modern LLMs
- Understanding Sequence Models
- Attention Mechanisms in Deep Learning
Learning Objectives
- Master the Transformer architecture and its components
- Understand self-attention and multi-head attention mechanisms
- Learn how positional encodings enable sequence modeling
- Implement key components of the Transformer architecture
- Compare Transformers with traditional RNN/LSTM approaches
Key Topics
- Self-attention mechanisms
- Multi-head attention
- Positional encodings
- Encoder-decoder architecture
- Scaled dot-product attention
Additional Resources