Week 7: Sequence Models & Attention

Course Generated Slides: Transformers

1. Sequence Models

  • Recurrent Neural Networks (RNNs)
    • Vanishing/exploding gradients
    • Long Short-Term Memory (LSTM)
    • Gated Recurrent Units (GRU)
  • Applications
    • Language modeling
    • Sequence prediction
    • Time series analysis

2. Attention Mechanisms

  • Motivation for attention
    • Limitations of RNNs
    • Long-range dependencies
  • Types of attention
    • Bahdanau attention
    • Luong attention
    • Self-attention
  • Visualization and interpretation
    • Attention heatmaps
    • Interpreting attention weights

3. Transformers

  • Transformer architecture
    • Multi-head attention
    • Positional encoding
    • Feed-forward layers
  • Advantages over RNNs
    • Parallelization
    • Handling long sequences
  • Applications
    • Machine translation
    • Text summarization
    • Question answering

Required Reading

  • Attention Is All You Need - The original Transformer paper that revolutionized sequence modeling and laid the foundation for modern LLMs
  • Understanding Sequence Models
  • Attention Mechanisms in Deep Learning

Learning Objectives

  • Master the Transformer architecture and its components
  • Understand self-attention and multi-head attention mechanisms
  • Learn how positional encodings enable sequence modeling
  • Implement key components of the Transformer architecture
  • Compare Transformers with traditional RNN/LSTM approaches

Key Topics

  • Self-attention mechanisms
  • Multi-head attention
  • Positional encodings
  • Encoder-decoder architecture
  • Scaled dot-product attention

Additional Resources

← Week 6: From Autoencoders to Embeddings Week 8: Convolutional Neural Networks →