Week 13: Evaluating LLMs - Metrics and Methods

1. Traditional NLP Evaluation

  • Quantitative metrics
    • BLEU, ROUGE, METEOR
    • Perplexity
    • F1 scores
  • Task-specific metrics
    • Question answering accuracy
    • Translation quality
    • Summarization evaluation
  • Limitations of traditional metrics
    • Context sensitivity
    • Semantic understanding
    • Human-like evaluation

2. LLM-Specific Evaluation

  • Benchmark suites
    • GLUE and SuperGLUE
    • BIG-bench
    • HELM framework
  • Capability testing
    • Reasoning assessment
    • Knowledge probing
    • Task generalization
  • Behavioral evaluation
    • Instruction following
    • Output consistency
    • Chain-of-thought accuracy

3. Human Evaluation Protocols

  • Evaluation frameworks
    • Likert scales
    • A/B testing
    • Expert review protocols
  • Annotation guidelines
    • Quality criteria
    • Rubric development
    • Inter-rater reliability
  • Systematic assessment
    • Blind evaluation
    • Control questions
    • Statistical significance

Critical Perspectives

Required Reading

  • Beyond the Imitation Game: Quantifying and Extrapolating LLM Capabilities
  • Holistic Evaluation of Language Models (HELM)

Learning Objectives

  • Understand different evaluation metrics for LLMs
  • Learn to design and implement evaluation frameworks
  • Compare and analyze LLM performance
← Week 12: Retrieval Augmented Generation (RAG) Week 14: LLMs as Decision Makers and Agents →