The Capacity for Moral Self-Correction in Large Language Models
rlhfmoral-self-correctionai-alignmentai-safetyanthropicllm-scaling
Abstraction: RLHF-trained LLMs can self-correct harmful outputs when instructed
Key points:
- Tests hypothesis that RLHF-trained language models can "morally self-correct" — avoiding harmful outputs when explicitly instructed to do so
- Strong evidence confirmed across three experiments covering stereotyping, bias, and discrimination
- Capability for moral self-correction emerges at 22B model parameters and typically improves with both larger model size and more RLHF training
- Two enabling capabilities at sufficient scale: (1) following instructions and (2) learning complex normative concepts of harm
- Results described as "cause for cautious optimism" regarding training LLMs to abide by ethical principles
- Authors include Dario Amodei, Samuel Bowman, and many Anthropic researchers (Ganguli, Askell, et al.)
Connections: Anthropic · Reinforcement Learning From Human Feedback · AI Safety · Large Language Models · AI Alignment
Source: https://arxiv.org/abs/2302.07459