Reinforcement Learning From Human Feedback
concepts · 4 notes linked
Related: Anthropic · AI Safety · Openai · Chatgpt · AI Alignment · Deepmind · Paul Christiano · Large Language Models
Notes
- 255 — Paul Christiano's defense of RLHF research positive net impact
- Perspectives on the Social Impacts of Reinforcement Learning with Human Feedback — Social and ethical impacts of RLHF across seven societal dimensions
- The Capacity for Moral Self-Correction in Large Language Models — RLHF-trained LLMs can self-correct harmful outputs when instructed
- The challenges of reinforcement learning from human feedback (RLHF) - TechTalks — RLHF limitations across feedback, reward modeling, and policy