The challenges of reinforcement learning from human feedback (RLHF) - TechTalks

rlhfalignmentreward-hackingllm-safety

Abstraction: RLHF limitations across feedback, reward modeling, and policy

Key points:

Connections: Openai · Chatgpt · Reinforcement Learning From Human Feedback · AI Safety

Source: https://bdtechtalks.com/2023/09/04/rlhf-limitations/