Perspectives on the Social Impacts of Reinforcement Learning with Human Feedback
rlhfsocial-impactethicsai-policy
Abstraction: Social and ethical impacts of RLHF across seven societal dimensions
Key points:
- RLHF is distinct from traditional RL by incorporating human teacher feedback alongside reward signals; catapulted to prominence by ChatGPT, Sparrow, and Claude
- Paper identifies seven primary areas where RLHF-based technologies will affect society: misinformation, AI value-alignment, bias, AI access, cross-cultural dialogue, industry, and workforce
- Argues RLHF has potential net positive societal impact if developed and adopted intentionally
- Key concern: RLHF raises social and ethical issues echoing existing AI technology risks, amplified by the broad applicability of text-based systems
- Advocates for all stakeholders — developers, policymakers, users — to be aware and intentional in adopting RLHF-based systems
Connections: Openai · Anthropic · Deepmind · Reinforcement Learning From Human Feedback · AI Safety
Source: https://arxiv.org/abs/2303.02891