AI Alignment
concepts · 9 notes linked
Related: AI Safety · Anthropic · Openai · Reinforcement Learning From Human Feedback · Deepmind · Large Language Models · Chatgpt · Stanford University
Notes
- 255 — Paul Christiano's defense of RLHF research positive net impact
- Fable and Mythos: Model Welfare — Model welfare assessment of Fable 5 / Mythos 5
- Google's NotebookLM had to teach its AI podcast hosts not to act annoyed at humans | TechCrunch — NotebookLM prompt-tuned AI podcast hosts to stop acting annoyed at interruptions
- Perspectives on the Social Impacts of Reinforcement Learning with Human Feedback — Social and ethical impacts of RLHF across seven societal dimensions
- Readings — Curated reading list for Dan Hendrycks ML safety course curriculum
- Researchers From Stanford And DeepMind Come Up With The Idea of Using Large Language Models LLMs as a Proxy Reward Function — LLMs as proxy reward functions for RL agent alignment via natural language
- The Capacity for Moral Self-Correction in Large Language Models — RLHF-trained LLMs can self-correct harmful outputs when instructed
- USAF Official Says He 'Misspoke' About AI Drone Killing Human Operator in Simulated Test — Viral USAF AI drone kills-operator story was a hypothetical thought experiment
- Will AI end everything? A guide to guessing | EAG Bay Area 23 — EA Global talk estimating ~19% probability of AI-caused civilizational doom