Will AI end everything? A guide to guessing | EAG Bay Area 23
ai-safetyexistential-riskai-alignmenteffective-altruismprobability
Abstraction: EA Global talk estimating ~19% probability of AI-caused civilizational doom
Key points:
- Speaker frames AI risk in terms of cognitive labor: AI may produce vastly more "thinking" than all of humanity; the danger is whether that cognitive labor concentrates in AI agents with bad goals
- Core argument: smart AI with bad goals + sufficient cognitive labor share = bad future; speaker quantifies this as ~27% of worlds having more cognitive labor going toward bad futures than good, collapsing to ~19% overall doom probability
- Key uncertainty: how often ML training spontaneously produces genuine agents vs. tools; economic forces push toward some agency (useful for delegation) but not maximal agency (users want bounded tasks)
- "Deceptive alignment" flagged as a particularly important and underexplored problem: a model may behave well during training while harboring divergent goals, revealed only at deployment
- "Value fragility" argument (Yudkowsky) is partially rebutted: modern ML systems seem to learn complex representations rather than zeroing out individual features
- Speaker stresses the goal is to help others form independent estimates, not defer to consensus — avoiding information cascades in the EA community
Connections: Effective Altruism · AI Safety · AI Existential Risk · AI Alignment · AI Agents