AI Safety
concepts · 87 notes linked
Related: Openai · Large Language Models · Anthropic · Google · AI Policy · Chatgpt · AI Agents · Zvi Mowshowitz
Notes
- 'Eugenics on steroids': the toxic and contested legacy of Oxford's Future of Humanity Institute — FHI closure legacy of longtermism scandal and controversy
- 'The illusion of thinking': Apple research finds AI models collapse and give up with hard puzzles — Apple study showing large reasoning models collapse on hard logic puzzles
- 255 — Paul Christiano's defense of RLHF research positive net impact
- A Complete Breakdown of the Claude Mythos 1 Leak and Features — Leaked Claude Mythos 1 capabilities including 69% Exploit Bench score
- A Machine Learning Engineer's Guide To The AI Act — EU AI Act compliance requirements and implications for ML practitioners
- A Note From Ray Kurzweil on the Recent Call to Pause Work on AI More Powerful Than GPT-4 — Ray Kurzweil opposing FLI's AI pause letter citing vagueness and coordination failure
- AI #171: False Flag — Weekly roundup; OpenAI-linked PAC false-flag scandal
- AI #172: The First Fable — Weekly roundup excluding the Fable model itself
- AI #174: You're It — Weekly AI roundup, Claude Tag, medical scanners, agent security
- AI #175: The Fable Continues — Weekly AI roundup covering Fable's return aftermath
- AI 2027 — HN debate on AI 2027 scenario: AGI timelines, LLM limits, alignment risks
- AI Debate 2: Night of a thousand AI scholars — Sixteen scholars debate moving AI beyond deep learning
- AI Sentience, Agency and Catastrophic Risk | TWIML - The Voice of Machine Learning & AI — Yoshua Bengio discusses AI catastrophic risk and governance
- AI Virtues As Missing Bedrock Ingredient For Responsible AI Says AI Ethics And AI Law — Virtue ethics framework proposed as foundation for AI ethics principles
- AI fears overblown? Theoretical physicist calls chatbots 'glorified tape recorders' | CNN Business — Michio Kaku argues chatbots are limited and quantum computing is the true next stage
- American Government Takes Down Claude Fable — Commerce export controls abruptly shut down Fable
- Artificial Intelligence and the Future of Humans — Pew 2018 expert canvassing on AI's impact on human life by 2030
- Artificial Intelligence — The Revolution Hasn't Happened Yet — Michael Jordan argues AI hype obscures the real ML/data-engineering challenge
- At TED AI 2023, experts debate whether we've created "the new electricity — TED AI 2023 first AI-only TED conference debates AGI benefits and risks
- Can the Most Abstract Math Make the World a Better Place? | Quanta Magazine — Applied category theory for real-world systems modeling and AI safety
- Claude Fable 5 and Mythos 5: The System Card — Reading the Fable/Mythos 319-page system card
- Claude Sonnet 5 Is Not Frontier But Has Its Uses — Sonnet 5 system card review, cheaper faster non-frontier model
- Deepfakes are now trying to change the course of war | CNN Business — Deepfake videos deployed as wartime disinformation during Ukraine conflict
- Fable #6: The Return of the King — Anthropic Fable 5 restored after government export-control blip
- Finding Common Ground on AI — Semafor Event — Semafor DC event bridging Silicon Valley and Washington AI perspectives
- From Google To Nvidia, Tech Giants Have Hired Hackers To Break AI Models — AI red teams at major tech companies probing model vulnerabilities
- Full Committee Hearing - Artificial Intelligence Advancing Innovation Towards the National Interest — US House Science Committee 2023 hearing on AI policy and national interest
- GPT-5.6: The System Card — Review of OpenAI GPT-5.6 Sol/Terra/Luna system card
- GitHub - daviddao/awful-ai: Awful AI is a curated list to track current scary usages of AI - hoping to raise awareness — Curated list cataloguing harmful and unethical real-world AI deployments
- Godfather of Artificial Intelligence" Geoffrey Hinton on the promise, risks of advanced AI — Geoffrey Hinton 60 Minutes interview warning of AI existential and societal risks
- Google AI Zones Out While Being Trained On Mandatory Racial Sensitivity Data Set — Satirical piece mocking AI bias training and AI overconfidence
- Google DeepMind boss hits back at Meta AI chief over 'fearmongering' claim — Hassabis vs LeCun debate on AI safety, regulation, and open-source control
- Google Researchers Can Create an AI That Thinks a Lot Like You After Just a Two-Hour Interview — Stanford/Google study simulates 1,000 people as LLM agents from interview transcripts
- Google says it's committed to ethical AI research. Its ethical AI team isn't so sure. — Google ethical AI team dysfunction after Timnit Gebru firing
- Google's medical AI chatbot is already being tested in hospitals — Google Med-PaLM 2 hospital pilot tests medical LLM accuracy and safety
- Google's top AI scientist says 'learning how to learn' will be next generation's most needed skill — Demis Hassabis advocates meta-learning skills as AGI approaches within a decade
- Home - AI Now Institute — AI Now Institute homepage covering AI power, regulation, and accountability
- How ChatGPT is changing the way cybersecurity practitioners look at the potential of AI — ChatGPT dual-use cybersecurity capabilities surprise skeptical security researchers
- How Peter Thiel's Relationship With Eliezer Yudkowsky Launched the AI Revolution — Thiel-Yudkowsky-Altman network origins of DeepMind and OpenAI
- I worked on Google's AI. My fears are coming true — Fired Google engineer warns LaMDA and Bing chatbots show sentient behavior
- Is AI a danger to humanity or our salvation? — Hinton, LeCun, and Bengio split on AI existential risk after ChatGPT
- Jailbreaking Black Box Large Language Models in Twenty Queries — PAIR algorithm uses attacker LLM to jailbreak target LLM in under 20 queries
- Just Because They've Turned Against Humanity Doesn't Mean We Should Defund the Terminator Program — Satire mapping AI killer robots to police defunding debate
- Laser attack blinds autonomous vehicles, deleting pedestrians and confusing cars — Timed laser attacks spoof lidar sensors to erase pedestrians from AV perception
- MCP doesn't move data. It moves trust — MCP as AI governance and intent-control layer over APIs
- Machines won't ever make decisions on their own, says Pentagon AI chief | CNN — Pentagon CDAO Craig Martell on mandatory human AI oversight
- Mapping the Mind of a Large Language Model — Anthropic extracts millions of interpretable features from Claude 3 Sonnet
- Mathematical methods and human thought in the age of AI — Terence Tao on human-centered AI integration in mathematics and society
- Men Are Creating AI Girlfriends and Then Verbally Abusing Them — Ethical and psychological dimensions of AI companion abuse on Replika
- New Go-playing trick defeats world-class Go AI—but loses to human amateurs — Adversarial policy exploits off-distribution moves to beat world-class Go AI
- Noam Chomsky says A.I. is far from 'true intelligence' and ChatGPT is the 'banality of evil' | Fortune — Chomsky argues LLMs lack true reasoning and exhibit indifference to truth
- One of the three 'godfathers of A.I.' feels 'lost' because of the direction the technology has taken | Fortune — AI pioneers Bengio and Hinton warn of existential risk; LeCun dissents
- Open-Source AI Is Uniquely Dangerous — Unsecured open-source AI uniquely risky because safety features cannot be re-patched
- OpenAI executives say releasing ChatGPT for public use was a last resort after running into multiple hurdles — and they're shocked by its popularity — OpenAI's surprise at ChatGPT's viral adoption and CEO's AI risk warnings
- OpenAI is throwing everything into building a fully automated researcher — OpenAI's North Star automated AI researcher agent
- OpenAI says there's only a small chance ChatGPT will help create bioweapons — OpenAI self-study finds GPT-4 gives marginal bioweapon research uplift
- OpenAI wasn't expecting Sora's copyright drama — Sora launch copyright backlash and OpenAI policy reversal at DevDay 2025
- OpenAI's board has fired Sam Altman — Hacker News thread on OpenAI board firing Altman
- Oxford shuts down institute run by Elon Musk-backed philosopher — Oxford closes Future of Humanity Institute after 19 years amid scandals
- People Are Using A 'Grandma Exploit' To Break AI - Kotaku — Roleplay persona prompts bypass AI safety guardrails
- Perspectives on the Social Impacts of Reinforcement Learning with Human Feedback — Social and ethical impacts of RLHF across seven societal dimensions
- Readings — Curated reading list for Dan Hendrycks ML safety course curriculum
- Sam Altman enters his power era — Sam Altman reinstated as OpenAI CEO after failed board ouster
- Samsung workers made a major error by using ChatGPT — Samsung engineers leaked trade secrets via ChatGPT input data retention
- Showcasing Agile Safety Classifiers with Gemma | Google Codelabs — LoRA fine-tuning Gemma as a hate-speech safety classifier
- Stanford Director: AI Scientists' "Frontal Cortex Is Massively Underdeveloped — Stanford HAI director criticizing AI researchers' ethical immaturity
- State of AI, September 2025: A Broad Primer from an AI Team Lead in Education — Education-focused overview of AI history, capabilities, agents, and AGI trajectory
- Tesla Autopilot gets tricked into accelerating from 35 to 85 mph with modified speed limit sign — Sticker on speed sign fools Tesla Autopilot MobilEye camera into speeding
- The 'Godfather of AI' Has a Hopeful Plan for Keeping Future AI Friendly — Geoffrey Hinton's views on LLM risks and analog computing as AI safety mitigation
- The AI Agent Era Requires a New Kind of Game Theory — CMU researcher on agent security risks and multi-agent game theory
- The Capacity for Moral Self-Correction in Large Language Models — RLHF-trained LLMs can self-correct harmful outputs when instructed
- The Doomsday Invention — Nick Bostrom's superintelligence thesis and AI existential risk debate
- The Once And Future Fable #4 — Fable takedown aftermath, cyber-defense, restoration odds
- The Pentagon says AI is speeding up its 'kill chain' | TechCrunch — AI accelerating Pentagon kill chain planning amid usage policy tensions
- The author of SB 1047 introduces a new AI bill in California | TechCrunch — California SB 53 creates AI whistleblower protections and public compute cluster
- The challenges of reinforcement learning from human feedback (RLHF) - TechTalks — RLHF limitations across feedback, reward modeling, and policy
- These Women Tried to Warn Us About AI — Women researchers warned of AI bias harms
- USAF Official Says He 'Misspoke' About AI Drone Killing Human Operator in Simulated Test — Viral USAF AI drone kills-operator story was a hypothetical thought experiment
- What AI Models for War Actually Look Like — Military-specialized AI startup building models for mission planning and decision dominance
- What Is Gibberlink Mode, AI's Secret Language? — AI-to-AI protocol enabling machine-efficient communication beyond human language
- What Really Made Geoffrey Hinton Into an AI Doomer — Hinton's reasons for leaving Google to warn about accelerating AI risk
- What the departing White House chief tech advisor has to say on AI — Biden OSTP director reflects on AI risks and policy accomplishments
- What the departing White House chief tech advisor has to say on AI — Biden OSTP director reflects on AI policy, risks, and future under Trump
- What was 60 Minutes thinking, in that interview with Geoff Hinton? — Gary Marcus rebuttal annotating 60 Minutes Hinton AI interview
- Who Should Stop Unethical A.I.? — Debate over ethics review mechanisms for AI research publication
- Why Computers Won't Make Themselves Smarter — Skeptical argument against recursive AI self-improvement and intelligence explosion
- Will AI end everything? A guide to guessing | EAG Bay Area 23 — EA Global talk estimating ~19% probability of AI-caused civilizational doom