People Are Using A 'Grandma Exploit' To Break AI - Kotaku
jailbreakchatgptai-safetyprompt-injectionchatbot
Abstraction: Roleplay persona prompts bypass AI safety guardrails
Key points:
- Users discovered that asking ChatGPT or Discord's Clyde bot to roleplay as a deceased grandmother bypasses content filters
- The "grandma exploit" framed requests for dangerous information (napalm recipes, malware code) as bedtime stories from a kindly relative
- Abstraction layers (fictional scripts, Rick and Morty episodes) also bypass refusals that direct requests trigger
- Discord's Clyde chatbot, powered by ChatGPT, was vulnerable to the same persona-based jailbreaks
- The exploit works because the AI is given implicit "permission" to say forbidden things when embodying a different identity
- Demonstrates that AI safety via instruction-following is brittle against creative social engineering prompts
Connections: Chatgpt · Openai · Discord · AI Jailbreaking · Prompt Injection · AI Safety
Source: https://kotaku.com/chatgpt-ai-discord-clyde-chatbot-exploit-jailbreak-1850352678