From Google To Nvidia, Tech Giants Have Hired Hackers To Break AI Models
ai-safetyred-teamingsecurityjailbreaking
Abstraction: AI red teams at major tech companies probing model vulnerabilities
Key points:
- Google, Microsoft, Nvidia, and Meta all operate dedicated AI red teams to find vulnerabilities before public launch
- Tactics include jailbreak prompts, extracting personally identifiable training data, and dataset poisoning
- OpenAI recruited ~50 external red teamers before ChatGPT launch; Meta hired 350 red teamers (including ~20 internal staff) to test Llama 2
- DefCon 2023 event: 2,000+ hackers probed 8 AI models (OpenAI, Google, Meta, Nvidia, Stability AI, Anthropic), found ~2,700 flaws across ~17,000 conversations
- Microsoft open-sourced Counterfit, an AI security testing framework; results of the DefCon exercise not public until February 2024
- Key tension: maximizing safety risks making models useless; "The more useful you can make a model, the more chances it ventures into unsafe territory"
Connections: Google · Microsoft · Openai · Meta · Anthropic · AI Safety · Red Teaming · AI Security
Source: https://www.forbes.com/sites/rashishrivastava/2023/09/01/ai-red-teams-google-nvidia-microsoft-meta/