📰 News OpenAI Built an AI That Hacks Its Own Models — And It Already Found Attacks No Human Has Seen
OpenAI trained an AI super-hacker called GPT-Red to attack its own models. It succeeds where human red-teamers fail — and GPT-5.6 is the safer for it. But the asymmetry between attacker and defender is now an arms race inside one company.