Researchers Earned $6,500 Breaching OpenAI With Anthropic's Claude
A rival‘s own AI model ended up doing the heavy lifting in a breach against OpenAI. Researchers at Hacktron AI used Anthropic’s Claude to write working exploit code. The entire intrusion took under 72 hours. OpenAI ultimately paid a $6,500 bounty once the team proved they had reached its private source code. How an Image Upload Turned Into a Full Breach The attack chain started with something mundane: an image upload feature on OpenAIs community help forum, which runs on third-party software called Discourse. A safety filter was supposed to screen uploaded files. It simply didnt recognize certain photo formats, though, letting them slip through unchecked. Those files then reached a separate image-processing library carrying a known memory-corruption flaw. Hacktrons three-person team, made up of Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini, attempted to weaponize that flaw in late July. Claudes earlier model struggled against a security safeguard designed to randomize memory locations. Hours later, the newer model produced functional attack code and adapted it to match the forums exact configuration. Follow us on X to get the latest news as it happens. How Claude Helped Researchers Earn $6,500 Breaching OpenAI. Source: Hacktron AI That alone granted access only to the forum‘s servers, not to OpenAI itself. A second, unrelated









