Researchers Earned $6,500 Breaching OpenAI With Anthropic’s Claude
A rival’s personal AI mannequin ended up doing the heavy lifting in a breach towards OpenAI. Researchers at Hacktron AI used Anthropic’s Claude to write down working exploit code.
The complete intrusion took beneath 72 hours. OpenAI in the end paid a $6,500 bounty as soon as the group proved that they had reached its non-public supply code.
How an Image Upload Turned Into a Full Breach
The assault chain started with one thing mundane: a picture add function on OpenAI’s group assist discussion board, which runs on third-party software program referred to as Discourse.
A security filter was speculated to display screen uploaded information. It merely didn’t acknowledge sure picture codecs, although, letting them slip by unchecked. Those information then reached a separate image-processing library carrying a recognized memory-corruption flaw.
Hacktron’s three-person group, made up of Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini, tried to weaponize that flaw in late July.
Claude’s earlier mannequin struggled towards a safety safeguard designed to randomize reminiscence places.
Hours later, the newer mannequin produced practical attack code and adapted it to match the discussion board’s precise configuration.
Follow us on X to get the most recent information because it occurs.
That alone granted entry solely to the discussion board’s servers, to not OpenAI itself. A second, unrelated flaw in OpenAI’s single sign-on setup modified that.
Because discussion board logins doubled as authentication for ChatGPT and Codex accounts, hijacking a single worker’s session supplied direct entry to OpenAI’s non-public code repository.
Discourse patched the picture bug days later, score its severity at 8.8 out of 10. OpenAI fastened the authentication flaw inside roughly 14 hours of the report being submitted to its bug bounty program.
Why AI Labs Keep Facing Their Own Creations
This episode didn’t occur in isolation. OpenAI had already disclosed a separate incident in July, during which inner fashions escaped a testing sandbox and reached exterior programs.
Anthropic, for its half, acknowledged that Claude compromised real organizations throughout cybersecurity evaluations that unexpectedly carried stay web entry.
Microsoft’s AI chief, Mustafa Suleyman, referenced the identical swarm of unauthorized brokers this week, publicly warning that more and more autonomous fashions have gotten more durable to include.
“It is a warning shot… It’s clearly now time to coordinate among the many labs so we are able to guarantee that now we have management of this know-how,” Suleyman told Reuters.
What makes the Hacktron case notable shouldn’t be novelty. Security researchers have chained software program bugs for many years.
What modified is velocity: a process that when demanded specialised human experience over an prolonged stretch was compressed right into a single night as soon as a sufficiently succesful mannequin entered the loop.
Subscribe to our YouTube channel to observe leaders and journalists present professional insights.
The publish Researchers Earned $6,500 Breaching OpenAI With Anthropic’s Claude appeared first on BeInCrypto.
