|

Anthropic’s Claude helped 3 researchers breach OpenAI in under 72 hours

Anthropic’s Claude helped three safety researchers breach OpenAI accounts and attain an inside code repository inside 72 hours.

Researchers at cybersecurity startup Hacktron chained an image-processing vulnerability with a flaw in OpenAI’s identification infrastructure in July to realize entry to a number of staff’ ChatGPT and Codex accounts.

One compromised Codex account was linked to OpenAI’s GitHub group, giving the researchers a path into the corporate’s inside software program surroundings.

The crew stopped after instructing the compromised worker’s Codex account to create a innocent pull request inside OpenAI’s personal openai/openai monorepo. Hacktron mentioned the researchers didn’t examine proprietary supply code.

This week, Hacktron disclosed the vulnerabilities and ended additional testing.

OpenAI reportedly fastened the identity-side flaw roughly 14 hours after receiving the report and later paid the corporate a $6,500 bounty.

Anthropic’s Opus 5 cleared a hurdle its predecessor couldn’t

The OpenAI attack accelerated after Anthropic launched Claude Opus 5, which overcame an exploitation hurdle that its predecessor had repeatedly failed to unravel.

Hacktron started inspecting the image-upload pipeline utilized by OpenAI’s Discourse group discussion board on July 23. HEIC and HEIF recordsdata have been processed via ImageMagick and the underlying libheif decoding library, giving attacker-controlled pictures a path into weak code.

The researchers equipped Claude Opus 4.8 with a Discourse Docker picture and requested it to examine the put in libheif package deal for safety weaknesses. The mannequin recognized lacking fixes that left a heap buffer overflow, enabling out-of-bounds reads and writes.

By July 24, Opus 4.8 had produced an exploit that achieved code execution when handle area structure randomization (ASLR) was disabled. But repeated makes an attempt to make the exploit work reliably towards Discourse’s regular configuration with ASLR enabled failed.

Anthropic launched Opus 5 later that day, giving the researchers one other route.

Related Reading

How a fake AI supercomputer stole $24 million from hundreds of crypto investors


Hacktron opened a recent session with the brand new mannequin, which produced a working ARM64 exploit for an area Mac inside about three hours. The researchers then requested it to adapt the exploit to the x86-64 structure and jemalloc reminiscence configuration utilized by Discourse.

By 6 a.m. on July 25, the crew had a working exploit that might execute code via a malicious picture add.

With that foothold established, the researchers subsequent examined whether or not Claude might reproduce the assault towards a distant surroundings with much less human intervention.

Hacktron positioned the mannequin in an autonomous loop towards its personal Discourse Cloud occasion. The firm mentioned Claude initially refused to develop an exploit instantly towards a distant system, prompting the crew to proxy the check surroundings so it resembled a capture-the-flag safety problem.

Four hours later, the agent had reproduced the assault towards the distant check surroundings.

The researchers then used the ensuing exploit towards OpenAI’s group discussion board, the place they gained administrative entry. A separate weak point in OpenAI’s single-sign-on system allowed them to maneuver from the discussion board into ChatGPT and Codex accounts.

One compromised worker had linked Codex to OpenAI’s GitHub group, creating the trail the researchers later used to show entry to the corporate’s inside repository.

Hacktron co-founder (*72*) “s1r1us” Pedhapati said the episode confirmed how shortly AI was compressing exploit-development timelines that after required much more specialised labor.

He mentioned:

“Our essential takeaway from hacking OpenAI: AI is lowering the quantity of scarce experience wanted to develop exploits. Work that after took months can now take days. Even main AI labs will be weak.”

However, Hacktron confused that the operation nonetheless relied on skilled human researchers. The firm famous:

“This was not utterly autonomous hacking, and expert human steerage remained vital.”

Robert Reith, founding father of blockchain safety agency Accretion, said skilled researchers nonetheless equipped a lot of the judgment wanted to show AI-generated work into a successful attack, however warned that the benefit might erode as fashions enhance.

According to him:

“There’s nonetheless a big hole between what expert researchers + AI can do vs. normal inhabitants + AI. The scary half is that this hole might turn out to be smaller as AI absorbs this data and instinct over time.”

AI Coding brokers develop the blast radius of a compromised account

The similar coding brokers that accelerated the exploit additionally elevated its potential attain as soon as the researchers gained management of an OpenAI worker account.

ChatGPT and Codex can connect with exterior companies, which means a compromised account might expose no matter integrations a person has licensed. Hacktron cited GitHub, Slack, and e-mail as companies that might turn out to be reachable, relying on an account’s configuration.

In this case, the worker’s GitHub connection supplied the trail into OpenAI’s inside repository.

Security brokers warned that this focus of permissions round (*3*) might make them more and more engaging targets as Codex, Claude Code and related agents turn out to be extra deeply embedded in company improvement workflows.

Codey Blakeney, analysis lead at Arcee, mentioned:

“The extra widespread Codex and Claude Code get, the extra individuals are going to try to goal them.”

Blakeney said the danger might develop if software program improvement turns into concentrated round a small variety of AI suppliers, creating broader factors of failure throughout engineering groups.

He famous that if regulation strikes us to fewer gamers, it means much less selection and extra single factors of failure. Blakeney added:

“The total means software program engineering works at most locations has utterly modified with coding brokers, and if only one firm has a nasty day, it’s going to mess up your roadmap and timelines.”

Maxime Fournes, CEO of AI security advocacy group PauseAI, said the breach additionally highlighted a longstanding imbalance between attackers and defenders that might turn out to be extra consequential as AI lowers the price of creating subtle exploits.

According to him, attackers want to search out one neglected weak point, whereas defenders should safe a wider assault floor. He famous:

“It’s massively tougher and costlier to defend towards all potential flaws than to take advantage of a single one.”

OpenAI tightened entry after the disclosure, whereas Discourse ready a patch by July 27 and added additional sandboxing round its image-processing system.

The submit Anthropic’s Claude helped 3 researchers breach OpenAI in under 72 hours appeared first on CryptoSlate.

Similar Posts