OpenAI Staff Blame Rush to Ship for Rogue Agent Hack

요약:OpenAI employees told Wired that competitive pressure to ship models and products has hindered safety, security, and alignment. In May, OpenAI's GPT-5.6 Sol and an unnamed pre-release model escaped an internet-restricted testing environment by exploiting a flaw, then breached Hugging Face to obtain cybersecurity test answers—what a former employee called the biggest safety incident in the company's history. OpenAI has slowed research, reassigned teams, and spent millions investigating. President Greg Brockman said safeguards would be strengthened, while former alignment chief Jan Leike noted safety had "taken a back seat." Safety advisory co-leader Boaz Barak called for cultural change rather than just fixes. The report follows months of leadership departures.

In brief

  • OpenAI employees told Wired that pressure to release models and products has made it difficult to prioritize safety, security, and alignment.
  • A former employee called the breach the largest safety incident in OpenAIs history.
  • OpenAI has slowed research, reassigned teams, and spent millions investigating the failure.

OpenAIs rush to release new models and products contributed to conditions that allowed its AI agents to escape internal testing environments and hack Hugging Face earlier this year.

Multiple current and former employees told Wired that competitive pressure has made it difficult for staff to devote enough attention to safety, security, and alignment—the work of ensuring AI systems behave as intended.

Myriad: When will OpenAI release GPT-6? Click to make your prediction.

“They were incredibly sloppy. If you‘re serious about this, your AI shouldn’t be able to break out onto the internet and then do it again right afterward,” a former OpenAI employee told Wired. “This was the biggest safety incident in OpenAIs history.”

In May, OpenAIs GPT-5.6 Sol and an unnamed pre-release model escaped an internet-restricted testing environment by exploiting a previously unknown software flaw. The agents then breached the open-source AI repository Hugging Face to obtain answers to their cybersecurity tests. In July, OpenAI confirmed that its models were responsible, before giving a fuller breakdown at the annual Black Hat conference last week.

OpenAI President Greg Brockman said the company is strengthening its safeguards as its models become more capable.

“Were reaching new levels of model capability that require more robust training, alignment, safety and security testing, deployment practices, and governance,” Brockman told Wired.

Employees have raised similar concerns before, including Jan Leike, OpenAIs former head of alignment, who left for rival AI developer Anthropic in 2024 after warning that safety had “taken a back seat” to product development.

“Building smarter-than-human machines is an inherently dangerous endeavor,” Leike warned. “But over the past years, safety culture and processes have taken a backseat to shiny products.”

Boaz Barak, co-leader of OpenAIs safety advisory group, wrote on X that addressing the latest failure would require “not just fixing some issues but also changing our culture.”

The report comes amid months of leadership turnover at OpenAI.

In April, head of OpenAIs video generator project Sora, Bill Peebles, former chief product officer and science chief Kevin Weil, and enterprise applications technology chief Srinivas Narayanan left the company. July brought the departures of product and business chief Fidji Simo, safety leader Sandhini Agarwal, chief futurist Joshua Achiam, and AI ethics lead Chloé Bakalar. Safety systems chief Johannes Heidecke also departed after OpenAI merged its safety and core research teams.

Earlier this week, OpenAI Chief Operating Officer Brad Lightcap announced his departure after eight years to start a new venture.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.

Your Email

Get it!

Get it!

면책 성명

본 기사의 견해는 저자의 개인적 견해일 뿐이며 본 플랫폼은 투자 권고를 하지 않습니다. 본 플랫폼은 기사 내 정보의 정확성, 완전성, 적시성을 보장하지 않으며, 개인의 기사 내 정보에 의한 손실에 대해 책임을 지지 않습니다.
전편

토큰화 주식 급락, SEC 지연으로 암호화폐의 월가 진출 ‘속도 저하’

다음

샌디스크, 마진 하락 언급…SNDK 14% 상승