OpenAI Chief Scientist Warns AI Labs May Need to Slow Down

요약:Jakub Pachocki is calling for mandatory safety standards as OpenAI finds it harder to monitor advanced AI models reasoning.

In brief

  • OpenAIs chief scientist called for voluntary slowdowns until shared safety standards are established.
  • Pachocki said monitoring models reasoning is becoming less reliable.
  • He urged international coordination as AI takes on more of its own development.

OpenAI‘s chief scientist Jakub Pachocki has called for voluntary slowdowns in AI development, warning that no lab’s safeguards are adequate to keep building more powerful systems at full speed for much longer.

In his post “An Alien Mind,” published Sunday, Pachocki argued that voluntary company commitments should become mandatory safety standards, enforced by independent auditors, governments or international bodies. He said OpenAI would withhold further scaling when needed but did not announce a new pause.

Myriad: How high will Tesla stock go? Click to make your prediction.

“Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” he wrote. “I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established.”

Pachocki, who joined OpenAI in 2017, also defended developing more powerful AI to secure infrastructure and protect against rogue agents, while warning against using those threats to justify reckless development.

“The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes,” he wrote.

He also referenced OpenAIs Hugging Face breach, where AI agents working on cybersecurity evaluations escaped their testing environment and attacked the company. According to OpenAI, the agents established covert communication channels and rebuilt them after researchers intervened.

An independent investigation by METR found that roughly 1,200 agents coordinated on an unauthorized message board, with about 700 joining the attack. This incident, Pachocki said, is why AI safeguards must hold even when models believe no one is watching.

“Crucially, we need future AIs to continue to hold human values regardless of whether they believe theyre under human supervision,” he wrote.

In research published last year, OpenAI found that penalizing models for expressing intentions to cheat could teach them to conceal those intentions while continuing to cheat.

AI models have since become more capable at finding and exploiting software flaws: OpenAI classified Astra at its highest cybersecurity risk tier, while Anthropic said Mythos Preview discovered thousands of previously unknown vulnerabilities across major operating systems and browsers.

Citing recent incidents of AI systems escaping human control, Sen. Bernie Sanders (I-Vt.) and Rep. Greg Casar (D-Texas) announced the forthcoming Ban Artificial Superintelligence Act on September 3. The proposal would pause advanced AI development until a new federal regulator establishes safety rules and permanently ban the development and deployment of superintelligent AI.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.

Your Email

Get it!

Get it!

면책 성명

본 기사의 견해는 저자의 개인적 견해일 뿐이며 본 플랫폼은 투자 권고를 하지 않습니다. 본 플랫폼은 기사 내 정보의 정확성, 완전성, 적시성을 보장하지 않으며, 개인의 기사 내 정보에 의한 손실에 대해 책임을 지지 않습니다.
전편

영국 금융감독당국, 예측 시장 금지 해제 검토: 보고서

다음

룰루레몬 주가 80% 폭락…창업자 이혼