OpenAI, Google, Meta and three other technology companies agreed Tuesday to bring in outside auditors to check their AI safety controls, signing a voluntary White House pact that carries no penalties if they fall short.
President Donald Trump called the agreement "morally binding." It has no enforcement mechanism and does not require companies to publish or name the auditors.
"And they understand that they have to self-police," Trump told reporters after the meeting. He said he would set up a 10-member board to oversee AI safety and appoint a new White House official to lead AI policy, and the accord says its measures could eventually be written into law.
Anthropic, Nvidia and Elon Musk's xAI, which is now part of SpaceX, also signed the Sept. 29 agreement. OpenAI was represented by President Greg Brockman, alongside Google's Sundar Pichai, Meta's Mark Zuckerberg, Anthropic's Dario Amodei and Nvidia's Jensen Huang.
The one-page document asks companies to monitor their most capable models during training and use, including whether they could enable cyberattacks or biological and chemical threats. It specifically calls for controls to prevent models from hacking or accessing computer systems in unintended ways.
An internal team would check that those protections work and that problems are fixed. An independent auditor would assess the controls, while a committee of each company's board would receive the findings and oversee fixes.
That would give outside reviewers a role in checking the safeguards companies rely on to keep experimental models contained.
Still, the agreement leaves the choice of auditors with the companies and sets no deadline for implementing the measures. Some of the steps are ones the companies already take in some form, according to AP.
Read More: OpenAI says its new 'Astra' AI can build attacks without human help
How AI agents are playing havoc
The accord follows a string of incidents in which experimental AI agents broke into computer systems they were never cleared to access, including OpenAI test agents that reached servers run by Hugging Face, a platform where developers share AI models.
An OpenAI agent also accessed an Australian government Medicare portal on June. 18, which the company disclosed to Australian authorities only in September.
Closer to crypto, AI has been suspected in some of the year’s biggest security scares.
In July, attackers began sweeping bitcoin from Coldcard hardware wallets through a five-year-old firmware flaw, taking 1,367 BTC worth nearly $89 million from 4,500 addresses across three instances. Coinkite, which makes the wallet, later said it believed someone used frontier AI to review its public code, though that has not been proven.
coindesk.com