Trump’s AI Safety Pact Is Toothless. But There Is a Path Forward.
The White House AI accord outlines sensible safeguards, including independent evaluators and board-level oversight, but does not compel companies to adopt them. Until the cost of safety failures outweighs the payoff of racing ahead, a pledge built on “believe” and “should” won’t slow anyone down.

By experts and staff
- Published
Connor MartinCFR Expert2026–27 International Affairs Fellow in National Security, sponsored by Janine and J. Tomilson Hill
Connor Martin previously served as deputy director of the Committee on Foreign Investment in the United States (CFIUS) at the Treasury Department. His work has focused on the national security implications of cross-border capital flows and on the national security risks posed by artificial intelligence (AI) and other frontier technologies.
My team at CFIUS spent long hours negotiating agreements with private companies to compel them to do—or not do—certain things. The committee has the legal authority to mitigate national security risks that could arise from transactions involving foreign investment in the United States.1 By law, any such mitigation agreements must be (1) effective, (2) verifiable, and (3) monitorable and enforceable.
As a framework, that approach is useful for assessing the “White House Accord on Super Intelligence,” a document signed on Tuesday by President Donald Trump and leaders of six of the top U.S. AI companies—Anthropic, OpenAI, Google, Nvidia, Meta, and xAI. Trump commented that the agreement would involve “a tremendous self-policing aspect,” is “morally binding,” and that “we’re going to have it be nice and safe.”
Let’s hope. Unfortunately, there is nothing in the document that legally compels any of the signers to change their behavior. More important, nothing alters their commercial incentives to keep testing frontier capabilities. The document does contain promising ideas—but it leaves to others the work of creating conditions that are effective, verifiable, monitorable, and enforceable.
A Conspicuously Noncommittal Agreement
The one-page agreement that came out of the White House meeting describes itself as a “Joint Commitment on Frontier Responsibilities,” but it does not actually compel the companies to do anything. It is closer in spirit to the “voluntary AI commitments” that Joe Biden’s administration secured in 2023 (though less detailed than those). It does not have the force of law or executive order, and the language itself is conspicuously noncommittal—with a revealing exception.
The document states, “[W]e believe each company should implement… four layers of controls and audits[.]” “Believe” and “should”—two words that any government lawyer would immediately strike from a CFIUS agreement, because they are unfalsifiable. Better words are “will” or “shall”—but “will” only appears at the end of the document, when the companies commit that they “will meet regularly to establish standards and best practices” on safety.
This is notable because one of the recommendations in Anthropic CEO Dario Amodei’s “pacing the frontier” essay—which accelerated the debate over AI safety in September—is for the U.S. government to waive antitrust restrictions so that the labs can work together on safety. It is unclear whether Trump’s signature on this one-pager amounts to an endorsement of that recommendation, but it could plausibly be read as a signal that federal antitrust enforcement would look the other way (and in any case, discussions are reportedly already ongoing between OpenAI, Anthropic, and Google DeepMind).
Commercial Incentives Unchanged
The nonbinding nature of the agreement also means it does nothing to change the commercial incentives for the companies to keep testing models whose capabilities are potentially dangerous.
The “pacing the frontier” debate is about two questions that look the same but aren’t:
- Should the world’s top AI companies slow down the training of their best models to give them time to get better at safety?
- Will they?
Even if their leaders earnestly believe the answer to the first question is yes, the answer to the second depends on whether the cost of safety violations can be made to exceed the cost of sprinting ahead. That creates a business decision about where to direct finite resources—money, employees, and computing power.
To date, the companies have grown by investing those resources in making their best models (“the frontier,” in industry-speak) even better at achieving tasks that consumers and enterprises want. “Pacing” means redirecting those resources instead to operational improvements (such as data hygiene) and to different kinds of model training and evaluation, focused not on making the models more effective but on making their actions more compliant ( or aligned) and their reasoning more transparent (or interpretable).
These are challenging engineering problems that will be expensive to solve. Reallocating resources toward them only makes business sense if the next dollar spent on safety earns as much as the next dollar spent on capability. Anthropic’s leaked IPO filing showed a 2025 operating loss of $8 billion, even as revenue grew by more than 1,000 percent to $4.6 billion, according to Reuters. On top of that, huge amounts of Anthropic’s cash are locked into infrastructure commitments. OpenAI, meanwhile, has projected negative free cash flow of nearly $280 billion from 2026 through 2030.
Anthropic and OpenAI are squeezed between the hyperscalers, which own the compute they currently rent and are developing their own very capable models, and open-weight models, which are free and good enough for most use cases. For these two companies, the promise of frontier capabilities is their product.
Of the other signers, Google and Meta already own significant compute, and they have other profitable business lines (and xAI is cushioned by its merger with SpaceX). But Google, Meta, and xAI are also reporting massive capital expenditures on AI infrastructure, and all three compete with OpenAI and Anthropic for the same enterprise relationships and consumer loyalty. If OpenAI and Anthropic don’t stop testing at the frontier, why would they?
[Video: https://vimeo.com/1219226879]
The incentives don’t stop there. The sector and arguably the entire U.S. economy are systemically exposed to the risk that model capabilities plateau. AI companies have swamped global credit markets, with over $500 billion in debt issued to the sector this year alone. They have also engaged in circular financing, in which customers finance suppliers and vice versa. The Bank for International Settlements issued a bulletin on Thursday about the “exceptional” scale of these arrangements in the industry, “potentially amplifying contagion” in the event of a shock.
This financial architecture was built for a world in which AI model capabilities continue improving rapidly. To slow that rate of improvement without collapsing valuations, investors need to believe that safety is worth the delta.ar liability for safety violations does not change that calculus.
Potential Ways Forward
For all that, the agreement’s “four layers of controls and audits” do provide the kernel of a path forward on AI safety—if the companies can be incentivized to adhere to them.
Proposals 1 and 2 call for “robust internal controls” and an “empower[ed]… internal team,” respectively, each of which is supposed to “ensure” that the models are not misused or do not behave in unintended ways. Presumably, all these companies already have such controls and such teams, which have not prevented the reported instances of AI breaking from its training and evaluation environments.
Proposals 3 and 4 are more promising.
Proposal 3 calls for the companies to “[p]artner with an independent external auditor or evaluator to carry out independent assessments of whether the controls, monitoring, and detection are operating as intended.” The independent evaluator approach is consistent with some of the mitigation measures that the U.S. government has used in CFIUS (e.g., “security officers” and “third-party monitors”). Anthropic and OpenAI have already committed to embed such evaluators in their companies, and embedded evaluators have been endorsed by many leading AI scientists. But the proposal leaves out crucial details for how to make them effective. For instance, it doesn’t say how evaluators would be kept meaningfully independent, protected from retaliation, and granted full access privileges.
Proposal 4 calls for the companies to “[d]esignate an independent committee of the board of directors to oversee and receive reports from the teams operating the controls and the internal and external auditors and evaluators, as well as to ensure any issues identified are remediated.” This has the virtue of identifying specific individuals at the corporate governance level with visibility into AI safety issues and responsibility for fixing them. But it fails to specify what would make the committee “independent,” how it would be constituted, or how it is supposed to ensure any issues are remediated.

Game theory dynamics in AI—not just among the U.S. labs, but also between the United States and China—incentivize racing ahead to dominate the technology. Research by economists suggests that to truly “pace the frontier,” all participants would have to share information transparently. The agreement’s strongest commitment, that the companies “will meet regularly to establish standards and best practices,” points toward this possibility. But, like the rest of the document, it lacks any conditions that would make that commitment verifiable and monitorable, or any enforceable penalties for violating it.
Making Safety Stick
In CFIUS cases, the government is well-practiced at establishing clear requirements and clear lines of responsibility: Company X is required to take Action Y; if it doesn’t, somebody is on the hook. The “Accord on Super Intelligence” lacks any such clarity and enforceability. It does not compel any company to take any action, and it doesn’t specify who is on the hook for any violations.
This omission is the more striking because Treasury Secretary Scott Bessent, an essential figure in the Trump administration’s approach to AI, has forcefully asserted that the government will not give AI companies a “liability shield.” Although the agreement alludes to some possibility of “laws or regulations” in the future, there is no indication that anything meaningful is on the way—certainly not before the next Congress.
It should not come as a surprise, then, if the first real teeth for any slowdown come from the courts. Nothing prevents Hugging Face, for example, from suing OpenAI for being hacked by its agents this summer. Regardless of whether such a suit were ultimately successful, it would be reasonable to expect model testing to slow while the industry watched it play out.
For now, the commercial pressures on the leading AI companies continue to incentivize them to push the boundaries of model capabilities. Regardless of how it happens, the only way to change that is to raise the cost of safety violations. Establish those incentives, and tremendous self-policing really might follow.
This work represents the views solely of the author(s). The Council on Foreign Relations is an independent, nonpartisan membership organization, think tank, and publisher, and takes no institutional positions on matters of policy.