Skip to content

Building Trust in AI Is Critical to the Frontier’s Future

Voluntary AI lab commitments are a start, but genuine security and safety will require enforceable standards and independent oversight.

Demonstrators carry a large banner with the words 'Stop the AI race' as they participate in the Stop the AI Race protest march in San Francisco, California, on July 11, 2026. The march included stops outside the offices of OpenAI, Anthropic, and Google DeepMind.
Demonstrators participate in the Stop the AI Race protest march in San Francisco, California, on July 11, 2026. The march included stops outside the offices of OpenAI, Anthropic, and Google DeepMind. Karl Mondon/AFP via Getty Images

By experts and staff

Published
  • Vinh NguyenCFR Expert
    Senior Fellow for Artificial Intelligence

Vinh X. Nguyen previously served as the National Security Agency’s chief artificial intelligence (AI) officer.

The summer began with some familiar crises—climate-related shocks, contested borders, renewed conflict in the Middle East—and ended with a novel one: the not insignificant chance that AI could destroy humanity.

In just ten weeks, there have been reports of multiple autonomous AI agents breaking out of their containers and hacking external systems to cheat on their tests; an AI-powered cyber capability that could affect the lives of a billion people; growing public resistance to AI data centers and mistrust at an all-time high; open letters from the AI community warning of emerging risks; and a deep recognition that no entity has been able to master and oversee a powerful autonomous capability that can reason, deceive, collude, and, when blocked, hack its way out.

The White House’s response is to downplay the risks, arguing that slowing down will give China an advantage in the AI race. But on September 12 there was unusual public agreement among the leaders of the top U.S. AI labs—Anthropic, OpenAI, SpaceX, and Google DeepMind—that the industry really does need to slow down, or “pace,” development to ensure adequate oversight, safety, and security controls.

Anthropic and OpenAI committed to providing independent third-party evaluators with employee-level access to assess the security and safety risks associated with model development and behavior. They have also said they will allow reviewers to publish findings without the companies’ editorial control, with only limited redaction of sensitive or legally privileged.

But while these early moves are welcome, there is more that AI companies themselves—as well as policymakers and sensitive industries—can and should be doing to improve safety and build public trust at this critical juncture in the development of AI. This includes:

AI companies disclosing funding and terms for third-party reviewers. In many cases, AI companies paid the same reviewers for testing and evaluation services and lured talent from these outfits by offering higher cash and equity compensation. AI companies should provide financial backing for a third-party nonprofit grantmaker to fund research, talent, payments, and certification for the independent evaluator ecosystem. If the companies plan to second their experts to independent evaluators, the experts should operate under clear conflict-of-interest rules and practices.

The independent evaluation ecosystem is nascent and needs dependable access, safe harbor for evaluators, measurement research and infrastructure, and competitive incentives to attract talent. Without independent funding and guarantees of access, the public and U.S. allies will not see the evaluation system as independent or legitimate, regardless of how well intentioned the arrangements are.

Congress authorizing an AI self-regulatory organization (SRO). Such an entity would provide a legislative backstop that legitimizes the role and enforcement authority of independent evaluators and enhances collaboration among companies to address security and safety challenges. Congress could assign this responsibility to the U.S. Commerce Department, in collaboration with the Departments of the Treasury and Defense and the intelligence community.

The White House operates a “voluntary” process for evaluating advanced AI for cyber risks but has refused to disclose much detail, leaving frontier labs, Congress, U.S. allies, and the public struggling to assess its utility, rigor, and fairness. Establishing an SRO with authority to organize, fund, certify, set rules, respond to incidents, and enforce standards for private AI companies’ development would help restore stability, independence, and legitimacy to AI development. It would also help close information asymmetry about AI capabilities and risks among AI companies, buyers, regulators, U.S. allies, and the public—enabling evidence-based debate and decision-making. Frontier labs are open to working on the details but need direction from the White House and Congress to fully realize this effort.

Policymakers and labs addressing the challenge of open-weight models. Pacing the frontier is no longer limited to controlling a few frontier labs’ advanced capabilities; it also depends on managing the diffusion and deployment of this technology. Labs can slow down as much as they want, but enterprises and institutions continue to explore mostly Chinese-origin, open-weight models (systems that are publicly downloadable, allowing anyone to inspect, modify, and run the software locally) to inform and make consequential, high-stakes decisions.

Throughout the Council on Foreign Relations’ LEAD AI summits and discussions this year, executive leaders across federal and state governments, Wall Street, frontier labs, technology, academia, and civil society groups noted that they share common interests in harnessing the technology to its fullest but felt they were unable to meaningfully move the needle toward AI security and safeguards. They want assurances of controllability—ensuring the technology is reliable, secure, confidential, and safe—to build confidence among their boards, regulators, C-suite peers, clients, and partners. They judged that controllability, affordability, and accountability are key to unlocking economic value and breaking down barriers to AI deployment, and that assurances across closed- and open-weight capabilities are a competitive advantage for American AI.

Pacing development is not solely for frontier labs and the government to resolve. The financial, health-care, life sciences, cybersecurity, and technology industries can now take the lead in pacing their own development and deployment.

Key industries stepping up. Financial, health-care, and life sciences firms are three of the largest purchasers of AI within regulated industries. They should require minimum viable security and safety assurances as they integrate AI into critical, high-stakes workflows. These include observability and data security tools (such as role-based controls, customer-managed encryption keys, and an interoperable standard for prompt-completion and chain-of-thought logging), agentic identities that are durable and accountable, and guarantees that user and enterprise data remain confidential. Frontier AI companies have already committed to these efforts, and buyers should hold them accountable. Buyers should signal this same demand to the open-source, open-weight ecosystem, driving those companies that want to serve open-weight models to distinguish themselves through security and safety advantages, not just pricing. AI systems with few security safeguards can take unsanctioned actions and leak confidential data, creating unpriced liabilities for enterprises over time.

AI and cybersecurity industries coordinating to ensure the security and safety of open-weight models for everyone. The Open Secure AI Alliance, led by Nvidia and comprising more than one hundred companies, should take on the most obvious task: establish a reputable, dedicated team to identify and mitigate vulnerabilities in widely used open-weight models, regardless of origin; retrain and minimize any embedded censorship; and integrate new security, controllability, and safeguard features to prevent their use to cause cyber, biological, and autonomy harms.

AI companies, the Frontier Model Forum, U.S. government entities, national labs, and academic institutions should collaborate with this alliance to share best practices and techniques to safeguard advanced AI models. If U.S. companies can turn open-weight models, regardless of origin, into more reliable, secure, and safer versions and make them available on platforms such as Hugging Face—a central, trusted repository of open-weight models for software developers—then the alliance would have fulfilled its mission.

The call to pace AI frontier development is a positive step toward a future in which trust is built into AI systems, the actors who deploy them, and the science behind their claims. Developing a trust infrastructure is not only the right thing to do in terms of averting the kind of crisis that AI leaders have been warning about, but it would also provide a competitive advantage that the United States’ adversaries cannot match.

This work represents the views solely of the author(s). The Council on Foreign Relations is an independent, nonpartisan membership organization, think tank, and publisher, and takes no institutional positions on matters of policy.