
Weak AI safety regulations may backfire and produce products that are potentially more dangerous than those created under no regulation, according to a new study published Monday in the Proceedings of the National Academy of Sciences. The research, led by Benjamin Laufer and colleagues from Cornell University and Carnegie Mellon University, uses theoretical economics and game theory to demonstrate how the structure of AI regulation matters more than its existence.
The central finding is that for true safety, regulation must be strict and target the companies that develop AI models—such as OpenAI, Google, and Anthropic—rather than focusing solely on downstream companies that apply the technology in real-world settings, like providers of AI medical diagnostic systems or e-commerce customer service chatbots. While regulating specific use cases might seem logical at first glance, the authors argue that doing so can backfire and reduce the overall safety of AI products.
The reason is rooted in what the researchers call “free-riding behavior.” When governments regulate downstream companies but let general-purpose AI model developers off the hook, those developers tend to cut corners on safety measures like third-party audits. They assume that downstream companies will ensure the safety of the final product, shifting the burden of safety oversight away from the place where the underlying model is created. As Laufer explained, “The regulation acts as a tool for the general provider to offload the safety burden onto the downstream specialist.”
Understanding the Game Theory Model
Game theory, the study of strategic decision-making where outcomes depend on the choices of multiple actors, provides the foundation for this analysis. The researchers constructed a model of the AI supply chain with two main players: general-purpose AI producers (like large frontier labs) and downstream domain specialists (such as companies building AI tools for healthcare or customer service). Each player can choose to invest heavily in safety or to minimise safety investments to save costs. The model shows that when both players face strict regulation demanding adequate safety investment, the outcome is optimal for everyone—better safety and better overall utility, defined as revenue minus investment cost.
But when regulation is weak or only targets downstream companies, a different dynamic emerges. General-purpose producers anticipate that downstream specialists will be held accountable for the end product, so they reduce their own safety efforts. Downstream specialists, meanwhile, may not have the technical capability or full visibility to thoroughly audit the foundation models they rely on. The result is a product that is less safe than if no regulation existed, because the division of responsibility creates gaps that neither party fully covers.
The Current Regulatory Debate in the United States
The study arrives at a moment when the U.S. government and Silicon Valley are intensely debating how artificial intelligence should be regulated. Two broad camps have formed. On one side are anti-regulation technologists who advocate for much lighter federal guardrails, largely aligning with the Trump administration’s approach to AI governance. This group argues that the AI industry should be free from unnecessary constraints to innovate as quickly as possible—and that this is the only way the United States can win the global AI race against China.
The self-proclaimed pro-innovation group often characterises supporters of stricter AI safety regulation as “doomers” at best or as attempting regulatory capture at worst. Supporters of stricter regulation counter that the AI industry, in pursuit of larger profit margins, is underestimating or downplaying the risks of under-regulated AI development. The list of purported downsides includes everything from AI-generated misinformation and “AI psychosis” to the community health consequences of data centres and a widely feared unemployment crisis that could follow broader AI adoption.
Yet the authors of the new study argue that safety versus revenue does not have to be an either-or choice. Their model suggests that stronger, well-placed regulation can benefit all players simultaneously, improving both the safety of the end product and the utility that general-purpose AI creators and downstream domain specialists derive from their investments. The key is to design regulation that demands sufficient investment from both types of actors, so that neither can rely on the other to shoulder the safety burden.
The Prisoner’s Dilemma of AI Safety
The situation described in the paper is a classic example of a prisoner’s dilemma, a foundational concept in game theory. In this problem, two rational decision-makers are given the option to cooperate or betray each other. If they both cooperate, they achieve the best possible collective outcome. But because neither knows what the other will choose, each fears being the one who cooperates while the other betrays—a scenario that yields the worst possible outcome for the cooperator. Rational actors therefore often choose to betray, ensuring their own benefit but guaranteeing a worse outcome for everyone than if they had cooperated.
In the context of AI regulation, cooperation means both general-purpose AI producers and downstream companies making meaningful investments in safety—such as third-party audits, red-teaming, bias testing, and transparency measures. Betrayal means cutting corners to save money, hoping that the other party will pick up the slack. Without strong coordination enforced by regulation, the rational choice for each actor is to under-invest, leading to a collectively unsafe outcome.
Strict regulation that covers the entire AI supply chain changes the incentive structure. It assures each actor that others are also required to invest in safety, reducing the temptation to free-ride. When both sides are confident that their competitors and partners are under the same obligations, cooperation becomes the rational choice, and the most ideal outcome for all becomes attainable.
Why Weak Regulation Backfires
The study’s most surprising conclusion is that weak regulation can actually be worse than none. Without any regulation, general-purpose AI developers might still invest in safety for reasons of reputation, liability, or voluntary commitment. But when weak regulation is introduced—especially if it only targets downstream applications—it can create a false sense of security while enabling developers to externalise safety costs.
For example, if a government mandates that AI medical diagnostic systems must receive regulatory approval, a hospital deploying such a system may be held responsible for ensuring its safety. The hospital, in turn, may demand safety assurances from the AI developer. But if the developer knows that the legal burden falls on the hospital, it may provide minimal documentation and avoid deeper safety testing. The hospital, lacking the technical expertise to independently audit a complex neural network, may accept these assurances. The result is a medical AI system that passes regulatory scrutiny but is actually less safe than one developed under no regulation, where the developer might have been more cautious to avoid liability.
The researchers emphasise that the problem is not regulation itself, but its design. Policies that focus on end-user industries without addressing the source of AI capabilities are doomed to fail because they ignore the incentives of the most powerful actors in the supply chain. Instead, regulation must be “strict and horizontally comprehensive,” covering model developers, infrastructure providers, and downstream deployers alike.
Implications for Policymakers
The findings carry significant implications for policymakers at both the federal and state levels. While some states have enacted their own AI laws, these often focus on consumer-facing applications such as hiring algorithms or deepfake disclosures, leaving model developers largely unregulated. The new research suggests that such piecemeal approaches may be counterproductive, as they create opportunities for developers to shift responsibility to less powerful downstream actors.
Federal efforts remain in flux. The Trump administration has advocated for light-touch regulation, while numerous congressional proposals have been introduced but not enacted. Meanwhile, the European Union’s AI Act takes a risk-based approach by imposing obligations on both providers and deployers, though its final rules are still being phased in. The new study suggests that other jurisdictions should learn from this example and design rules that require safety investments at every stage of the AI value chain.
Coordination is also crucial on the international stage. Because AI development is global, a country with strict regulation may simply push AI development elsewhere, unless there is some form of international agreement. The authors argue that the logic of the prisoner’s dilemma extends to nations as well: if all countries cooperate to enforce safety standards, everyone benefits; if one country defects to attract AI investment, the entire global safety ecosystem suffers. This makes international alignment on AI safety standards a pressing collective action problem.
At the same time, the researchers are careful to note that their model is a simplification. Real-world regulations involve courts, agencies, technical standards, and political compromise. But the underlying incentives they identify are robust: when the burden of safety can be shifted to another actor in the supply chain, it will be shifted, unless regulation prevents it.
As Laufer put it, “People think of AI as a single object, but actually AI involves a very complicated set of stakeholders and actors that each have their own contributions to the technology. To regulate in a thoughtful way, we need to consider the whole supply chain, not just a single provider or entity.”
That is the central lesson of the study: AI safety is not a one-dimensional problem, and neither is its solution. Effective regulation must be strict, comprehensive, and designed with an understanding of the strategic interactions that shape the behaviour of every player in the AI ecosystem. Anything less may not just fail to protect the public—it may make the world less safe than doing nothing at all.
Source:Gizmodo News
