A Cornell team models who carries the safety bar in a two tier AI stack: the foundation model lab that trains the underlying AI, and the vendor that wraps it for end users. A weak bar on the lab can leave users worse off than no rule at all.
When a regulator tries to protect the public by aiming AI safety rules at the company closest to the user (the chatbot vendor, the AI tutor, the medical-diagnostic integrator), the result can be a system less safe than if the regulator had done nothing. A peer-reviewed PNAS paper by Benjamin Laufer and colleagues at Cornell, built on a game-theoretic model, shows why.
The model has two players and one rule-setter. A "generalist," the kind of foundation-model lab that produces a base model, moves first. A "specialist," a downstream deployer that adapts that model for a customer-facing product, moves second. The regulator sets a minimum safety standard with non-compliance penalties, and both players split revenue. Each side can invest in safety and performance.
The problem appears when the regulator sets a weak or zero bar for the generalist. The generalist knows what the specialist will be required to do, and knows its own fine is small. It free-rides: under-invests in safety and offloads the work onto the specialist. The specialist cannot fully compensate, and equilibrium safety ends up below the no-regulation baseline, the point where, by the model's assumptions, both players invest in safety on their own. SingularityHub's coverage frames it bluntly: rules aimed at the wrong tier can leave the final product less safe than no rule at all.
In real-world terms, the generalist is the OpenAI-style base-model provider; the specialist is the customer-service chatbot vendor, the AI tutor, or the medical-diagnostic systems integrator that wraps that base model for a specific use. The paper does not name a company. It does not collect deployment data. Its evidence is the equilibrium of the model itself, under the assumptions spelled out in the arXiv preprint.
The finding lands while US state and federal AI bills are still being argued, and while EU AI Act implementing rules are being negotiated. The US, as the Cornell Chronicle notes in its write-up of the paper, has no strong federal AI legislation, so states are producing a patchwork. Bills in that mix vary in which tier (base-model provider or downstream deployer) carries the enforceable safety duty, and what bar the other tier must clear. The model says that choice is not neutral: a low bar at one tier with a high bar at the other can build a less-safe outcome into the design.
The constructive lesson is narrow. Clean design means one of two shapes: either a single clearly-liable party that the rest of the stack can rely on, or symmetric, well-set bars at both tiers. Anything in between, one tier high and the other low or zero, is the trap the paper names. That is a watching-item for readers who follow pending AI bills, not a recommendation of any specific draft.
There are limits. The model assumes both players see the safety standard in advance, that revenue splits in a specific way, and that the regulator's penalties are the only enforcement channel. Laufer, as quoted in the Cornell press release, frames the result as a warning about bill design rather than a forecast of any one law. The published PNAS version may also differ in framing or quantitative content from the arXiv preprint, and any specific numbers from the model should be treated as preprint-level until the journal version is checked.
The next move is in the bills. Watch which tier the drafters set low, and what that does to the other tier's incentives. A bill that puts a high bar on the chatbot vendor while leaving the base-model lab with a low or zero bar is, by the model's logic, the worst case. A bill that picks one tier to be clearly liable, or sets comparable bars on both, is the case the model says actually protects the public.