Policymakers are increasingly concerned about “frontier models,” a term meant to identify a small group of highly-capable general-purpose models that could pose unusually serious risks, such as mass casualties or billion-dollar damage. Lawmakers are seeking to define the category so they can impose additional safety requirements on these models. But drawing that line is not straightforward: Set it too broadly, and relatively benign models face unnecessary regulation; set it too narrowly, and genuinely risky models may escape it.
Three U.S. states and the European Union now have statutes governing the most capable general-purpose artificial intelligence (AI) models, and a bipartisan House draft would add a federal definition. California’s SB 53 defines a frontier model as a foundation model trained using more than 10^26 integer or floating-point operations (FLOP) but applies its strictest rules only to “large” frontier model developers with $500 million or more in revenue. New York’s RAISE Act used a similar training threshold and initially added a $100 million compute-cost floor and covered distilled models, but the March 2026 amendments removed both in favor of California’s language. Illinois SB 315 matches it, as does the Great American AI Act draft. The EU takes a different approach: Article 51 of the EU AI Act sets a lower threshold of 10^25 FLOP for presuming that a general-purpose AI model poses systemic risk, but allows developers to rebut that presumption by demonstrating that their model’s capabilities are inferior to those of the most advanced models.
There is no strong technical basis for these definitions. The Biden Administration first used the 10^26 FLOP threshold in its 2023 executive order, at a time when no AI model on the market had yet crossed it. Grok 3, released in February 2025, was the first to pass this threshold, and Epoch AI projects roughly 30 notable models above 10^26 FLOP by 2027 and over 200 by 2030. A line drawn to capture a handful of systems is on course to capture a couple hundred which was not the original intention.
Another problem is that as training efficiency improves, developers will be able to achieve a similar level of model capability with significantly less training compute. Developers can also reach high capability without a large training run by deriving a model from one that already has it: distillation, quantization, pruning, and sparse upcycling all yield smaller models that inherit much of a larger model’s behavior at a fraction of its cost. Training compute is also a poor measure for newer reasoning models since much of their cost (and compute) comes from answering each query. As a result, models with identical capabilities may be subject to different rules.
The consequences run in two directions. Over-inclusion is a compliance tax, imposing risk frameworks, audits, and reporting duties on developers whose models pose nothing resembling catastrophic risk. It also dilutes the distinction between frontier models and the much larger universe of capable AI systems, making an exceptional regulatory regime less targeted. Conversely, under-inclusion leaves potentially risky systems outside a regime built to regulate exactly that.
The definition of frontier models should evolve in response to advances in capabilities and efficiencies, keeping the category focused on the relatively small number of models whose capabilities warrant exceptional regulation. The EU already carries this idea in statute, defining high-impact capabilities in Article 3(64) of the AI Act by reference to the most advanced models rather than a fixed figure. Article 51(3) of the AI Act lets the Commission amend thresholds by delegated act, citing algorithmic and hardware efficiency explicitly, while SB 53 requires only annual reassessment.
Over time, various legal definitions of frontier models may begin to diverge even more, creating additional complexity in AI governance. Congress should address this risk by directing the Center for AI Standards and Innovation to develop—in consultation with the private sector—a dynamic definition based on capabilities relevant to the risks these laws are intended to address, along with a process for regularly updating it as technology advances. Otherwise, regulators will continue to use definition unmoored from technical realities and fail to achieve their purpose.
Frontier AI laws are not specifically designed to protect privacy or civil rights. Catastrophic risk, as these statutes define it, means mass casualties or billion-dollar damage, and the EU’s code of practice narrows systemic risk to chemical, biological, radiological, and nuclear (CBRN) incidents, loss of control, cyber attacks, and harmful manipulation. Discrimination, surveillance, and data misuse sit in a separate, but important, layer—Colorado’s ADMT statute, Illinois’s Human Rights Act amendments, California’s ADMT rules—and those rules attach to how people use a system, rather than how they build it. Frontier AI laws should therefore remain focused on models whose capabilities create the exceptional risks those laws are designed to address, while other laws address harms arising from how AI systems are deployed. While separate in focus, these laws can and should cross reference each other where appropriate. Keeping these regulatory layers distinct would allow policymakers to address both frontier risks and more widespread AI harms without turning the frontier category into a catch-all.
The goal should not be to regulate everything that crosses yesterday’s compute threshold, but to identify the models that actually represent today’s technological frontier and tomorrow’s greatest risks. A dynamic, capability-based definition would give policymakers a better chance of doing both.


