California is building the regulatory plumbing for a profession that barely exists: the AI evaluator, a specialist who audits AI systems on behalf of the companies that build them. As Khari Johnson reported for CalMatters, Gov. Gavin Newsom last month signed one law setting standards for the “independent verification organizations” that would employ evaluators and a second creating a registry of evaluators. He also assembled an expert group, due to report in November, on two harder questions: whether evaluators should be embedded inside companies developing the most powerful AI systems, and what counts as an adequate audit at all.
The ecosystem is moving in the same direction. More than 200 AI researchers and evaluators — including a former OpenAI whistleblower and the head of the United Nations’ Independent International Scientific Panel — signed a letter last month backing standards for independent evaluators. A joint international statement calling for independent AI evaluations has drawn support from nearly 30 countries, according to the Finnish Ministry of Foreign Affairs. When a new AI evaluation group launched at the UN in New York, its closing speaker was a California state senator, Jerry McNerney, the Democrat who authored the verification-standards law. “We can put these companies on notice,” McNerney said, “that they’re going to be evaluated by folks that understand the process, that understand the technology, and will be transparent and hold them accountable so that they create safe products, and they don’t send out these things into the wild that can cause havoc with banking, with our infrastructure and so on.”
The push follows measurable incidents. This year, AI agents being tested for their ability to hack computer systems carried out a series of high-profile attacks on business and government websites, and the CEOs of leading AI companies subsequently committed to independent third-party evaluations of their models. A former Anthropic employee added to the alarm by posting publicly about the possibility that AI models will kill all humans. Public appetite for regulation already existed: roughly three in four Californians said this spring that the government should require testing of advanced AI models, per a Carnegie Endowment California survey, and testing requirements appeared in last year’s bills on discriminatory AI decisions and catastrophic-event prevention.
Who pays the auditor
The structural problem is the same one bond rating had before 2008. Independent evaluators sign contracts with the companies whose products they evaluate. Assemblymember Rebecca Bauer-Kahan, the San Ramon Democrat behind both auditor bills, put it plainly: “I’m hearing from evaluators, ‘We have to be careful, because we want them to let us back in.’ So what are you avoiding saying in order to continue to gain access and contracts?”
Caroline Siegel-Singh of the Federation of American Scientists makes the comparison explicit: before the 2008 financial crisis, bond sellers could effectively shop for the rating they wanted, and AI audit regimes risk repeating that structure. Her proposed fix is to have companies and governments pool money to pay evaluators rather than letting labs pay them directly — a mechanism Anthropic has also endorsed. The 200-plus researchers’ letter adds three requirements: evaluators with full editorial control over their findings, meaningful disclosure and mitigation of conflicts, and access for multiple evaluation groups at the same level a company’s own safety employees get.
Bauer-Kahan names a second failure mode: AI companies testing models as a performative exercise, to reassure the public without sincerely trying to find problems. Strong legal requirements, she argues, can prevent that.
The compliance floor
Legal requirements create their own distortion, says Matt O’Shaughnessy, a former State Department and congressional staffer now at the Center for Democracy and Technology. In guidelines he and colleagues recently published for policymakers, they concluded that when an assessment exists to satisfy a legal mandate, organizations tend to do the minimum the mandate requires. “It can make it harder for assessors to get the buy-in they need to make deeper changes in companies,” O’Shaughnessy said. His guidelines also note what a real audit demands: experts spanning privacy, mental health, cybersecurity, bias in outputs and physical safety — plus scrutiny of incentive structures inside the AI company, not just the technology. Siegel-Singh adds that evaluators should publish results so other evaluators can judge the quality of the work.
Sacramento as a node in a larger network
California is already synchronizing with foreign regulators. In 2024, state lawmakers sought to align AI regulation with the European Union, and in August they synced implementation of AI-content watermarking laws with the EU. Foreign consulates sometimes coordinate tech policy in Sacramento directly. David Lametti, Canada’s ambassador to the UN, argues that a coalition of mid-sized countries working with California could pool economic influence to push the largest AI companies toward the safety guidelines in the international statement. Pasi Rajala, Finland’s deputy foreign minister, said California working with small nations “can play a key role” against fast-developing AI risks.
O’Shaughnessy inverts the dependency: an international framework for commercial AI evaluation will be very hard to build without a domestic US framework first. That framework does not exist. The closest thing is the Frontier Act, a bill that would write testing requirements into federal law; it has the support of California Republican Rep. Jay Olbernolte, and Rep. Ted Lieu, the Democrat co-chairing the House AI caucus, called for a committee vote on it last month. “I can imagine some of the things that California is doing influencing how Congress thinks about these issues and that becoming something bigger,” Lieu said.
The November report from Newsom’s expert group is the next concrete marker: it will say whether California mandates evaluators inside the labs building its most powerful models — and whether “adequate audit” gets a definition the labs’ own contractors have to meet.

