Mapping Evaluation Cards to Emerging AI Governance Requirements in California, the EU, and the UK

Mapping Evaluation Cards to Emerging AI Governance Requirements in California, the EU, and the UK

A push for independent assessment

On September 18, California Governor Gavin Newsom issued an executive order accelerating the state’s work on independent AI oversight. Among other things, the order asks state agencies to develop recommendations for embedding independent verification organizations within frontier AI companies to conduct regular audits and evaluations, verify safety frameworks and risk assessments, and maintain independent verification of emergency shutdown mechanisms. California has also recently created frameworks for independent verification organizations (SB 813) and a registry for AI auditors (AB 1405).

This builds on a broader shift toward technical evaluation and assurance around the world. The UK’s AI Security Institute (AISI) has been conducting independent evaluations of advanced AI systems since 2023, including pre- and post-deployment assessments of potentially harmful capabilities, and Parliament’s Joint Committee on Human Rights (JCHR) recently called for new statutory safeguards against the human rights risks posed by AI in its report on Human Rights and the Regulation of AI. In the EU, the AI Act requires providers of general-purpose AI models that pose systemic risk to conduct and document state-of-the-art model evaluations, including adversarial testing, while the GPAI Code of Practice further develops expectations for pre-deployment evaluation, post-market monitoring, and external evaluations. The European AI Office is also developing independence and qualification requirements for external evaluators of systemic-risk GPAI models. Similar efforts elsewhere, such as Singapore’s AI Verify, are pushing toward more structured approaches to testing, documentation, and independent scrutiny.

The evidence question

Across these approaches, a practical implementation question keeps coming up: once an evaluation is conducted, how should its evidence be recorded and communicated?

Policymakers need to know what was tested and how. Evaluators and auditors need enough methodological detail, including completed metadata, to interpret or reproduce results. Developers increasingly face requests for similar evidence in different formats. Shared reporting conventions can streamline that evidence across research, assurance, and governance processes.

Shared reporting can also reduce the burden on oversight bodies. When providers report evaluation evidence in different formats and at varying levels of detail, regulators have to spend additional time locating, interpreting, and comparing the information they need. Standardized, structured reporting can make that process easier and also enable more reliable automated parsing and analysis of evaluation evidence.

Such a common baseline can also support other governance processes. In California, for example, the state will need to evaluate and compare prospective Independent Verification Organizations as it develops its designation regime. SB 813 requires applicants to provide information about the benchmarks, technologies, metrics, and methodologies they propose using, and directs the state to develop the regime with attention to consistency, comparability, and avoiding duplicative requirements where practicable. A common format for describing evaluation evidence could make that information easier to assess across applicants and over time.

Mapping requirements to Evaluation Cards

This is where our work at EvalEval fits into a broader open evaluation ecosystem. Through Every Eval Ever and the wider Evaluation Cards effort, we are working on shared, open infrastructure for documenting evaluation results in ways that make them easier to find, compare, analyze, reproduce, and reuse.

We are currently working with public-sector evaluators to put this infrastructure into practice. In our collaboration with the UK AI Security Institute, AISI is making publicly reported evaluation methods and findings available through Evaluation Cards where appropriate, including verified results, context, and configuration information. Feedback from AISI has also helped shape the Every Eval Ever (EEE) schema itself. This gives us a concrete example of how open reporting infrastructure can support independent evaluation: results produced by an evaluator can be published in a shared structure and compared with evidence from the wider ecosystem.

To understand how far this infrastructure could support emerging governance requirements, we reviewed 40 requirements across the California (AB 1405, SB 813), EU (AI Act, GPAI Code of Practice), and UK (JCHR report) instruments in our crosswalk and broke each requirement into the discrete pieces of evidence needed to satisfy it. This produced 281 evidence elements in total.

We then classified each element according to its unit of record. We treated an element as evaluation-reporting evidence when it primarily describes an evaluation run, its methodology, configuration, benchmark, results, or the conditions under which it was conducted. We classify this as information that can reasonably be attached to an evaluation record. Evidence about an evaluator’s institutional independence or qualifications, a model as a whole, an incident, organizational governance, or a regulatory process was instead assigned to the corresponding evaluator-, model-, incident-, organization-, or regulator-level record. On this basis, 60 of the 281 elements fall within the scope of evaluation reporting, while the remaining 221 belong in other kinds of records.

For each of those 60 elements, we then asked whether EEE or AutoBenchmarkCards contains a field whose unit and meaning match the required evidence, and whether that information can be represented in a structured field, in free text only, or not at all. Evaluation Cards unifies evaluator-reported information from EEE with benchmark metadata automatically supplied by AutoBenchmarkCards, as shown in the visualization.1

What the mapping found

Of the 60 evidence elements in scope for evaluation reporting, EEE can already represent 56. Fifty-one are captured through dedicated, typed fields; another five can currently be recorded through free-text fields supplied by the evaluator or other reporter.

For example, requirements in the EU’s General-Purpose AI (GPAI) Code of Practice to document an evaluation’s tooling and resource constraints map onto existing EEE fields for tools available during an evaluation and token limits, while requirements to record whether an evaluation was conducted internally or by an external evaluator map to source_metadata.evaluator_relationship.

Evaluation Cards can combine those submissions with benchmark-level information automatically pulled from AutoBenchmarkCards. In our crosswalk, AutoBenchmarkCards turns two of the five elements that EEE captures only in free text (benchmark data type and benchmark languages) into structured metadata. This increases the number represented through structured fields from 51 to 53.

Where other standards fit

Our research on evaluation reporting standards is one part of the broader push towards independent assessment. Questions about evaluator independence, conflicts of interest, access arrangements, credentials, and regulator-facing processes belong in standards focused on evaluators themselves. AEF-1, for example, addresses operating conditions for independent third-party evaluators. Model-level information has a natural home in model cards, while incident and flaw reporting is being developed through efforts such as FLARE-AI.

These efforts can work as complementary layers: evaluator standards establish expectations for who performs an assessment and under what conditions; model cards document the system being assessed; evaluation reporting standards record what was tested, how, and with what results; and incident-reporting standards capture failures that emerge after deployment. Making those layers interoperable can allow the same underlying evidence to travel more easily between evaluators, developers, researchers, and oversight bodies without asking each actor to reinvent its documentation from scratch.

Explore the map

The interactive visualization below presents our current mapping. It shows which pieces of governance evidence an Evaluation Card can address, with each field color-coded by where it comes from: reported fields that evaluators fill in and submit through Every Eval Ever, and auto-pulled benchmark metadata that Evaluation Cards brings in from AutoBenchmarkCards. What a reporter controls is what they submit through EEE. Open it in a new tab for the most room, and hover over any band for the source text and field definitions.

Governance evidence crosswalk
Each band links one piece of required evidence to an Evaluation Card field that could record it: blue for fields the evaluator reports through Every Eval Ever, violet for benchmark metadata pulled in from AutoBenchmarkCards. Hover a band or node for the source quote, field definition, and mapping rationale; click to pin.

Evaluation Cards are designed to be living documents that grow with multi-stakeholder input over time. We therefore believe that the relationship between governance requirements and evaluation infrastructure should be bidirectional: new policy requirements can help identify where Evaluation Cards need to become more expressive, while existing schemas can give policymakers a concrete, machine-readable way to specify and compare the evaluation evidence they are asking for. Our mapping shows that the same reporting infrastructure can already support a substantial share of the evaluation-specific evidence appearing across California, the EU, and the UK, while making clear where complementary standards or new fields are still needed. We hope regulators, evaluators, developers, and researchers will use the crosswalk as a practical reference and share feedback from real-world use so that Evaluation Cards can continue to evolve alongside emerging governance requirements.

References
  1. Mitchell et al. (2019). Model Cards for Model Reporting. arXiv:1810.03993.
  2. UK AI Safety Institute (2024). AI Safety Institute approach to evaluations. gov.uk.
  3. European Union (2024). Regulation (EU) 2024/1689 (Artificial Intelligence Act). eur-lex.europa.eu; Article 55.
  4. European Commission (2025). The General-Purpose AI Code of Practice. digital-strategy.ec.europa.eu.
  5. Hofmann et al. (2025). Auto-BenchmarkCard: Automated Synthesis of Benchmark Documentation. arXiv:2512.09577.
  6. Batzner et al. (2026). Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results. arXiv:2606.14516; GitHub.
  7. California Legislature (2026). AB 1405: Artificial intelligence: auditors: registration. leginfo.legislature.ca.gov.
  8. California Legislature (2026). SB 813: Independent verification organizations. leginfo.legislature.ca.gov.
  9. Ghosh et al. (2026). Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting. arXiv:2606.09809.
  10. Governor of California (2026). Governor Newsom issues executive order to accelerate independent oversight and advance the creation of an AI kill switch. gov.ca.gov.
  11. Longpre et al. (2026). FLARE-AI: Flaw Reporting for AI. arXiv:2606.31567.
  12. AI Evaluator Forum (live). AEF-1: Minimum Operating Conditions for Independent Third-Party AI Evaluations. aievaluatorforum.org.
  13. AI Verify Foundation (live). What is AI Verify?. aiverifyfoundation.sg.
  14. EvalEval Coalition & UK AI Security Institute (2026). How UK AISI and EvalEval Are Making Benchmark Results Reproducible. evalevalai.com.
  1. AEF-1 appears separately as an optional overlay because it primarily describes the evaluator rather than the evaluation itself, although some of its requirements overlap with evaluation-reporting evidence. ↩

Get Involved

Join Our Community

Researchers, practitioners, and students are welcome to contribute to our mission. Send us an email to learn more about getting involved.

[email protected]

Hosted By