cs.CY cs.AI

A taxonomy built to make AI audits executable — at least 74 risk taxonomies exist, and almost all stop at naming things

cs.CY Gemma Galdon Clavell, Pablo Accuosto, Usman Gohar Jul 2026

Many taxonomies organize the risks of AI, yet most go no further than listing them, leaving unsaid how a risk becomes a measured value and a defensible grade. This paper presents the layer that turns one risk into a test, a measurement and a grade, and then the taxonomy that lets the method scale.

Paper overview (our summary)

  • Field (arXiv category)cs.CY(+1)
  • AuthorsGemma Galdon Clavell, Pablo Accuosto, Usman Gohar
  • Submitted2026-07-02
  • arXiv ID2607.02201v1

Key points

  • At least 74 AI risk taxonomies exist, and almost all go no further than listing risks rather than showing how an audit is conducted.
  • The authors argue the hard part is converting a risk into an executable procedure, a figure, a severity on a calibrated scale and a defensible grade.
  • The paper demonstrates this end to end on one risk, personal information disclosure, against a public benchmark.
  • In the example shown, disclosure moves from 0 percent to 51 and then 84 percent as adversarial conditioning increases.
  • The taxonomy organizes 76 subcategories across 10 categories and 20 sub-groups, with mappings to 18 external frameworks.

1Naming a thing and measuring it

Many catalogues have been assembled setting out what can go wrong with AI. The authors count at least 74. Having a catalogue and being able to audit, however, are different. What is hard in an audit lies past the point of naming.

  1. 1NamingPlacing an item such as personal information disclosure into a catalogue
  2. 2Turning it into a testMaking it a procedure that can be run against a real system
  3. 3MeasuringObtaining a result as a number
  4. 4GradingProducing a defensible assessment against calibrated severity

Most taxonomies stop at the first step, on the authors account. A catalogue supplies shared vocabulary, but a catalogue alone cannot stop two people reaching different conclusions about the same risk.

2Change the conditions and the number changes

The risk shown in the examplePersonal information disclosureRun against a public benchmark
With weak adversarial conditioning0 percent disclosureThe same model, the same risk
As conditioning strengthens51 percent, then 84 percentMeasurement design changes the conclusion

What stands out in the example is that the same risk on the same model moves from 0 to 84 percent as the conditions change. There is no single answer to whether a model discloses; the number means nothing without how it was tested. Comparing audit results therefore requires the method of measurement to match as well.

3Open infrastructure by design

The conceptual scaffold of the taxonomy is published with stable identifiers and machine-readable formats, while the methodological part such as severity calibration sits in a practitioner layer. The scaffold is shared while the operational build-out is left to each practice.

4The link to regulation

This site covers many articles on how rules and audit procedures are made. Where regulation calls for risk assessment, what is called for is not fitting things into a catalogue but measuring and showing. The layer the authors call a bridge is exactly the gap between regulatory text and a real system. This article is our own summary of public research information and does not warrant its contents.

Why it matters

Comparing AI assessments requires matching not only numbers but methods. Unless the conditions of measurement are recorded, results can be neither compared nor reproduced by either the party requiring an audit or the party undergoing one.

FAQ

Is a catalogue of risks enough to audit?
The paper says no. A catalogue supplies vocabulary but does not show how to measure and grade, so conclusions about the same risk can diverge.
Why does disclosure vary on one model?
Because results move with how the test is conditioned. In the example it ranges from 0 to 84 percent, so a number means little without the method.

Sources (primary)

Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.

#AI#arXiv#Research papers#AI auditing#Risk assessment
Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.