cs.CR cs.AI

Comparing security scanners for AI models — whether a judgment is right and whether one was reached should be measured separately

cs.CR Qianlong Lan, Vinothini Pandurangan, Anuj Kaul, et al. (4) Aug 2026

A study comparing three static scanners that check machine learning artifacts for unsafe content. One scanner was perfect on the cases where it reached a judgment, yet reached no judgment at all on half the subjects. Conventional metrics hide that gap.

Paper overview (our summary)

  • Field (arXiv category)cs.CR(+1)
  • AuthorsQianlong Lan, Vinothini Pandurangan, Anuj Kaul, et al. (4)
  • Submitted2026-08-27
  • arXiv ID2608.27424v1

Key points

  • Three static scanners for machine learning artifacts were compared on a controlled corpus across 145 specimen families.
  • Conventional metrics capture only the cases where a scanner yields a usable security judgment.
  • Definitive judgments were reached on 100 percent, 81.5 percent and 49.6 percent of the 135 labelled families by tool.
  • Conditional on judging, the tool at 49.6 percent achieved 100 percent precision, recall and F1.
  • For 48 malicious families one tool could not complete, the other two detected consistently with ground truth.

1Being right is not enough

A scanner tends to be judged on whether it gets things right. Yet rightness can only be measured where a judgment was reached at all. Subjects on which none was reached drop out of the calculation.

Judgment accuracyJudgment availability
How often the judgment issued is correctHow often a judgment issued at all
What conventional metrics captureWhat conventional metrics omit
It means not being wrongIt means being able to answer

The numbers here make the difference vivid. One tool achieved 100 percent precision, recall and F1 on what it judged — never once wrong. Yet of the 135 labelled families, it reached a definitive judgment on 67, or 49.6 percent. On half the subjects, no answer was obtained at all.

2The three tools

Labelled specimen families135Drawn from 170 artifacts across 145 families
Share reaching a definitive judgment100 percent, 81.5 percent, 49.6 percentVarying widely by tool
The 48 malicious families one tool could not completeThe other two detected consistently with ground truthA complementary relationship

The share on which a judgment was reached ranges from 100 percent down to 49.6. Moreover, for the 48 malicious families where one tool could not complete its analysis, the other two produced detections consistent with ground truth. One tool gap was filled by another.

3Increment against redundancy

The authors raise a second point: separating incremental detection coverage from redundancy between tools. In this study one tool identified no malicious family beyond those found by the other two combined. Adding tools does not necessarily widen coverage.

4The link to other records on this site

This site covers many records of vulnerabilities actually exploited, where whether something can be detected and whether the detection is correct likewise appear as separate questions. The point that evaluation design drives the conclusion bears directly on choosing tools. This article is our own summary of public research information and does not warrant its contents.

Why it matters

Comparing scanners requires separating accuracy from whether an answer arrives at all. Where several are used together, whether an added tool actually widens coverage has to be checked separately.

FAQ

Does an F1 of 100 percent mean the tool is best?
That figure covers only what it judged. The tool reached definitive judgments on 49.6 percent of subjects, with the rest dropping out of the calculation.
Does adding tools widen coverage?
Not necessarily. In this study one tool identified no malicious family beyond those found by the other two combined.

Sources (primary)

Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.

#AI#arXiv#Research papers#Security#Evaluation metrics
Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.