Comparing security scanners for AI models — whether a judgment is right and whether one was reached should be measured separately
A study comparing three static scanners that check machine learning artifacts for unsafe content. One scanner was perfect on the cases where it reached a judgment, yet reached no judgment at all on half the subjects. Conventional metrics hide that gap.
Paper overview (our summary)
- Field (arXiv category)cs.CR(+1)
- AuthorsQianlong Lan, Vinothini Pandurangan, Anuj Kaul, et al. (4)
- Submitted2026-08-27
- arXiv ID2608.27424v1
Key points
- Three static scanners for machine learning artifacts were compared on a controlled corpus across 145 specimen families.
- Conventional metrics capture only the cases where a scanner yields a usable security judgment.
- Definitive judgments were reached on 100 percent, 81.5 percent and 49.6 percent of the 135 labelled families by tool.
- Conditional on judging, the tool at 49.6 percent achieved 100 percent precision, recall and F1.
- For 48 malicious families one tool could not complete, the other two detected consistently with ground truth.
1Being right is not enough
A scanner tends to be judged on whether it gets things right. Yet rightness can only be measured where a judgment was reached at all. Subjects on which none was reached drop out of the calculation.
The numbers here make the difference vivid. One tool achieved 100 percent precision, recall and F1 on what it judged — never once wrong. Yet of the 135 labelled families, it reached a definitive judgment on 67, or 49.6 percent. On half the subjects, no answer was obtained at all.
2The three tools
The share on which a judgment was reached ranges from 100 percent down to 49.6. Moreover, for the 48 malicious families where one tool could not complete its analysis, the other two produced detections consistent with ground truth. One tool gap was filled by another.
3Increment against redundancy
The authors raise a second point: separating incremental detection coverage from redundancy between tools. In this study one tool identified no malicious family beyond those found by the other two combined. Adding tools does not necessarily widen coverage.
4The link to other records on this site
This site covers many records of vulnerabilities actually exploited, where whether something can be detected and whether the detection is correct likewise appear as separate questions. The point that evaluation design drives the conclusion bears directly on choosing tools. This article is our own summary of public research information and does not warrant its contents.
Why it matters
Comparing scanners requires separating accuracy from whether an answer arrives at all. Where several are used together, whether an added tool actually widens coverage has to be checked separately.
FAQ
Does an F1 of 100 percent mean the tool is best?
Does adding tools widen coverage?
Sources (primary)
Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.
- arXiv abstract page (original, official)
- PDF (arXiv)
- arXiv ID: 2608.27424