Telling benign from malignant adnexal masses with ultrasound AI — a 12,000-patient international IOTA study — a clinical trial (ClinicalTrials.gov)
An international multicenter observational study testing how well machine-learning models that incorporate quantitative ultrasound features (radiomics) can distinguish benign from malignant adnexal masses before surgery. Target enrollment is 12,000.
Trial overview (primary data)
- StatusNot yet recruiting
- ConditionsAdnexal Masses
- SponsorFondazione Policlinico Universitario Agostino Gemelli IRCCS
- Target enrollment12,000 participants
- Period2026-09-01 〜 2029-09-01
Key points
- An international multicenter observational study testing ultrasound machine learning with radiomic features for preoperative classification of adnexal masses.
- Target enrollment is 12,000 — a scale that follows from validating finer categories: borderline, primary invasive and metastatic.
- AI outputs do not influence patient management or alter the diagnostic and therapeutic pathway; they are compared with real clinical decisions only post hoc.
- Evaluation extends past accuracy to decision curve analysis, resource use, cost-effectiveness and patient-reported measures.
- This describes the objectives and hypotheses of the study; effectiveness has not been established.
1An everyday problem that stays hard
Masses on the adnexa — the ovaries, fallopian tubes and surrounding tissue — turn up often in gynecology. The difficulty is how far benign can be told from malignant before surgery. Missing a malignancy delays referral to a specialized center; reading a benign mass as malignant leads to surgery that was never needed.
Ultrasound is the most accessible and widely available examination, but interpreting the images still varies with the reader's experience, and a share of cases stays genuinely hard to call. This study asks how far machine learning that incorporates quantitative features extracted from the ultrasound image (radiomic features) can support that preoperative classification.
2What a target of 12,000 signals
Across the AI-related clinical trials this site tracks, the median target enrollment is 200. A target of 12,000 is sixty times that, and large even among the minority aiming above 5,000. Malignancy in adnexal masses is not common to begin with, so gathering enough cases for each finer category — benign, borderline, primary invasive, metastatic — is far beyond the reach of any single center.
The scale is not incidental; it follows from the aim of validating a subdivided classification.
3A condition: do not change care
Equally notable is that the AI output is designed not to move actual care. The post hoc analysis lines up AI predictions against the preoperative assessments clinicians actually made, but the record states explicitly that this will not alter patient management or the diagnostic and therapeutic pathway.
The study is therefore at the stage of asking how AI judgments differ from the judgments being made today, not whether using AI improves outcomes. The proportion of potentially avoidable operations and comparisons of healthcare resource use are handled as a reconstructed scenario rather than as the result of an intervention.
4Measures placed beyond accuracy
The study also refuses to stop at accuracy. Clinical utility is assessed through decision curve analysis, resource use is compared between standard management and a reconstructed AI-supported scenario, and cost-effectiveness is brought into view. Patient-reported measures are included as well: satisfaction with how the diagnosis was communicated, the clarity of the information provided, and perceived quality of care.
It is one example of how the evaluation of imaging AI is widening from hit-rate figures toward how care actually changes and how patients experience it.
Why it matters
Adnexal masses are a frequent gynecological finding, so better preoperative classification could avoid unnecessary surgery and route suspected malignancy to specialized centers sooner. Measuring clinical utility, resource use and patient experience alongside accuracy offers a reference for how imaging AI should be evaluated.
FAQ
Would taking part change the care a patient receives?
Has ultrasound AI been shown to tell benign from malignant accurately?
Sources (primary)
Source: ClinicalTrials.gov (U.S. NIH/NLM, public domain). This site does not provide medical advice. Verify the latest and exact details with the official source. This site is not endorsed or certified by the NIH/NLM.
- ClinicalTrials.gov (study record, original)
- NCT ID: NCT07793201