Showing general-purpose AI the same photographs the specialist sees — three models judged against anesthesiologists on predicting a difficult airway — a clinical trial (ClinicalTrials.gov)
A prospective observational study in adults scheduled for elective surgery requiring intubation. Standardized eight-view preoperative airway photographs are assessed by ChatGPT, Gemini and Grok under one structured prompt, and their predictions are compared with expert anesthesiologist assessments, conventional airway evaluation and what actually happened during surgery. Target enrollment is 319.
Trial overview (primary data)
- StatusActive, not recruiting
- ConditionsDifficult Intubation, Difficult Laryngoscopy
- SponsorMemorial Atasehir Hospital
- Target enrollment319 participants
- Period2026-06-25 〜 2026-09-01
Key points
- A prospective observational study of 319 participants testing whether general-purpose multimodal AI (ChatGPT, Gemini, Grok) can predict difficult intubation and difficult laryngoscopy from eight-view preoperative airway photographs.
- All three models receive the same structured prompt, so differences between them are measurable on the same footing before any comparison with people.
- Three comparators are set: expert anesthesiologist image assessment, conventional airway evaluation findings, and prospectively recorded intraoperative outcomes.
- Only 12 of the 803 records this site holds as of 2026-09-04 (1.5%) mention comparison against specialists.
- The record positions the work as a clinician-supervised screening tool, not as replacement.
1What anesthesia wants settled beforehand
General anesthesia secures the airway by passing a tube into the trachea, and depending on the shape of the throat and the movement of the neck that can be difficult. Knowing in advance allows staff and equipment to be arranged; being surprised by it carries real danger.
Preoperative assessment therefore looks at how far the mouth opens and how much of the throat is visible, yet conventional evaluation is known to be an unreliable predictor. This study takes eight-view preoperative airway photographs as its material and measures how far commonly available general-purpose AI can predict from them.
2Not a purpose-built model, but general-purpose ones
The material is the same set of eight photographs, and ChatGPT, Gemini and Grok are each given the identical structured prompt. That makes the differences between models measurable on the same footing before any comparison with people. Because these models were never built for medical images, a result showing that they do not work carries just as much information.
3Where the truth is placed
What matters in this design is that three comparators are set, not one: an anesthesiologist's assessment of the same photographs, the findings of conventional airway evaluation, and the intraoperative airway outcomes recorded prospectively.
Agreeing with a specialist and correctly predicting what turned out to be difficult are different achievements, and on the rung where AI is compared with people, the choice of what counts as truth drives the conclusion itself.
4Positioned as supervised screening
The record states the intent as exploring whether image-based AI assessment may support preoperative airway risk stratification as a clinician-supervised screening tool. Not replacement, but help with catching cases. In any attempt to bring general-purpose models into clinical judgment, where in the process they sit and with what authority has to be settled before the performance figures arrive. This registration settles that first.
Why it matters
Bringing general-purpose models into clinical judgment requires deciding on the comparator and the placement before the accuracy figures arrive. This registration sets three reference points, expert assessment, conventional evaluation and the actual outcome, and states the placement as supervised screening. Running several general-purpose models over the same material under one prompt is a pattern that carries into evaluations well outside medicine.
FAQ
Why use general-purpose AI?
What counts as the truth when performance is measured?
Will this AI be used in actual anesthesia decisions?
Sources (primary)
Source: ClinicalTrials.gov (U.S. NIH/NLM, public domain). This site does not provide medical advice. Verify the latest and exact details with the official source. This site is not endorsed or certified by the NIH/NLM.
- ClinicalTrials.gov (study record, original)
- NCT ID: NCT07700485