Completed OBSERVATIONAL NCT07459491

Medical-AI trial: agreement between ChatGPT-5 and anesthesiologists on preoperative risk class (ASA-PS) — preoperative triage (NCT07459491)

Damla Kaytancı Özçelik Updated 2026-07-02

An observational study in adults scheduled for elective surgery comparing the agreement between the ASA-PS physical-status class assigned by anesthesiologists and the one generated by ChatGPT-5 from the same anonymized information. It also explores differences in lab recommendations and links to perioperative red-blood-cell (transfusion) use. 703 patients; completed.

Trial overview (primary data)

  • StatusCompleted
  • ConditionsArtifical Intelligence, Preoperative Evaluation
  • InterventionsOTHER: No intervention (observational study)
  • SponsorDamla Kaytancı Özçelik
  • Target enrollment703 participants
  • Period2026-01-10 〜 2026-05-01

Key points

  • An observational study measuring the agreement between anesthesiologist- and ChatGPT-5-assigned ASA-PS class in adults for elective surgery.
  • ASA-PS is the global standard for preoperative status but depends on clinician judgment, with inter-observer variability.
  • Primary endpoint is agreement of the two classifications; secondary looks at differences in lab recommendations and links to perioperative red-blood-cell use.
  • Both clinicians and the AI receive the same anonymized preoperative clinical information.
  • Target 703; completed in 2026. The focus is consistency, not AI superiority.
  • What is measured is not whether AI outperforms physicians but how far its classification agrees with theirs.

1Preoperative assessment as the entry point

Before surgery, the overall condition of a patient is assessed to plan anesthesia and the operation and to prepare transfusion and testing. The global reference for this is the ASA-PS classification, which grades patients from healthy to those with severe systemic disease. Because the class depends on clinician interpretation, judgments can differ between evaluators (inter-observer variability).

2Measuring agreement with anesthesiologists

This trial applies a large language model (LLM), ChatGPT-5, to that preoperative assessment and, as an observational study, measures how well the AI agrees with anesthesiologists. Per the registry summary, the same anonymized preoperative clinical information is given to both clinicians and the AI, and the agreement between the ASA-PS classes they assign is the primary endpoint.

Secondarily, it examines how the recommended preoperative labs differ between AI and clinicians, and how the assessments relate to perioperative red-blood-cell (transfusion) use. Participants are adults scheduled for elective surgery, with a target of 703; the study is at the completed stage.

3Not superiority, but agreement

What this trial measures is not whether AI outperforms physicians. It is how far the AI's classification agrees with theirs.

Measuring which performs betterMeasuring how far they agree (this trial)
Requires a reference truthPlaces physician judgement as the comparator
Allows discussion of replacementStops at confirming consistency with clinical judgement
The conclusion is superiority or non-inferiorityThe conclusion is agreement, and the difference in tests recommended

ASA-PS classification depends on clinician interpretation and is known to vary between assessors. The same anonymised preoperative information goes to both physician and AI, with the agreement between their classifications as the primary endpoint. Preoperative assessment is performed in volume daily, and risk classification carries through to test ordering and transfusion preparation.

Why it matters

Risk classes like ASA-PS drive resource decisions such as testing and transfusion prep. Measuring on real data how closely an LLM agrees with clinicians is a useful example of validity checking before bringing LLMs into care (this trial evaluates agreement, not efficacy).

FAQ

What is the ASA-PS classification?
A preoperative physical-status grading defined by the American Society of Anesthesiologists, from healthy patients to those with severe systemic disease. It is used worldwide to guide anesthetic and perioperative planning.
Did it show AI is more accurate than anesthesiologists?
No. This observational study evaluates how closely ChatGPT-5 agrees with clinicians; it does not establish AI superiority or safety.
Why link it to red-blood-cell (transfusion) use?
Preoperative risk class influences resource decisions such as testing volume and transfusion preparation, so the study secondarily examines how AI- and clinician assessments relate to actual perioperative red-blood-cell use.

Sources (primary)

Source: ClinicalTrials.gov (U.S. NIH/NLM, public domain). This site does not provide medical advice. Verify the latest and exact details with the official source. This site is not endorsed or certified by the NIH/NLM.

#Medical AI#Clinical trial#Large language model#Anesthesia#Preoperative assessment#ChatGPT
Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.