Completed OBSERVATIONAL NCT07673991

Medical-AI trial: can ChatGPT classify acetabular (pelvic) fractures? compared with orthopedic residents (multimodal LLM, 184 cases) (NCT07673991)

Ankara City Hospital Bilkent Updated 2026-06-29

A retrospective observational study evaluating the diagnostic reliability of the multimodal AI model ChatGPT-4o in Letournel-Judet classification of acetabular fractures from pelvic radiographs (Judet views), compared with two fourth-year orthopedic residents and a reference standard from CT and intraoperative findings. 184 cases. Completed.

Trial overview (primary data)

  • StatusCompleted
  • ConditionsAcetabular Fractures, Pelvic Injury
  • InterventionsOTHER: Diagnostic Assessment by ChatGPT-4o
  • SponsorAnkara City Hospital Bilkent
  • Target enrollment184 participants
  • Period2026-02-01 〜 2026-05-12

Key points

  • A retrospective observational study of ChatGPT-4o reliability in Letournel-Judet classification of acetabular fractures from pelvic (Judet) radiographs.
  • Compared against two fourth-year orthopedic residents and a reference standard from CT and intraoperative findings.
  • A systematic radiographic checklist assesses whether the AI identifies anatomical landmarks and integrates them into a fracture pattern.
  • Primary endpoint is the diagnostic accuracy of the ChatGPT-4o classification; 184 cases.
  • Aim is feasibility data for LLMs as decision support in complex trauma; not an establishment of clinical adoption or specialist replacement.

The acetabulum is the socket part of the pelvis at the hip joint, and fractures there have complex shapes for which accurate classification is essential to plan surgery. The widely used Letournel-Judet system requires reading how fracture lines run across radiographs from several angles and demands expertise.

This study examines whether a multimodal large language model (LLM) that can also handle images, ChatGPT-4o, can be used for that task.

Per the registry summary, the study retrospectively analyzes standard radiographs (anteroposterior, iliac oblique, and obturator oblique, the Judet views) from 184 patients with pelvic injuries, and compares the ChatGPT-4o classification against the independent assessments of two fourth-year orthopedic residents and a reference standard set by an experienced trauma surgeon using multiplanar CT and intraoperative findings.

A systematic radiographic checklist is used to assess whether the AI can identify key anatomical landmarks (fracture lines, columns, walls) and integrate them into a final fracture pattern. The primary endpoint is the diagnostic accuracy of the ChatGPT-4o classification.

The positioning of this study is to measure, against two yardsticks (residents and a confirmed CT/intraoperative reference), how far an LLM can currently serve as a decision-support tool in complex orthopedic trauma. How well a general-purpose LLM holds up on specialist image classification needs evidence, and providing a checklist to structure the steps is part of that.

It is a feasibility and reliability evaluation; it does not establish clinical adoption of ChatGPT-4o or replacement of specialist judgment.

Why it matters

Whether a general-purpose large language model can support specialist image-classification decisions is a high-interest topic. Measuring it against two yardsticks (residents and a confirmed reference) is a useful reference for evaluating LLM feasibility in care (this trial evaluates feasibility and reliability, not clinical adoption).

FAQ

What is a multimodal AI?
An AI that can take inputs beyond text, such as images. Here it reads radiographs and uses them to judge fracture classification.
Does ChatGPT replace orthopedic surgeons?
No. This study compares LLM classification reliability against people and a confirmed reference; it does not establish clinical adoption or replacement of specialist judgment.

Sources (primary)

Source: ClinicalTrials.gov (U.S. NIH/NLM, public domain). This site does not provide medical advice. Verify the latest and exact details with the official source. This site is not endorsed or certified by the NIH/NLM.

#Medical AI#Clinical trial#Large language model#Orthopedics#Imaging#ChatGPT
Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.