Medical-AI trial: can ChatGPT classify acetabular (pelvic) fractures? compared with orthopedic residents (multimodal LLM, 184 cases) (NCT07673991)
A retrospective observational study evaluating the diagnostic reliability of the multimodal AI model ChatGPT-4o in Letournel-Judet classification of acetabular fractures from pelvic radiographs (Judet views), compared with two fourth-year orthopedic residents and a reference standard from CT and intraoperative findings. 184 cases. Completed.
Trial overview (primary data)
- StatusCompleted
- ConditionsAcetabular Fractures, Pelvic Injury
- InterventionsOTHER: Diagnostic Assessment by ChatGPT-4o
- SponsorAnkara City Hospital Bilkent
- Target enrollment184 participants
- Period2026-02-01 〜 2026-05-12
Key points
- A retrospective observational study of ChatGPT-4o reliability in Letournel-Judet classification of acetabular fractures from pelvic (Judet) radiographs.
- Compared against two fourth-year orthopedic residents and a reference standard from CT and intraoperative findings.
- A systematic radiographic checklist assesses whether the AI identifies anatomical landmarks and integrates them into a fracture pattern.
- Primary endpoint is the diagnostic accuracy of the ChatGPT-4o classification; 184 cases.
- Aim is feasibility data for LLMs as decision support in complex trauma; not an establishment of clinical adoption or specialist replacement.
- ChatGPT-4o classification is measured against two standards: two residents, and confirmation from CT and intraoperative findings.
1Acetabular fractures are a complex call
The acetabulum is the socket part of the pelvis at the hip joint, and fractures there have complex shapes for which accurate classification is essential to plan surgery. The widely used Letournel-Judet system requires reading how fracture lines run across radiographs from several angles and demands expertise.
This study examines whether a multimodal large language model (LLM) that can also handle images, ChatGPT-4o, can be used for that task.
2Tested on 184 pelvic trauma cases
Per the registry summary, the study retrospectively analyzes standard radiographs (anteroposterior, iliac oblique, and obturator oblique, the Judet views) from 184 patients with pelvic injuries, and compares the ChatGPT-4o classification against the independent assessments of two fourth-year orthopedic residents and a reference standard set by an experienced trauma surgeon using multiplanar CT and intraoperative findings.
A systematic radiographic checklist is used to assess whether the AI can identify key anatomical landmarks (fracture lines, columns, walls) and integrate them into a final fracture pattern. The primary endpoint is the diagnostic accuracy of the ChatGPT-4o classification.
3Measured against two standards
The study compares the classification by ChatGPT-4o against two different references. Each measures something different.
| The comparator | How the standard is formed |
|---|---|
| Two orthopaedic residents | Fourth-year residents assessing independently — a comparison against human judgement |
| The confirmed determination | An experienced trauma surgeon confirming from multiplanar CT and intraoperative findings — a comparison against truth |
Acetabular fractures are complex in form, and the widely used Letournel-Judet classification requires reading the course of fracture lines across radiographs in several projections, which takes experience. A systematic radiographic checklist is used to assess whether the AI can identify the key anatomical landmarks and integrate them into a final fracture pattern, across 184 pelvic trauma cases analysed retrospectively.
Why it matters
Whether a general-purpose large language model can support specialist image-classification decisions is a high-interest topic. Measuring it against two yardsticks (residents and a confirmed reference) is a useful reference for evaluating LLM feasibility in care (this trial evaluates feasibility and reliability, not clinical adoption).
FAQ
What is a multimodal AI?
Does ChatGPT replace orthopedic surgeons?
Sources (primary)
Source: ClinicalTrials.gov (U.S. NIH/NLM, public domain). This site does not provide medical advice. Verify the latest and exact details with the official source. This site is not endorsed or certified by the NIH/NLM.
- ClinicalTrials.gov (study record, original)
- NCT ID: NCT07673991