Two large language models measured against a blinded specialist consensus on Turkish emergency department notes — 600 cases (NCT07632859)
A retrospective observational study measuring how far the diagnoses assigned by two large language models (GPT-4.1 and Claude Sonnet 4.6) agree, at ICD-10 chapter level, with the majority consensus of three blinded board-certified emergency medicine specialists, across 600 anonymized Turkish-language emergency department notes. Recorded as completed in August 2026.
Trial overview (primary data)
- StatusCompleted
- ConditionsEmergency Medicine, Diagnostic Errors, Artificial Intelligence (AI) in Diagnosis
- SponsorMarmara University Pendik Training and Research Hospital
- Target enrollment600 participants
- Period2026-05-01 〜 2026-08-07
Key points
- The reference standard is the majority consensus of three board-certified emergency medicine specialists coding independently and blinded; cases without chapter-level majority agreement are excluded without replacement.
- Both models are queried once per note with a single locked prompt at temperature 0 in stateless API calls, with no retrieval augmentation, external tools, or extended-reasoning mode.
- The primary outcome is the proportion of cases where the rank-1 diagnosis matches the reference at ICD-10 chapter level, with a Wilson 95 percent confidence interval.
- The ICD-10 code entered at case closure is not a comparator; it describes current documentation practice, and no test of superiority or inferiority is performed.
- The analysis plan was frozen before any accuracy computation and reporting follows STARD-AI 2025. The material is 600 anonymized Turkish-language notes.
- Truth is a blinded majority of three specialists, with a single prompt, temperature 0 and a frozen analysis plan.
1Deciding what counts as correct, before measuring
The hardest part of measuring how well a large language model diagnoses is deciding what counts as the right answer. The reference here is the majority of ICD-10 codes assigned independently by three board-certified emergency medicine specialists, each blinded to the others. The code the treating physician entered at case closure and the subsequent clinical course are withheld from them as well.
Further, cases where the specialists do not reach chapter-level majority agreement are excluded without replacement — a decision not to force cases with an unsteady answer into the scoring.
2Fixed conditions so the measurement can be repeated
The conditions of measurement are pinned down in detail: one query per note, a single locked prompt, temperature 0, stateless application programming interface calls, with no retrieval augmentation, no external tools and no extended-reasoning mode. And the analysis plan was finalised and frozen before any accuracy computation began.
Generative models vary in output even on identical input, and the numbers move when prompts or settings change. Writing the conditions down and freezing them is the brake that stops that variability from being used selectively after the fact. Reporting is stated to follow STARD-AI 2025, a reporting guideline for diagnostic accuracy studies.
3Observational work makes up 44 percent of the field
The careful part is how the ICD-10 code entered at case closure is handled. The record states explicitly that it is not a comparator; it is characterised against the same reference standard as a description of current documentation practice, and no test of superiority or inferiority is performed. Of the 750 medical-AI clinical trials this site holds as of 2026-08-28, 333 are observational, or 44 percent of the total.
With 417 interventional trials making up the other 56 percent, trials that use AI to change something and trials that measure AI itself are close to evenly matched. The result figures are not part of this registry record.
4Fixing conditions so nothing drifts
Generative models vary in output on the same input, and the numbers move with the prompt and the settings. This study fixes conditions in advance so that drift cannot be used conveniently after the fact.
- 1Settle the truth firstMajority of ICD-10 codes assigned independently and blind by three board-certified emergency physicians
- 2Exclude the unstable casesCases without chapter-level majority agreement are excluded without replacement
- 3Fix how it is measuredOne call per record, a single fixed prompt, temperature 0, no external tools
- 4Freeze the analysis planFinalised and frozen before accuracy is computed
Reporting follows STARD-AI 2025, the guideline for diagnostic accuracy studies. The ICD-10 code entered at case closure is stated explicitly not to be a comparator: it is described as depicting current documentation practice and held to the same standard, with no test of superiority or non-inferiority.
Why it matters
In clinical evaluation of large language models the numbers move with prompt, temperature and reasoning mode. Fixing the conditions and writing them down, freezing the analysis plan before computation, and following a reporting guideline is a concrete answer to how a performance claim about AI can be made checkable. Measuring in a language other than English also matters for thinking about where evaluation is skewed.
FAQ
Why use the majority of three specialists as the reference?
What does freezing the analysis plan mean?
Which model was shown to be better?
Sources (primary)
Source: ClinicalTrials.gov (U.S. NIH/NLM, public domain). This site does not provide medical advice. Verify the latest and exact details with the official source. This site is not endorsed or certified by the NIH/NLM.
- ClinicalTrials.gov (study record, original)
- NCT ID: NCT07632859