The largest box is other — residents testing four language models against standard references
A study in which family medicine, internal medicine and psychiatry residents use OpenEvidence during actual practice, with their conclusions assessed for clinical appropriateness and compared against three other large language models. The registered intervention is other.
Trial overview (primary data)
- StatusActive, not recruiting
- ConditionsAI (Artificial Intelligence), Large Language Model, Generative Artificial Intelligence
- InterventionsOTHER: AI clinical reference tool
- SponsorCambridge Health Alliance
- Target enrollment20 participants
- Period2025-10-01 〜 2026-09-30
Key points
- OpenEvidence aggregates and synthesizes peer-reviewed medical studies and answers user questions using generative artificial intelligence.
- Though clinicians use it, there is little published data on whether its outputs are accurate or appropriately inform clinical decisions.
- The study asks whether its use leads to clinically appropriate decisions among family medicine, internal medicine and psychiatry residents.
- It also compares the output against ChatGPT from OpenAI, Claude from Anthropic and Gemini from Google.
- Residents also consult a standard reference tool to mitigate safety risks, and target enrolment is 20.
1The largest box says the least
Each of the four boxes so far carried a meaning: observation, diagnostic test, device, behavioural. This article takes up the box chosen more often than any other in these registrations, and at the same time the one carrying almost no information as a classification. Other.
One registration in four says other. What the AI actually does cannot be read from that field.
2What is being tested
Using questions that arose in real practice rather than exam items is the point of the design. The registration notes that few studies have examined performance in a real-world clinical setting and fewer still have compared models.
3An observational study that lists an intervention
This registration is typed observational, yet the intervention field carries an AI clinical reference tool. A clinician using a tool is an action, but not an intervention on a patient. That tension reads as one reason the registration settles into the box marked other.
Target enrolment is 20, and the subjects here are residents rather than patients. The next article takes up an AI registered as a procedure.
Why it matters
The intervention box chosen most often in these registrations is other, covering one in four, and it is also the box that says least about what the AI does. Studies where a clinician uses a tool without any intervention reaching a patient tend to collect there, and testing language models in real practice is one such case.
FAQ
Why use real clinical questions rather than exam items?
Are the subjects patients?
Why is the intervention listed as other?
Sources (primary)
Source: ClinicalTrials.gov (U.S. NIH/NLM, public domain). This site does not provide medical advice. Verify the latest and exact details with the official source. This site is not endorsed or certified by the NIH/NLM.
- ClinicalTrials.gov (study record, original)
- NCT ID: NCT07199231