Completed NA INTERVENTIONAL NCT07597499

Testing whether a structural verification layer on AI-assisted care reduces safety failures — two questions in one three-arm randomized trial (Waymark, NCT07597499)

Waymark Updated 2026-07-08

A three-arm 1:1:1 patient-level randomized trial in real high-risk multidisciplinary encounters testing whether adding a clinical AI structural verification layer to an AI-assisted physician workflow reduces clinically meaningful safety failures, and whether the AI-assisted workflow itself improves the same endpoint against unassisted standard care. It enrolled 240 patients across three states over 12 weeks.

Trial overview (primary data)

  • StatusCompleted
  • ConditionsHigh-Risk Multidisciplinary Care, Clinical Decision Support, Artificial Intelligence-Assisted Care, Telemedicine
  • InterventionsBEHAVIORAL: Gemini 3.1 Pro with Safety Prompt, BEHAVIORAL: ANCHOR Clinical AI Verification Layer (with Gemini 3.1 Pro)
  • SponsorWaymark
  • Target enrollment240 participants
  • Period2026-05-15 〜 2026-07-06

Key points

  • A three-arm randomized trial testing both the effect of adding a structural verification layer to an AI-assisted workflow and the effect of AI assistance itself.
  • The arms are unassisted standard care, AI-assisted, and AI-assisted with the verification layer, randomized 1:1:1 at patient level with encounter-level mixed-effects analysis.
  • The verification layer combines a Logical Neural Network certificate, six specialist agents and concept-decomposed output with citation provenance, and is physician-facing only.
  • The primary endpoint is a per-encounter binary composite including failure to mention a do-not-miss diagnosis, under-triage and contraindication.
  • It enrolled 240 patients across three states over 12 weeks and is recorded as completed on 2026-07-06.
  • Three arms — standard care, AI support and AI plus a verification layer — test checking output rather than using a better model.

1Layering on AI rather than adding AI

How this trial frames its question is distinctive. It compares not only whether AI is used but whether a further layer of verification is placed on top of AI. The three arms are unassisted standard care, AI-assisted, and AI-assisted with the verification layer. The effect of introducing AI and the effect of a mechanism that mechanically inspects its output are measured separately within one trial.

Against the question of how to reduce AI error, it tests an answer that places a separate layer over the output rather than reaching for a better model.

2Defining what counts as a safety failure

The primary endpoint is defined as a per-encounter binary composite, including failure to mention a do-not-miss diagnosis, under-triage and contraindication. Evaluations of AI tend to reach for accuracy, whereas what is counted here is the subset of errors that carry clinical meaning. Unless what counts as a failure is settled beforehand, whether assistance helped cannot be judged. The record also notes the trial was pre-registered.

3Twelve weeks from start to completion

Enrolment began on 2026-05-15 and completion is recorded as 2026-07-06, under a design allowing a 12-week active window — short among the 803 medical-AI clinical trials this site holds as of 2026-08-31. The same sponsor has registered other trials as well; across the 803 records, 135 sponsors have more than one trial and the records belonging to them number 417, or 52 percent of the total. How far safety failures fell in either arm is not part of this registry record.

4Not adding AI, but layering on top of it

The way the question is posed marks this trial. It compares not only whether AI is used but whether a verification layer is added on top of it.

  1. 1Standard care without supportServes as the comparator
  2. 2AI supportMeasures the effect of introducing AI
  3. 3AI support plus a structured verification layerMeasures the effect of mechanically checking the AI's output

Against the question of reducing AI error, it tests an answer that places a separate layer verifying output rather than using a better model. The primary endpoint is a binary composite per encounter, covering absence of mention of a must-not-miss diagnosis, under-triage and contraindications. Unless what counts as failure is settled first, whether support helped cannot be judged.

Why it matters

Addressing AI error is not only a matter of switching to a better model. A design that layers mechanical inspection over the output can be measured independently of model improvement. Defining failure in advance and separating unassisted, assisted and verified arms is a useful reference for structuring how the effect of an AI deployment is tested.

FAQ

Why three arms?
To separate the effect of using AI from the effect of layering verification on its output, comparing unassisted care, AI-assisted care, and AI-assisted care with the verification layer.
What counts as a safety failure?
A per-encounter binary composite including failure to mention a do-not-miss diagnosis, under-triage and contraindication, per the record.
What were the results?
The trial is recorded as completed, but how far failures fell in either arm is not part of this registry record.

Sources (primary)

Source: ClinicalTrials.gov (U.S. NIH/NLM, public domain). This site does not provide medical advice. Verify the latest and exact details with the official source. This site is not endorsed or certified by the NIH/NLM.

#Clinical trials#AI#Healthcare#Decision support#Safety
Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.