The part that starts after deployment — a registry linking AI pathology scores from routine care to patient outcomes over twelve years — a clinical trial (ClinicalTrials.gov)
A multi-site prospective observational registry that links the results of AI-based pathology testing performed as part of routine clinical care in patients with solid tumors to their short- and long-term outcomes, evaluating real-world performance and clinical utility. The design allows further analyses and investigator-initiated ancillary studies to be added later, and completion is planned for December 2038.
Trial overview (primary data)
- StatusNot yet recruiting
- ConditionsSolid Tumor Malignancy, Breast Cancer
- InterventionsDIAGNOSTIC_TEST: Ataraxis testing
- SponsorAtaraxis AI, Inc.
- Target enrollment1,000 participants
- Period2026-07-31 〜 2038-12-01
Key points
- A multi-site prospective observational registry linking AI pathology test results produced during routine care in patients with solid tumors to their short- and long-term outcomes, with a target enrollment of 1,000.
- With no control arm, it cannot show that using the AI is better; it shows how closely the scores produced correspond to what actually followed.
- The adaptable design allows sponsor-led analyses and investigator-initiated ancillary studies to be added later. Of the 803 records this site holds as of 2026-09-04, this is the only one mentioning such an extension.
- Running from July 2026 to December 2038, about 12.3 years. Only 7 of those same 803 records (0.9%) have a completion date in 2038 or later.
- This is not the rung that proves performance but the rung that watches whether proven performance survives in practice.
1Doing well inside a trial guarantees nothing outside one
Every rung so far served to measure performance under arranged conditions: choose the cases, fix the comparator, fix the order. In actual practice nothing is chosen. Specimen quality varies, and so does how and when the test gets ordered. What this registry gathers is not testing performed for a study but testing already performed as part of routine care, linked afterwards to what happened to those patients.
2What changes
There is no control arm. So this design cannot say that using the AI is better. What it can say is how closely the scores that were actually produced correspond to what actually followed. It is not the rung that proves performance; it is the rung that watches whether proven performance survives contact with the floor.
3Twelve years long
Cancer outcomes surface over years, so linking to long-term outcomes takes a corresponding stretch of time. The registry also uses an adaptable design in which analyses and investigator-initiated ancillary studies can be added after it opens. Among those same 803 records, exactly one mentions that kind of later extension, and it is this registry.
The longer something runs, the more questions arrive that were not the original question. Room for them has been written into the design.
4The eight rungs, side by side
Across these articles, trials of medical AI have been read as eight rungs. Laying out how often each rung appears in the records here shows where the field is thick and where it is thin. The denominator is the 803 records this site holds as of 2026-09-04.
| Rung | The question | Records mentioning it, of those same 803 |
|---|---|---|
| 1 Gather | Can the material for measurement be built? | 74 (9.2%) |
| 2 Look back | Could it have been caught earlier in past data? | 75 (9.3%) |
| 3 Check separately | Does it hold outside where it was built? | 13 (1.6%) |
| 4 Compare with people | Can it match specialists on the same footing? | 12 (1.5%) |
| 5 Randomize | Do outcomes change when care changes? | 170 (21.2%) |
| 6 Ask about replacement | Can it be called not worse? | 12 (1.5%) |
| 7 Measure in the field | Does it hold inside daily operations? | 13 (1.6%) |
| 8 Watch | Does it hold after deployment begins? | 11 (1.4%) |
The thick rung is randomized comparison; the thin ones begin at external checking and continue to the end. Most medical AI climbs as far as measuring accuracy, while whether it works elsewhere, whether it may replace anything, and whether it holds after deployment are still barely asked. That distribution is itself a description of where medical AI currently stands.
Why it matters
Performance shown under arranged conditions does not automatically survive where conditions are not arranged. A registry that keeps linking results produced in routine care to actual outcomes is the mechanism for watching that gap. Few of the clinical trials held here reach this rung, which shows how heavily the evaluation of medical AI is weighted toward accuracy and how lightly toward monitoring after deployment.
FAQ
Will this registry prove the AI is effective?
Why does it take twelve years?
What is an adaptable design?
Sources (primary)
Source: ClinicalTrials.gov (U.S. NIH/NLM, public domain). This site does not provide medical advice. Verify the latest and exact details with the official source. This site is not endorsed or certified by the NIH/NLM.
- ClinicalTrials.gov (study record, original)
- NCT ID: NCT07706842