Models that read medical images perform well in published evaluation and appear in routine care far more slowly. The delay is caused by requirements that sit outside the model's performance entirely.

Why imaging was the first target

Medical images are digital, standardised in format, produced in enormous volume and stored with an associated diagnosis.

That combination gives exactly the labelled dataset supervised learning requires, which almost no other part of medicine supplies.

The task is also well specified, since detecting a defined abnormality in a defined image type is a bounded problem with a clear answer.

What approval requires beyond accuracy

Regulators assess a device for a stated purpose in a stated population, which means the claim determines the evidence required.

A narrow claim about flagging one finding on one scanner type is approvable on modest evidence, and a broad diagnostic claim requires prospective clinical study.

Most cleared systems therefore make narrow claims, which is why the approved uses look far less impressive than the published research.

Why performance drops on arrival

A model trained on images from a small number of institutions learns the characteristics of their equipment, protocols and patient population.

Deployed elsewhere, on different scanners with different settings and a different case mix, accuracy falls, sometimes substantially.

This is why local validation before deployment has become standard practice rather than an optional check, and it takes months at each site.

How workflow determines whether it is used

A radiologist reading a high volume of studies will not use a system requiring a separate login or an extra step.

Systems that succeed appear inside the existing reading environment, marking findings on the image the clinician is already looking at.

Integration with the imaging systems in use is consequently a larger engineering effort than the model, and it is specific to each institution.

Where liability slows adoption further

A clinician remains responsible for the diagnosis, so a flagged finding must be reviewed and an unflagged region cannot be skipped.

That preserves accountability and removes the time saving, since the reading still happens in full and the alert is additional information rather than a substitute.

The uses that have spread fastest reflect this precisely: triage systems that reorder a worklist so urgent cases are read sooner, changing when a human looks rather than whether one does.