What kind of AI tool are you actually being sold?
Three families, and they are not interchangeable. A triage tool changes the order of a worklist. A detection aid marks a location and asks a human to look. A generating tool produces text or a measurement that can end up in the report. Risk, oversight, and the documentation you should demand all follow from which one you have.
A triage tool says nothing about what the study shows. It moves a case up a list so a human sees it sooner. Vendors sell this as free upside, since nothing is removed and no interpretation is offered, and the pitch is genuinely attractive on a night shift. The part that goes unsaid is the other end of the list. A worklist has a finite length, and a case pushed up pushes another case down. If the tool is silent on a study that mattered, that study is now read later than it would have been under a simple chronological queue. Ask the vendor what happens to unflagged cases, and ask your own team whether anyone would notice a systematic delay.
A detection aid puts a mark on the image and hands the decision back. The reader accepts it, ignores it, or calls it something else. That sounds like a clean division of labor until you ask where the disagreement goes. In most installations it goes nowhere. The radiologist dismisses a box, signs the report, and the only record that the tool was wrong lives in one person's memory until the end of the shift. Nothing reaches the vendor, nothing reaches your quality committee, and whoever reviews the contract later has no local evidence to argue with. Build the capture before the tool goes live. Afterward nobody has time to retrofit it, and the cases you could have learned from are already signed.
The third family is the one to be strictest about. A tool that populates a measurement field, drafts an impression, or writes a sentence into the report is producing content that a physician then signs. Where the boundary sits between decision support and something closer to interpretation is argued about constantly, and the argument is not settled by the vendor's marketing category. Decide locally what your readers are permitted to accept without independent verification, write that decision down, and make sure the reporting system does not quietly default to the opposite, because the shipped default is what will govern in practice.
What does regulatory clearance tell you, and what does it not?
It tells you a regulator reviewed the tool for a stated indication, in a stated population, under stated conditions of use, and that the maker met a bar for that claim. It does not say the tool is good, that it will perform on your patients, or that using it outside the cleared indication becomes someone else's problem.
Read the indications-for-use statement rather than the brochure. That sentence names the modality, the body region, the intended role of the software, and often the patient population. It will say whether the output is an adjunct, whether it is meant for notification only, and whether it may be used as a primary reading tool. This is the sentence that governs, and it is usually shorter and narrower than the sales conversation around it. Facilities get into trouble by buying a device cleared for prioritization and then treating its silence as a negative read, or by running a tool cleared in adults across a pediatric service because the software does not refuse.
Clearance pathways differ, and the evidence behind two cleared products can differ enormously. One may rest largely on a comparison with a device already on the market plus a retrospective performance study; another may carry a reader study designed to show something about how radiologists actually behave with the tool on screen. Both are legitimate. They are not the same evidence, and neither is a prospective study at a site that resembles yours. Ask which pathway applied and what was submitted, in plain terms, and be skeptical of an answer that stops at the word cleared.
Then ask about versions, because this is where quiet trouble accumulates. Models get retrained. Thresholds get tuned. A vendor may push an update the way any software vendor does, and the facility can end up running a different model against the same clinical policy without a single meeting having taken place. Require that version changes are announced before they land, that someone at your organization approves clinical use of a new version, and that the version in force is recorded against the case. If the contract does not say this, the default is silence.
What was the model validated on, and does that population look like yours?
This is the question buyers skip. Ask for the number of sites, the scanner makes and configurations, the acquisition protocols, the case mix, how the reference standard was established, and who read the ground truth. Then ask how many of those cases came from a place that looks anything like yours.
Test sets are often enriched with positive cases so that performance can be measured without collecting an impractical number of studies. That is standard practice, not misconduct, and any competent vendor will say so if asked directly. What follows from it is arithmetic that gets left out of the slide deck. Sensitivity and specificity travel between populations reasonably well. Positive predictive value does not, because it depends on how common the finding is in the group being scanned. A tool that looks precise on an enriched set can produce a great many false alarms in a screening population where the finding is rare, and your readers will feel that in the first weeks, long before anyone writes it down.
The other mismatch is technical, and it is easier to check. Scanner manufacturer, model, reconstruction kernel, slice thickness, contrast timing, portable versus fixed equipment, the age structure of the people who actually come through your department: each can differ from the validation sites in ways that matter. A model built on studies from a handful of large academic centers has seen a particular kind of acquisition, and your technologists can usually tell you in a minute how far your own protocols sit from that. Put the request in writing and treat the gaps as information. A vendor who cannot break performance out by site or by equipment is telling you something real about how the evidence was assembled. Refusing to describe the sites and the equipment at all is a different answer, and you should hear it as one.
- The number of independent sites that contributed cases, and whether any resemble yours in equipment and case mix.
- Scanner manufacturers, models, and the acquisition protocols the validation studies were produced under.
- How the reference standard was established, by whom, and whether those readers could see the model output.
- Whether the test set was enriched with positive cases, and what the finding's frequency was in the source population.
- Performance broken out by subgroup: age, sex, body habitus where relevant, equipment, and site.
- Whether any validation site was also involved in developing the model.
What happens to the tool's performance after go-live?
It moves, and usually nobody notices. A protocol edit, a scanner replacement, a new reconstruction setting, a shift in who walks through the door: any of these can pull a model away from the conditions it was built for. The software keeps producing confident output the whole time. Local monitoring is the only thing that catches it.
Here is the failure mode in its most ordinary form. A technologist team standardizes on a thinner slice, or switches to a different reconstruction kernel, for a perfectly good clinical reason. The change is discussed inside the department, documented in the protocol book, and never mentioned to anyone thinking about the software. The model was tuned on the older configuration. Over the following weeks the flag rate drifts, and because it drifts slowly, it reads as a quiet month rather than a problem. Nobody connects a protocol decision in one meeting to a detection tool in another. That is not carelessness. It is what happens when the tool has no owner who sits in both rooms.
Monitoring does not require a data science team. It requires a number, a chart, and somebody whose job it is to look. Track the flag rate over time, split by modality and by scanner, and know what it looked like in the first month. Sample cases both ways: flagged studies against the final report, and unflagged studies against the final report, because the second sample is the one that tells you about misses. Record reader agreement where the workflow allows it. Set a value at which somebody has to ask a question, and name that somebody. A chart that nobody is responsible for reading is decoration.
Decide in advance who can switch the tool off. It should be a named clinical role with authority to suspend use immediately, not a decision that waits for a procurement cycle or a call to the vendor. Write the trigger conditions next to the name, so the person is not improvising under pressure. Then plan for the outage that will eventually happen. Once readers have adapted to a tool being present, its sudden absence changes how a shift runs, and the department that has never worked a night without it will find that out on the worst possible night. Treat that as a workflow question with a written answer, not as an IT ticket.
How does the output reach the reader, and can the reader ignore it?
Look at the actual screen before you sign anything. A colored box burned into the images, a separate series, a worklist badge, and a line pre-populated in the report are four different levels of pressure on a radiologist. The question is not whether a reader may disagree with the tool. It is whether disagreeing costs them anything.
Automation bias runs in both directions, and the second direction is the dangerous one. A visible mark pulls the eye toward the marked spot and, by the same mechanism, away from the rest of the study. That much is intuitive. The harder half is silence: readers begin treating the absence of a flag as reassurance, even for a tool that was never designed or cleared to rule anything out. A tool that is right most of the time is exactly the tool that teaches this habit, and it teaches it faster to a tired reader at the end of a long list. No policy sentence about maintaining independent judgment survives contact with that effect. The display design has to do the work instead, and that includes the numbers: a confidence score printed to a decimal implies a precision the model may not have.
Ask the concrete questions during the demonstration, with a radiologist in the room. Can the reader see the study before any overlay appears? Can the overlay be toggled off and back on? Is the output stored as a separate series, so annotations never contaminate the diagnostic images? Does the reporting integration insert text the radiologist must delete to disagree, or text the radiologist must actively add to agree? That last one looks like a preference setting and behaves like a policy. Opt-out text produces signed reports nobody quite intended to sign. Then give the reader somewhere to put a disagreement: one field, two clicks, no free text required. Without it, your only evidence about the tool's behavior at your site is a vendor dashboard built from the tool's own output.
Who is responsible for a missed finding when a tool was in the loop?
The interpreting physician signs the report, and that does not change because software was on screen. What changes is that a second set of records now exists, and someone will read it later. So decide in advance what the tool's role is, write it down, and keep the documents that show how the decision to deploy it was made.
The awkward part is that the same log supports two mirror-image arguments. If the tool marked something and the reader dismissed it, the log shows the reader was pointed at it. If the tool said nothing and the finding was missed, the reader's belief that the software agreed does not move responsibility anywhere. Facilities sometimes buy a tool half expecting it to absorb risk. It does not work that way, and a contract clause about vendor limitations of liability is worth reading with counsel before signature rather than after an incident. This article is general education, not medical or legal advice.
What protects a facility is an ordinary governance record, kept from the beginning. Who approved this tool, on what evidence, for which studies, and with what stated limits. What training the readers received and when. What is being monitored and by whom. There is also a set of data questions that tend to get answered in a sales conversation and never in the contract: where inference happens, who holds the images while it happens, how long they are retained, and whether your studies may be used to develop future models. Get those into the agreement in the same words the vendor used out loud.
A boundary about this article. DLA Imaging publishes no artificial intelligence product, no validation claim, and no deployment claim on this website. What is described above is how a buyer evaluates a category of software, not an account of anything DLA Imaging operates. If a reading or teleradiology service tells you a tool sits somewhere in its workflow, the same documents apply to that service, and you are entitled to see them before the first case rather than after a problem. A vendor relationship and a professional services relationship are not the same thing, but the paperwork question is.
- A written statement of intended use at your site, naming modality, body region, and the populations the tool is and is not approved for.
- The indications-for-use language from the clearance, placed next to your own statement so any difference is visible.
- The validation evidence you were given, including a note on what the vendor could not provide.
- The monitoring plan: what is measured, by whom, how often, and the value that triggers a review.
- A named clinical role with authority to suspend use, plus the procedure for working without the tool.
- Data handling terms: where inference happens, retention, and whether your images may be used for future model development.
- Version control: how model updates are announced, approved, and recorded against individual cases.
Sources and scope
- American College of Radiology, professional guidance on the assessment and clinical use of artificial intelligence in medical imaging
- Radiological Society of North America, professional education on the evaluation and reporting of artificial intelligence in radiology
- American Association of Physicists in Medicine, professional guidance on imaging equipment performance and quality control
- DICOM Standards Committee, Digital Imaging and Communications in Medicine standard
- IHE International, profiles for imaging workflow and results integration