What is radiology peer review supposed to accomplish?
Peer review looks for patterns in reading that no single case would reveal: a modality where misses cluster, an anatomy that gets skimmed at the end of a list, a report style that leaves the ordering clinician guessing. It is a quality instrument aimed at the group. It is not a re-diagnosis of the patient, and it is not a ranking of radiologists.
Two questions get tangled together constantly. The clinical question belongs to the patient's care team and gets settled by follow-up imaging, biopsy, surgery, or time. The quality question asks whether the reading process is working across hundreds of studies. Peer review answers the second one. When a reviewer opens a study read months earlier and sees something the first reader did not describe, the patient's situation has usually already resolved one way or another through some other route. What the program is there to extract is whether that miss belongs to a category worth acting on, and whether anyone else in the group is making the same one.
The ranking problem deserves saying out loud. As soon as scores are attached to individual radiologists and lined up next to each other, the program stops measuring reading quality and starts measuring who submits their hard cases and who reviews generously. Groups that publish per-reader tables get exactly what they built: a very clean spreadsheet and no usable signal. The committees that hold onto honest data are the ones that treat the group aggregate as the unit of analysis and leave the individual pattern to a private conversation between colleagues. Framing at the first meeting decides most of this. A group that opens by putting individual scores on a screen has already told everyone what the program is for.
Is random sampling or targeted review the better design?
Neither works alone. Random sampling estimates how the group reads on an ordinary day and gives a baseline you can defend. Targeted review goes where suspicion already exists: a modality drawing complaints, a subspecialty stretched thin, cases whose outcome turned out differently. Random sampling produces a rate. Targeted review produces a cause. Reporting them as one number destroys both.
Here is what quality committees skip. A random sample is only random with respect to the pool it came from, and the pool is usually whatever the review tool could reach without extra work. If it fills up with routine chest radiographs and follow-up studies that have clean priors sitting next to them, the discrepancy rate will come out low and it will say almost nothing about how the group handles a difficult CT at three in the morning. A low rate can be a genuine finding. It can also be an artifact of an easy denominator. Before anyone celebrates, the committee should be able to describe the case mix behind the number by modality and by acuity.
Targeted review has the mirror problem. Because it starts from suspicion, it finds what it went looking for, and no rate it produces means anything. The reviewer also knows the case was flagged, which colors what they see before they open it. Outcome-linked review, where a study gets pulled because surgery or a later scan showed something, teaches more per case than anything else in the program and carries the heaviest hindsight bias in the program. Both of those are true at once. What most groups settle on is a standing random sample large enough to trend, plus targeted pulls discussed on their own terms and kept out of the headline number entirely.
One practical detail sinks more programs than either design flaw. The review tool sits in a separate worklist that nobody opens, or it opens only when an administrator sends a reminder near the end of the quarter. What arrives then is a batch of cases scored in one sitting by someone trying to finish. A sample collected that way is not random in any useful sense, and the scores carry whatever mood the reviewer was in. Putting the review step inside the reading worklist, so a case appears the way any other case does, fixes more than a better scale ever will.
- Random sample: state the pool, the case mix, and the sampling interval, not just the percentage reviewed.
- Targeted pulls: record why the case was selected, so the hindsight is visible during the discussion.
- Outcome-linked cases: excellent for learning, unusable as a rate.
- Report the two streams separately. Blending them makes both numbers meaningless.
How do peer review scores work, and why are they contested?
Most scales ask the reviewer to place a case on a short ordinal ladder running from full agreement, through a minor difference, to a discrepancy likely to have affected care. The idea is reasonable. Reliability is the problem: hand the same case to two reviewers and they often pick different rungs, so small differences between readers or between quarters are mostly noise.
The weakness is structural, so picking a better scale does not fix it. An ordinal category compresses two separate judgements into one number: how large the perceptual or interpretive difference was, and how much it mattered to the patient. Those two come apart all the time. A reader can miss a small nodule that changes nothing for years, and a reader can hedge an obvious finding into language so vague that the clinician goes the wrong direction, without technically missing anything. A single one to four rating forces the reviewer to average two unlike things, and every reviewer averages them differently. Scoring discrepancy and clinical impact on separate axes is more work at the workstation and produces something a committee can actually read.
Calibration helps and almost nobody does it. A program that wants scores capable of surviving scrutiny should periodically send the same handful of cases to every reviewer and compare what came back. The groups that try this are usually unhappy with what they find. The fix is not a longer rubric with more adjectives in it. It is worked examples: a small library of scored cases with the reasoning written out underneath, so a reviewer can see what a 2 looks like instead of inferring it from a word like minor. The same library orients a new reviewer without anyone having to convene.
Who scores the case matters as much as the scale. A general radiologist reviewing a musculoskeletal MRI read by a fellowship-trained musculoskeletal radiologist will score conservatively, or will flag something the subspecialist would not. Run that mismatch across a quarter and the numbers drift for reasons that have nothing to do with reading quality. Programs that match reviewer to subspecialty get cleaner data and immediately hit the second problem, which is that in a small group the obvious reviewer is the person who reads the same list and knows exactly whose case it is.
When is a difference an error, and when is it a defensible difference of opinion?
Ask what a reasonable radiologist could have concluded from the same images, the same history, and the same priors at that hour. If the finding is visible in retrospect but the original reading was defensible on the information available, that is disagreement. If the finding was there to be seen and the report does not account for it, that is a miss.
Hindsight is the whole difficulty. A reviewer who already knows how the case turned out will see the finding faster and will find it genuinely hard to believe anyone missed it. The effect is well described and it does not switch off because the reviewer is aware of it. Serious programs push back in small ways: reviewing without the outcome attached where the workflow allows it, recording whether the reviewer knew the result, and requiring the reviewer to write down what information the first reader actually had. A study read with no priors, no usable history, and a four word indication line is not the same study the reviewer is looking at with the chart open beside it.
What the report said matters here as much as what the radiologist saw. A report listing three possibilities without ranking them can be technically unfalsifiable and still leave the clinician unable to do anything. Some of the most useful findings a review program produces are not misses at all. They are reports that were correct and unusable. Committees that only count misses never surface those, because the ordinal scale has no rung for a correct report nobody could act on. This is one reason the free-text comment field, if somebody actually reads it, is worth more than the score sitting beside it.
How does a discrepancy meeting turn a finding into something a reader can use?
By splitting two jobs that usually get muddled together: establishing what happened in the case, and deciding what the group is going to change. The first is a short factual discussion. The second produces a named owner and a date. A meeting that only does the first generates minutes and no change, and attendance drops.
Here is the failure almost every program hits. The reviewer scores a case, the committee discusses it, the minutes record a thoughtful conclusion, and the radiologist who read the study never hears a word about it. Sometimes that is deliberate, because someone decided anonymity protects the process. Sometimes it is nobody's job and everyone assumes it is somebody else's. Either way the loop is open, and an open loop cannot change how anyone reads. The reader learns nothing, repeats the pattern, and the same category turns up in next quarter's summary looking like a stubborn systemic problem when it is one person who was never told.
Closing the loop is unglamorous administrative work. Someone has to be named as the person who delivers feedback. There has to be a route that reaches the reader in weeks rather than at the annual review. And the message has to include the images. An email saying a case was scored a 3 teaches nobody anything. Sending the study, the original report, the reviewer's comment, and an invitation to look at it again teaches a great deal, including the occasions when the original reader turns out to have been right and the reviewer wrong. Record those too. A program that has never once overturned a reviewer is not being careful with its own data.
Then there is the part people avoid saying in the room. The moment radiologists believe review scores feed into contracts, privileges, or a credentialing file, submissions dry up. Nobody flags their own uncertain case. Reviewers score gently, because they know what a bad score costs a colleague they will sit next to on Monday. The numbers improve and the program gets worse. Learning review and performance management have to stay apart, and readers have to be told plainly which one they are sitting in. Where a real competence concern shows up it belongs in a separate documented process with its own safeguards, and how that process is built is a question for the facility's own governance and its counsel.
What should a facility expect a reading service to describe about its own review process?
A facility should be able to hear, in specific terms, how cases are selected, who reviews them, what scale is used, how a discrepancy reaches the original reader, and what happens when a review turns up something serious. A service that answers with a percentage and nothing else has told you nothing you can check.
The useful questions are procedural, and anyone running a real program answers them without preparing. How is a case pulled, and out of what pool? Does the reviewer see the outcome? What happens when the reviewer and the original reader disagree about whether there was a discrepancy at all? Who tells the reader, and how long does that take? What triggers an amended report, and how does the facility learn that one was issued? Where does a serious finding go, and who on the facility side receives it? If any of those answers have to be improvised, the process is probably not written down anywhere.
Then ask about the numbers with the case mix attached to them. A rate quoted without a denominator, a modality breakdown, and a sampling method should be read as marketing copy. Ask what the rate was for the categories the facility actually sends, which is often a much narrower slice than the whole pool. Ask whether targeted and random pulls are reported separately. Ask how many reviewed cases ended in an addendum, and by what route. None of this asks a service to hand over identifiable cases or internal personnel matters, and a reasonable service says where that line sits instead of declining the conversation. This guide describes general practice and states nothing about DLA Imaging's own review design, committee structure, or results.
- Case selection: the pool, the sampling method, and how targeted pulls are flagged.
- Reviewer assignment: subspecialty match, and whether the outcome is visible during review.
- The scale in use, and whether discrepancy and clinical impact are scored on separate axes.
- The route from a scored discrepancy back to the radiologist who read the study, with a time expectation attached.
- What triggers an amendment or a direct call to the facility, and who receives it.
- How random and targeted results are reported separately.
Sources and scope
- American College of Radiology, RADPEER peer review program
- American College of Radiology, practice parameter for communication of diagnostic imaging findings
- The Joint Commission, standards for ongoing professional practice evaluation
- American Board of Radiology, practice quality improvement within continuing certification
- Radiological Society of North America, quality and safety education resources