What is BI-RADS actually for?
It is the American College of Radiology's reporting and data system for breast imaging. It fixes the vocabulary a report uses to describe findings, requires a single final assessment category, and attaches a management recommendation to that category. The purpose is a report a referring clinician can act on without having to interpret the radiologist's prose.
Radiology reports drift toward description. A mass is round, circumscribed, equal in density to the surrounding tissue, unchanged in position. All true, and all useless to the family physician who has to tell the patient something this afternoon. Breast imaging solved this earlier and more strictly than the rest of radiology, because the volume is high, the population is largely asymptomatic, and the consequences of an ambiguous last line are carried by someone who will never speak to a radiologist. The structured lexicon does part of the work. The assessment category does the rest, and it is the part with operational weight, because the systems that receive the report read the category and not the paragraph above it.
The lexicon matters more than it looks. Shape, margin, density, distribution, and the composition of the breast tissue itself all have defined terms, and a facility with four readers gets four different reports without them. Composition is the clearest case: it is reported on a defined scale, it governs what the reader can and cannot see on that study, and it is the sort of statement that gets compared across years and across readers. Where the vocabulary is loose, comparison stops working. A finding one reader calls prominent and another calls a focal asymmetry is the same finding in two languages, and nobody can tell whether it changed.
What BI-RADS does not do matters too. It is not a risk model, it is not a diagnosis, and a category is not a probability the reader calculated for that patient. It is an assessment with a recommendation attached, made under a shared definition so that the next reader, the audit, and the referring office all mean the same thing by it. Facilities occasionally treat the category as a number that speaks for itself and stop reading the report. The category tells you what happens next. The report tells you why.
What do the assessment categories do that a description cannot?
They end the report in a decision. Each category states how suspicious the study is and, with it, what should happen: nothing beyond routine screening, a short-interval follow-up, additional imaging, or tissue sampling. One report gets one final assessment. That constraint is what makes the report trackable and usable by systems that never read the prose.
The incomplete category generates most of the operational traffic. It says the reader cannot finish the assessment from what is in front of them and needs something more: additional views, an ultrasound, or a prior study that has not arrived. In screening, that is a normal and expected outcome, and it is the trigger for a recall. In the diagnostic setting it should be rare, because the whole point of a diagnostic study is that the additional imaging is happening now, with the patient still in the department. A diagnostic report that ends incomplete usually means something in the workflow broke, and the report is where the break becomes visible.
The rule that a report carries one final assessment sounds bureaucratic until you see what happens without it. The body of the report describes a suspicious mass and the report closes with a benign category, because the reader dictated the category first out of habit and then talked through the finding. The prose and the code now disagree. The tracking system reads the code. The patient letter template reads the code. The audit reads the code. Whoever eventually notices the disagreement is reading the chart for an unrelated reason months later. This is testable in an afternoon: pull reports whose text contains suspicious language and whose assigned category is benign or negative. The sample is rarely empty.
- Incomplete: the assessment cannot be finished without additional imaging or a prior comparison.
- Negative: nothing to report, and routine screening continues.
- Benign: a finding is present and it is benign, and routine screening continues.
- Probably benign: very low suspicion, with short-interval follow-up rather than immediate sampling.
- Suspicious: subdivided by degree, with tissue diagnosis recommended.
- Highly suggestive of malignancy: the category carries a recommendation that appropriate action follow.
- Known biopsy-proven malignancy: used for studies of a cancer already established by tissue.
Why does breast imaging carry requirements that other reading areas do not?
Because mammography is regulated at the federal level, and the requirements land on the individual physician. Initial qualification, continuing experience in the form of a minimum volume of interpretations, and continuing education specific to mammography are counted per physician over defined intervals. A practice cannot satisfy them collectively on a reader's behalf, and a facility has to hold the documentation.
Three layers stack here and they are not the same layer. Federal regulation sets the floor for who may interpret a mammogram and what a certified facility must do. The accrediting body the facility works through adds its own program requirements and its own review. Facility policy and medical staff bylaws sit on top and can be stricter than either. State law and payer contracts sometimes add a fourth. The counts, the intervals, and the documentation formats all live in those instruments and they get revised, which is why nothing in this article should be read as the current threshold. Confirm the numbers against the rule as it stands and against your accrediting body's current program requirements.
The failure is almost always the same, and it is a gap rather than a mistake. A facility contracts with an outside reading group. The facility assumes the group verified its readers' mammography qualification, because the group employs them. The group assumes the facility verified it, because the facility holds the accreditation and the credentialing file. Both assumptions are reasonable. Neither is written anywhere. The mismatch surfaces at the worst time: during an accreditation review, or when a study lands on a worklist and someone finally asks who is allowed to open it. Ask for the qualification record by name, physician by physician, with dates, and put a re-verification interval in the agreement. Continuing experience is a rolling count, so a reader who qualified last year is not automatically qualified this year.
One boundary belongs here in plain words, because this is a public website. Nothing on this page describes any DLA Imaging reader's qualification, credentialing status, or breast imaging scope, and nothing here should be read as a claim that a particular study type falls within scope. Those questions are answered by a facility agreement and a credentialing file, with documents rather than marketing pages. Accreditation works the same way. A website can describe what a program is required to hold; it cannot stand in for the certificate, and a facility should not let it. A facility evaluating any reading arrangement should ask for the qualification record directly, and it should arrive per physician.
Why does a missing prior change the read?
Because a large part of screening interpretation is change detection. A finding that has been sitting in the same place, the same shape, and the same size for years reads differently from the same finding seen for the first time. Without a comparison, the reader has one time point, and a study that would have been resolved on sight becomes a recall.
The cost of that recall is not evenly distributed. The facility absorbs a slot, the reader absorbs the second read, and the patient absorbs a phone call asking her to come back, which she will spend the intervening days interpreting for herself. Nobody in that chain wanted the recall. It happened because an archive did not deliver a study that existed. The reverse error is quieter and worse: a reader who assumes the comparison was performed when no prior was ever loaded, and writes a stability statement the images cannot support.
Here is where it actually breaks. The priors exist. They are in the archive, they are indexed, and someone will tell you with complete confidence that they were available. What did not happen is delivery. The pre-fetch rule matched on an identifier that changed when the patient's name changed, or the outside facility's release process takes days, or the images arrived and the prior reports did not, which is a partial comparison at best. A remote reader has no way to distinguish a patient with no prior imaging from a patient whose prior imaging failed to route. The worklist should say which one it is. Present, absent, or still transferring, stated explicitly, with the reader able to hold a study rather than read it blind.
Two practices help more than they should have to. Give the reader a documented way to request a prior and to record that the request was made, so a study read without comparison shows why on its face. And write down what happens when the prior arrives after the report was signed. That is an addendum in some cases and an amendment in others, and a breast imaging program that has not settled the difference in advance will settle it under pressure. Five o'clock on a Friday is a poor time to be deciding it for the first time.
Why do screening and diagnostic studies route differently?
They answer different questions with different exams. Screening looks for disease in a person with no symptoms, uses a standard set of views, and can be interpreted after the patient has gone home. A diagnostic study investigates something specific and is built as it goes, which usually needs a radiologist who can change the exam while the patient is still there.
That difference decides almost everything about routing. Screening volume batches well. It can sit in a worklist, it can be read in a block, and the turnaround expectation is measured against a schedule rather than against a patient waiting in a gown. Diagnostic work does not batch. A directed study frequently becomes an ultrasound, then a targeted view, then a decision about sampling, and each step depends on what the last one showed. If the reader cannot influence the next step in something close to real time, the patient goes home and comes back, and a workup that could have finished in one visit spreads across two.
What breaks here is the booking, and it arrives disguised as a reading failure. A patient with a palpable lump gets booked into a screening slot, because the phone call said mammogram and the screening slot was the one that was open. The standard views are performed. The reader sees a screening study, sees nothing, and reports a negative assessment, which is accurate for the images and wrong for the patient, whose symptom was never worked up and does not appear in the history field. Every breast program knows this can happen. Fewer have a stop between the booking and the machine: an intake question that changes the exam type, and a technologist authorized to change it without restarting the order cycle.
Any arrangement that reads across locations has to settle this in the scope document rather than on the fly. Decide which study types actually belong in the remote workflow instead of assuming the whole modality does. Then write down what happens when a screening study turns out to need diagnostic work the same day. That second question is the harder one, and the answer comes from the local team's availability rather than from the reading arrangement. A facility that cannot produce a physician for the tailored part of a workup has a coverage problem, and no amount of remote reading capacity fixes it. Say so up front instead of discovering it on the first recall.
What does outcome tracking require of a breast imaging program?
A closed loop between what the reader assessed and what turned out to be true. The program has to collect biopsy and surgical results, match them back to the interpretation that recommended the sampling, and analyze the result by facility and by individual interpreting physician. The audit is a requirement for a certified facility, and it is also the only real feedback a reader gets.
The loop breaks at the return path, and it breaks for a structural reason. The pathology result goes to whoever ordered the biopsy, and the surgical finding goes to the operating team. Neither of them owes the imaging department a copy. So the interpretation that started the whole sequence sits in the radiology system with no outcome attached to it, and the audit gets assembled at the end of the year by someone calling offices and reading operative notes. Programs that run well have a named person whose actual job includes chasing those results, week by week, while the case is still recent. Programs without that person produce an audit built on whatever came back on its own, which is a biased sample by construction.
Denominators cause more argument than the numerator ever does. A positive predictive value calculated over everything a reader sent to sampling is a different number from one calculated over everything recalled from screening, and the two get compared to each other in meetings anyway. Fix the definitions in writing before anyone sees a result, and publish the denominator next to every rate. The per-physician requirement adds a second problem when readers work across more than one facility: each site holds a fragment of that reader's work, and no site sees the whole picture. Somebody has to decide whether the reader's own aggregate is tracked anywhere, and who is entitled to see it.
The honest objection is that audits become filing. They do, wherever the output is a number that goes into a binder. The version that changes anything is small and frequent: a reader sees the outcomes of their own recalls and their own sampling recommendations close enough in time to remember the case, and the review committee looks at the cases that presented clinically after a negative or benign assessment. That second group is uncomfortable to look at, which is exactly why it is the group that teaches. A program that only reviews its true positives is measuring its filing, not its reading.
- The interpretation and its assessment category, with the reader identified.
- Whether the patient was recalled, and what the diagnostic workup found.
- The pathology result for every case sent to sampling, including the benign ones.
- Surgical findings where they differ from the core sample.
- Cases that presented clinically within the interval after a negative or benign assessment.
Sources and scope
- American College of Radiology, Breast Imaging Reporting and Data System and breast imaging accreditation program requirements
- U.S. Food and Drug Administration, Mammography Quality Standards Act requirements for certified facilities and interpreting physicians
- Society of Breast Imaging, professional guidance on breast imaging practice and reporting
- Radiological Society of North America, professional education on breast imaging interpretation and reporting
- American Association of Physicists in Medicine, guidance on mammographic image quality and equipment performance testing