How RADS Categories and Scoring Work in Radiology

RADS, short for Reporting and Data Systems, are standardized classification frameworks used by radiologists to score what they see on imaging studies and translate those findings into clear risk categories. The original and best known is BI-RADS (Breast Imaging Reporting and Data System), developed by the American College of Radiology in the late 1980s to bring consistency to mammography reports. Since then, the RADS concept has expanded to cover the lungs, prostate, liver, thyroid, ovaries, bladder, and soft tissues, with each system following a similar logic: assign a numbered category to an imaging finding, link that category to a recommended next step, and give every doctor reading the report a shared vocabulary.

Where the RADS Concept Came From

Before BI-RADS existed, mammography reports varied wildly from one radiologist to another. Two doctors could look at the same image and describe it in entirely different terms, leaving the referring physician to guess how worried they should be. The ACR launched its BI-RADS initiative in the late 1980s specifically to fix that problem, creating a lexicon of standardized descriptors for specific imaging features so that a “spiculated mass” meant the same thing in every report, everywhere.1PubMed Central. The ACR BI-RADS experience: learning from history The first formal edition was published in 1993, and the system has been updated several times since, most recently to a fifth edition that extended coverage beyond mammography to include breast ultrasound and MRI.2PubMed. A Pictorial Review of Changes in the BI-RADS Fifth Edition

BI-RADS proved so useful that the ACR began applying the same template to other organs. The basic architecture repeats: a dictionary of descriptors, a numbered scale of assessment categories, management recommendations tied to each category, and a built-in framework for auditing results over time.3PubMed. BI-RADS(®) fifth edition: A summary of changes That template now underpins at least a dozen organ-specific systems.

How the Numbered Categories Work

Most RADS systems use a scale from 0 to 5 or 0 to 6, though the exact number of categories varies. In BI-RADS, for example, category 0 means the images need additional evaluation, category 1 is negative (nothing abnormal), category 2 is benign, category 3 is probably benign with short-interval follow-up recommended, category 4 is suspicious enough to warrant biopsy, and category 5 is highly suggestive of malignancy. The newer Soft Tissue-RADS, still designated as a work in progress, uses seven categories (0 through 6) with diagnostic flowcharts to guide the radiologist through category selection.4PubMed. Soft Tissue-RADS: An ACR Work-in-Progress Framework for Standardized Reporting of Soft-Tissue Lesions on MRI

The key insight is that each number carries an action. A BI-RADS 3 does not just communicate “probably fine”; it communicates “bring this patient back in six months for a follow-up image.” A BI-RADS 4 does not just mean “suspicious”; it means “this needs a biopsy.” That linkage between category and recommended management is what makes RADS systems more than just a vocabulary exercise. They shape clinical decisions in a way that free-text reports never could.

BI-RADS and the Likelihood of Cancer

The categories in BI-RADS are designed so that as the number goes up, the probability of malignancy rises sharply. In one study, the positive predictive value was essentially zero for categories 1 and 2, around 3% for category 3, 23% for category 4, and 92% for category 5.5PubMed. Positive predictive value of the Breast Imaging Reporting and Data System Category 4 covers a wide range of suspicion, though, so the fifth edition subdivided it into 4a, 4b, and 4c. Research on those subcategories shows the spread clearly: about 6% of 4a lesions turned out malignant, 15% of 4b, 53% of 4c, and 91% of category 5.6PubMed. BI-RADS lexicon for US and mammography: interobserver variability and positive predictive value

Those numbers matter for patients because category 4 is where a biopsy is recommended, and category 4a accounts for a lot of biopsies that come back benign. This is the inherent trade-off in any screening system: catch cancers early versus put people through unnecessary procedures. Deep learning models are now being tested to help sort 4a from the higher subcategories, with the goal of reducing biopsies that turn out to be unnecessary while still catching the true cancers.7PubMed Central. Reducing unnecessary biopsies of BI-RADS 4 lesions based on a deep learning model for mammography

Beyond the Breast

The RADS family has grown rapidly. Each system adapts the core template to the anatomy and imaging characteristics of a particular organ.

Lung-RADS was created to standardize reporting for low-dose CT lung cancer screening. Before its introduction, the National Lung Screening Trial had a false-positive rate of about 23%. Applying Lung-RADS criteria retrospectively to that same trial data cut the false-positive rate roughly in half at baseline, down to about 13%, and to around 5% on follow-up screens.8PubMed Central. Performance of Lung-RADS in the National Lung Screening Trial: a retrospective assessment A prospective validation in a diverse urban screening program confirmed the reduction, reporting a false-positive rate of about 10%.9PubMed. Effectiveness of Lung-RADS in Reducing False-Positive Results in a Diverse, Underserved, Urban Lung Cancer Screening Cohort The trade-off is a modest drop in sensitivity, meaning Lung-RADS misses a few more early cancers than the original trial criteria did. The system has already gone through two versions, with challenges remaining around how to handle cavitary nodules, measure growth, and decide when to return patients to routine screening after a follow-up evaluation.

PI-RADS (Prostate Imaging Reporting and Data System) classifies lesions seen on multiparametric MRI of the prostate, scoring them from 1 (very low suspicion) to 5 (very high suspicion of clinically significant cancer). The system influences whether a patient undergoes a prostate biopsy and, if so, whether that biopsy targets a specific MRI-visible lesion. One study of patients with PI-RADS 4 or higher lesions found that about 12% had clinically significant cancer detected only on systematic (random) biopsy cores rather than targeted cores, suggesting that even with high-scoring MRI lesions, targeted biopsy alone can miss some cancers.10The French Journal of Urology. PIRADS ≥ 4 MRI lesion: Is performing systematic biopsies still essential for detecting clinically significant prostate cancer?

LI-RADS (Liver Imaging Reporting and Data System) categorizes observations in patients at risk for hepatocellular carcinoma, running from LR-1 (definitely benign) to LR-5 (definitely HCC). The system uses specific imaging hallmarks like arterial-phase hyperenhancement and washout appearance to place a liver finding on that scale, and includes separate categories for treatment response and tumor in the portal vein.11PubMed Central. Diagnostic Criteria and LI-RADS for Hepatocellular Carcinoma

TI-RADS (Thyroid Imaging Reporting and Data System) uses a points-based approach. Each ultrasound feature of a thyroid nodule, such as its shape, echogenicity, margins, and the presence of calcifications, receives points. The total determines the TI-RADS level and whether a biopsy is recommended. A multi-institutional analysis confirmed that the risk of malignancy rises steadily as TI-RADS points accumulate.12PubMed. Multiinstitutional Analysis of Thyroid Nodule Risk Stratification Using the American College of Radiology Thyroid Imaging Reporting and Data System The ultrasound features that most reliably indicate malignancy are spiculated margins, microcalcifications, a taller-than-wide shape, and marked hypoechogenicity.13PubMed Central. Risk Stratification of Thyroid Nodules: From Ultrasound Features to TIRADS

O-RADS (Ovarian-Adnexal Reporting and Data System) applies the same logic to ovarian and adnexal masses found on ultrasound, scoring them based on morphologic features to estimate malignancy risk and guide management.14PubMed. O-RADS US v2022: An Update from the American College of Radiology’s Ovarian-Adnexal Reporting and Data System US Committee

Do Radiologists Actually Agree on the Scores?

A standardized system only works if different radiologists assign similar scores to the same image. The evidence here is mixed and worth understanding honestly. For sonographic BI-RADS, one study found that agreement between two observers on the final category was only fair, with a kappa of 0.35. Individual descriptors ranged from fair to substantial agreement. The same radiologist reading the same image twice did much better, with substantial to near-perfect self-consistency.15PubMed Central. Interobserver and Intraobserver Agreement of Sonographic BIRADS Lexicon in the Assessment of Breast Masses That gap between inter-observer and intra-observer agreement shows up across RADS systems: radiologists are consistent with themselves, but less so with each other.

PI-RADS shows stronger numbers. One study reported excellent overall agreement between readers, with a kappa of 0.81.16Egyptian Journal of Radiology and Nuclear Medicine. Interobserver agreement of Prostate Imaging–Reporting and Data System (PI-RADS–v2) Updates to the system have also helped. In a comparison of PI-RADS version 2 versus 2.1, agreement for category 4 or higher lesions in the peripheral zone improved from a kappa of 0.51 to 0.64 with the newer version.17PubMed. PI-RADS Versions 2 and 2.1: Interobserver Agreement and Diagnostic Performance in Peripheral and Transition Zone Lesions Among Six Radiologists These improvements through version updates are a recurring theme across RADS systems: each revision tries to tighten language and reduce the ambiguities that let two readers reach different scores.

Competing Guidelines for the Same Organ

TI-RADS is an interesting case because multiple countries and professional societies have developed their own versions. The ACR-TIRADS is widely used in the United States, but there are also Korean (K-TIRADS), European (EU-TIRADS), Chinese (C-TIRADS), and other variants, each weighting ultrasound features somewhat differently. A head-to-head comparison of six TI-RADS guidelines on the same set of over 800 thyroid nodules found that all six provided effective differentiation between malignant and benign nodules, but with meaningful differences in sensitivity and specificity. The ACR-TIRADS had near-perfect sensitivity at about 99.8% but specificity of only 27%, meaning it flagged most benign nodules too. The Chinese TIRADS had the highest specificity at about 93% but much lower sensitivity.18PubMed. A real-world comparison of the diagnostic performances of six different TI-RADS guidelines, including ACR-/Kwak-/K-/EU-/ATA-/C-TIRADS A systematic review and meta-analysis similarly found that K-TIRADS tended to classify nodules at higher risk levels, ACR-TIRADS at moderate risk, and EU-TIRADS at lower risk.19European Thyroid Journal. Head-to-head comparison of American, European, and Asian TIRADSs in thyroid nodule assessment: systematic review and meta-analysis

For patients, this means your thyroid nodule might be recommended for biopsy under one guideline but not another. The choice of which TI-RADS a clinic uses reflects both regional practice norms and a deliberate trade-off between catching every possible cancer versus sparing patients from biopsies of nodules that are almost certainly benign.

How Long It Takes to Learn a RADS System

RADS systems look straightforward on paper, but applying them accurately to real images takes training. A study of trainees learning the VI-RADS system (for bladder MRI) found that agreement with an expert reference improved steeply from about 65% correct in the first batch of cases to 82% after the second batch, eventually reaching 87% by the fifth batch. Significant improvement came after about 100 to 150 cases.20PubMed Central. The learning curve in bladder MRI using VI-RADS assessment score during an interactive dedicated training program A study of 54 trainees learning O-RADS for ovarian ultrasound found that common stumbling blocks included classic benign lesions (which trainees overcategorized as suspicious), assessing blood-flow patterns, and evaluating solid-appearing masses.21Ultrasonography. The learning curve and difficult points of the O-RADS ultrasound risk stratification system in 54 trainees

These learning curves have real implications. A RADS score generated by a radiologist still relatively early in their experience with that system may be less reliable than one from a more seasoned reader. Institutions that adopt a new RADS framework typically invest in dedicated training programs and mentored case review before expecting consistent results.

RADS Scores Combined with Blood Tests

An emerging theme in the literature is pairing RADS scores with laboratory biomarkers to improve accuracy. For ovarian tumors, combining O-RADS ultrasound scores with serum HE4 (a protein marker) yielded an area under the curve of 0.959, with sensitivity of about 95% and specificity of about 93%, outperforming either tool alone.22PubMed. Comparing and combining O-RADS ultrasound, serum CA125, and HE4 for the preoperative diagnosis of ovarian tumors For bladder cancer, combining VI-RADS scores with serum CYFRA 21-1 levels improved prediction of overall survival. Patients who scored high on both measures had one-year survival of only about 43%, compared to significantly better outcomes in patients who were low on either measure.23PubMed Central. Prognostic Utility of Combining VI-RADS Scores and CYFRA 21-1 Levels in Bladder Cancer: A Retrospective Single-Center Study These combinations suggest that RADS scores are increasingly being used not just to decide whether something is cancer but to forecast how aggressively it might behave.

Artificial Intelligence and RADS

AI tools are being built to either assign RADS categories automatically or to assist radiologists in doing so. During the COVID-19 pandemic, a CO-RADS system was developed for chest CT findings, and an automated AI system achieved an area under the curve of 0.95 in discriminating COVID-positive from COVID-negative patients on an internal test set, with moderate-to-substantial agreement with human observers.24PubMed Central. Automated Assessment of COVID-19 Reporting and Data System and Chest CT Severity Scores in Patients Suspected of Having COVID-19 Using Artificial Intelligence For prostate imaging, a deep-learning system showed moderate agreement with an expert radiologist on PI-RADS scoring (kappa of 0.40), but when it came to detecting clinically significant cancer in patients who went on to biopsy, the AI’s performance was not significantly different from the expert’s.25PubMed Central. Deep-Learning-Based Artificial Intelligence for PI-RADS Classification to Assist Multiparametric Prostate MRI Interpretation: A Development Study

AI integration into RADS is promising but carries complications. Newer imaging techniques for prostate MRI, such as multi-shot echo-planar imaging and reduced field-of-view diffusion-weighted imaging, could improve scan quality for both humans and algorithms, but they add technical complexity and are hard to standardize across different hospitals and scanner brands. Folding AI-derived risk scores and clinical data into existing RADS frameworks could sharpen risk stratification, but it also makes the system harder to use and interpret at the bedside.

When Patients See Their RADS Scores

With the rise of patient portals, more people now see their imaging reports, including the RADS category and, increasingly, any AI-generated scores attached to them. Research on how patients react to this information reveals some unintended consequences. In a study of mammography reports, when patients with a BI-RADS 1 (negative) result were also shown an AI abnormality score that happened to be just below the system’s threshold for concern, their desire to follow up with their doctor jumped from about 25% to as high as 80%, even though the AI agreed with the radiologist that the result was normal.26npj Digital Medicine. Accessing AI mammography reports impacts patient follow-up behaviors: the unintended consequences of including AI in patient portals

The same research group found a parallel effect on cancer worry. Concern about breast cancer rose sharply among patients with normal results once any AI score was included in the report, regardless of whether the AI actually flagged the case as abnormal.27medRxiv. Accessing AI mammography reports impacts patient interest in pursuing a medical malpractice claim: The unintended consequences of including AI in patient portals This raises real questions about how much raw scoring information should be visible to patients without context. A number that means “everything is fine” to a radiologist can look alarming to a patient who does not understand the threshold system.

Errors, Liability, and the Limits of Standardization

RADS systems reduce ambiguity, but they do not eliminate diagnostic error. Missed or misclassified findings in breast imaging remain a significant source of malpractice claims. Common causes include perception errors (the radiologist did not see the finding), interpretation errors (they saw it but underestimated its significance), and communication failures (the finding was documented but the referring physician was not adequately alerted). A standardized report does not fix a perception error, and a BI-RADS category only helps if the right category is assigned in the first place.

The broader limitation of any RADS system is that it compresses complex imaging findings into a single number. That number is powerful for communication and decision-making, but it inevitably loses nuance. A BI-RADS 4a lesion with a 6% chance of malignancy and a BI-RADS 4c lesion with a 53% chance of malignancy both get the same top-line recommendation: biopsy. The subcategories help, but they still group together lesions with quite different risk profiles. Each version update tries to sharpen the distinctions, and AI may eventually enable more granular risk estimates, but for now, RADS scores are best understood as useful simplifications rather than precise probabilities.