Medical AI Can Explain Its Decisions—but Can Clinicians Trust Its Reasons?

| Input:

A Korea University team analyzed 40,000 brain MRIs and found that "robustness" most closely matches radiologists' clinical judgment

A doctor reviews a skull X-ray. For medical AI rationales to be trusted, systems must consistently highlight the same target areas even when training conditions change. The photo is for illustrative purposes. Photo=Getty Image Bank
A doctor reviews a skull X-ray. For medical AI rationales to be trusted, systems must consistently highlight the same target areas even when training conditions change. The photo is for illustrative purposes. Photo=Getty Image Bank

Artificial intelligence in medical imaging can analyze brain magnetic resonance imaging (MRI) scans not only to estimate disease likelihood, but also to visually highlight—in color—the specific brain regions that influenced its decision. By placing this visual "evidence" alongside its diagnostic conclusion, the system appears more transparent and trustworthy.

However, clinicians face a crucial underlying question: Is the AI’s highlighted evidence stable? If retraining the same AI architecture causes it to point to entirely different brain regions, the initial visual explanation may simply be a fluke of that specific training run.

A research team led by Professor Christian Wallraven of Korea University’s Department of Artificial Intelligence, in collaboration with a radiology team led by Professors Yoo Sung-hye, Kim Bo-gyu, and Park A-rim at Korea University Anam Hospital, recently evaluated multiple metrics used to assess medical AI explanations. The team announced on the 24th that "robustness"—a metric measuring how consistently an explanation holds up across repeated training runs—most reliably aligns with clinical evidence.

A Plausible Explanation Is Not the Same as a Reliable One

Deep-learning AI generates outcomes through complex computational pathways, making it nearly impossible for humans to trace every step in practice. Consequently, even when AI flags an abnormal finding, clinicians often cannot verify precisely why it reached that conclusion—a challenge widely known as the "black box" problem.

Explainable artificial intelligence (XAI) was developed to address this issue by mapping the brain regions the AI prioritized, allowing clinicians to visually inspect the rationale behind a diagnosis.

The primary challenge, however, is the lack of a standardized benchmark to determine which explanations are genuinely trustworthy. While researchers have proposed metrics to measure "fidelity" (how accurately an explanation reflects the AI’s internal decision process) and "complexity" (how simple and readable the explanation is), whether these metrics correlate with clinically meaningful regions has remained unvalidated.

(From left) Cho Dae-hyun, an integrated master’s-doctoral student in Korea University’s Department of Artificial Intelligence (first author); Prof. Yoo Sung-hye of Radiology at Korea University Anam Hospital (co-second author); Prof. Kim Bo-gyu of Radiology at Korea University Anam Hospital (co-second author); and Prof. Christian Wallraven of Korea University’s Department of Artificial Intelligence (corresponding author). Photo=Korea University Anam Hospital
(From left) Cho Dae-hyun, an integrated master’s-doctoral student in Korea University’s Department of Artificial Intelligence (first author); Prof. Yoo Sung-hye of Radiology at Korea University Anam Hospital (co-second author); Prof. Kim Bo-gyu of Radiology at Korea University Anam Hospital (co-second author); and Prof. Christian Wallraven of Korea University’s Department of Artificial Intelligence (corresponding author). Photo=Korea University Anam Hospital

Why Robustness Aligns Best with Clinical Evidence

The robustness metric evaluated by the research team measures how consistently an AI model flags the same "important" regions when its underlying architecture is retrained multiple times from different initializations. A high robustness score indicates that the AI's visual explanation is a stable, reproducible finding rather than a random outcome of a single training run.

To test this, the researchers analyzed approximately 40,000 brain MRI scans sourced from the UK Biobank and the Alzheimer’s Disease Neuroimaging Initiative (ADNI). They executed over 10,000 simulations combining nine distinct deep-learning architectures with eight XAI explanation methods.

They then evaluated whether each metric was clinically meaningful by comparing them against clinical benchmarks. After dividing the MRI scans into 3D grid units called "voxels" to analyze subtle variations in tissue morphology, they compared the metrics against region-by-region disease volume associations and the expert judgments of practicing radiologists.

The analysis revealed that conventional metrics like fidelity and complexity showed low correlation with clinical evidence. Furthermore, fidelity and complexity scores varied wildly for the exact same explanation method when applied across different AI architectures.

In contrast, robustness demonstrated consistent agreement with radiologists' clinical judgments and morphological brain analyses. Crucially, robustness scores remained stable even when changing the underlying AI architecture, while also offering favorable computational efficiency.

"This study is significant because it systematically connects evaluation criteria for medical AI explanations with concrete clinical evidence," said Professor Wallraven.

Professor Yoo Sung-hye added, "For medical AI to be deployed in real-world clinical settings, it must demonstrate not only high diagnostic accuracy, but also that its underlying rationale is clinically valid and reproducible."

The authors noted that the study focused specifically on structural brain MRIs across two analytical tasks. Further research will be needed to confirm whether robustness demonstrates similar validity in other imaging modalities, such as computed tomography (CT), X-rays, and digital pathology.

Under the European Union’s AI Act, medical device AI is classified as a high-risk technology, requiring strict standards for transparency, accuracy, and robustness. The team's findings underscore that evaluating medical AI requires looking beyond diagnostic accuracy to ensure its explanations remain stable and consistent.

The study was published online in the journal Medical Image Analysis on June 16.

×