AI Generates and Self-Checks Specialist Study Materials—Yet 1% Still Contain Critical Errors

| Input:

Twenty radiologists re-verified AI-cleared content; critical errors missed by the automated check exceeded the research team’s pre-set safety threshold of 0.3 percent

As generative AI expands in medical education, a new study underscores the need for final expert oversight after finding critical errors in content that cleared automated verification. Photo=Getty Images Bank
As generative AI expands in medical education, a new study underscores the need for final expert oversight after finding critical errors in content that cleared automated verification. Photo=Getty Images Bank

Artificial intelligence can now generate study materials for medical specialist board exams and run automated checks for errors. But if materials clear that automated process, can residents preparing for board exams trust them as-is?

For radiology board-exam preparation, the answer is closer to “not yet.” About 1 percent of the study materials that passed AI’s automated checks still contained critical errors that could distort a physician’s clinical judgment.

Asan Medical Center in Seoul announced on August 24 that a research team had developed a system that mass-produces radiology education materials using generative AI and automatically filters out errors, followed by a verification study conducted by practicing physicians.

The study was led by Prof. Kim Nam-guk of the Department of Convergence Medicine at Asan Medical Center, researcher Nam Yoo-jin, and Prof. Hong Pa of the Department of Radiology at Samsung Changwon Hospital. The findings were published in the international journal npj Digital Medicine.

6,000 Flashcards Created from the Specialist Curriculum

Based on the Korean Society of Radiology’s specialist training curriculum, the researchers instructed AI to select core content needed for board exams and produce study tools.

The system first created flashcards consisting of question-and-answer pairs designed for repeated review of key concepts. It then refined the text and created accompanying infographics. AI re-checked the quality of both text and images, and if an output failed to meet predefined criteria, it was sent back to the previous step to be regenerated.

Through this process, the team produced 6,000 flashcards and 833 infographics across 11 radiology subspecialties.

The researchers were not merely testing how much material AI could generate. The crucial question was how far medical education could rely on materials that AI had checked independently and deemed "problem-free."

AI Gave It a "Pass," But Doctors Found Critical Errors in 1 Percent

Of the generated flashcards, 1,284 underwent human verification. Twenty physicians—nine radiology residents and 11 board-certified radiologists—reviewed whether the facts were accurate, whether the materials were appropriate for specialist education, and whether they contained blocking errors that required correction.

In this study, a "blocking error" did not refer to a simple typo or awkward design, but to a factual error that, if learned as-is, could steer a physician’s clinical judgment in the wrong direction. Examples included mislabeling anatomical locations, linking imaging findings to the wrong disease, or presenting incorrect differential diagnoses. In one case among materials approved by the AI, the content failed to properly identify a venous thrombus appearing alongside a renal mass.

When the team analyzed 1,100 expert evaluations across 980 flashcards that had passed AI’s automated checks, they found 11 instances of such critical errors—representing a rate of about one per 100 evaluations, or 1 percent. In short, while AI concluded there were "no major problems," human physicians found remaining errors that made the materials unsafe for direct educational use.

Before starting the study, the researchers set a target of keeping critical errors missed by AI below 0.3 percent—meaning fewer than three missed errors per 1,000 evaluations. In statistics, missing a real problem by judging it as problem-free is called a "false negative."

The authors noted that 0.3 percent is not an official safety standard set by governments or academic societies, as no accredited error-tolerance standard yet exists for AI in medical education. Instead, the researchers established this study benchmark by referencing patient-safety thresholds in other clinical areas, such as medication-dispensing errors or missed diagnostic tests. Therefore, the 1 percent error rate should not be broadly generalized as a universal error rate for all medical AI materials, but rather viewed as the outcome of evaluating this specific radiology system under the study's methodology.

A chart details the process of generating radiology education content using artificial intelligence. Graphic=Asan Medical Center
A chart details the process of generating radiology education content using artificial intelligence. Graphic=Asan Medical Center

Beyond AI "Hallucinations": Why Plausible Errors Slipped Through

Why did AI miss these errors during its automated re-check?

Generative AI is known for "hallucinations"—generating plausible-sounding statements that lack factual basis. However, the researchers noted that the remaining errors in this study could not be attributed solely to AI hallucinations.

To maximize automated accuracy, the team repeatedly refined prompt engineering and enabled the AI to search and reference designated external evidence sources during verification. Even so, out of 39 total critical errors identified across all human evaluations, the AI caught 28, while 11 slipped through undetected.

A key factor was the sequential AI-on-AI verification structure. One AI model's output was passed directly to the next model in a linear workflow, rather than being subjected to a cross-examination system where models challenged earlier judgments. As a result, if a plausible wrong answer was generated initially, subsequent AI stages accepted it without suspicion.

Furthermore, the AI exhibited high "confidence." For the 11 missed errors found by human experts, the checking AI assigned top scores for both accuracy and educational quality. These were not borderline items that barely passed; they were materials rated as flawless by the system.

The visual nature of radiology also contributed. Approximately three-quarters of the 11 missed errors were image-related. Even if written descriptions are correct, radiology materials become invalid if an arrow points to the wrong anatomical structure or if spatial relationships between organs and blood vessels are rendered incorrectly. While generative AI can produce convincing medical images at a glance, it has not yet achieved consistent precision in fine anatomical details.

Ultimately, these errors resulted from a combination of factors: AI hallucinations, subsequent AI models trusting plausible false answers, a sequential verification structure that prevented re-evaluating earlier steps, and hardware/software limitations in depicting detailed medical imaging. Wrong answers that appeared plausible to both human eyes and AI checks were the most likely to survive the entire pipeline.

From left: Prof. Kim Nam-guk of the Department of Convergence Medicine at Asan Medical Center, researcher Nam Yoo-jin, and Prof. Hong Pa of the Department of Radiology at Samsung Changwon Hospital. Photo=Asan Medical Center
From left: Prof. Kim Nam-guk of the Department of Convergence Medicine at Asan Medical Center, researcher Nam Yoo-jin, and Prof. Hong Pa of the Department of Radiology at Samsung Changwon Hospital. Photo=Asan Medical Center

AI as a "Second Set of Eyes"—With Final Verification Left to Humans

The study also revealed an interesting dynamic regarding human review. Clinicians first evaluated the flashcards without seeing the AI’s assessment, and later re-examined the same materials after reviewing the issues flagged by the AI.

Upon seeing the AI’s feedback, the number of evaluations identifying critical errors rose from 39 to 54, and scores for accuracy and educational quality became noticeably stricter. This suggests AI can serve effectively as a "second set of eyes," prompting human experts to catch details they might otherwise overlook.

However, the authors noted that this increase cannot definitively prove that AI directly improved human error-detection. Reviewers may have caught additional issues simply by inspecting the materials a second time, or knowing that AI had flagged potential problems may have made them more critical.

Even if expert review is necessary, inspecting every piece of generated content remains a bottleneck. In the study, the median time a physician spent reviewing a single flashcard was 53 seconds. At that pace, reviewing all 6,000 flashcards would require approximately 88 hours—roughly 11 full workdays at eight hours a day.

Automated processing also encountered technical limits: for a portion of the generated flashcards, processing errors occurred, leaving no automated quality-check results at all.

While AI excels at producing vast quantities of educational content rapidly, delegating the entire pipeline to automated systems allows critical errors to persist. Conversely, having human experts create and verify every item manually negates the efficiency gains of AI.

The researchers propose a "human-in-the-loop" framework to bridge this gap. Under this approach, AI handles initial content generation and preliminary automated screening, while domain experts perform targeted final verification before materials are deployed in clinical education.

"Through this large-scale study, we demonstrated that expert verification is essential to safely deploy generative AI in real educational settings," said Prof. Kim Nam-guk. "Further research is needed to systematically manage errors missed by automated screening and establish clear safety standards."

Researcher Nam Yoo-jin added, "While we confirmed the potential of generative AI to mass-produce medical imaging study materials, our findings show that rather than using generated content as-is, it is crucial to analyze at what stage and what types of errors occur, and to filter them out thoroughly."

×