
A joint research team from Samsung Medical Center and AI startup ZeroOneAI has developed a domain-specialized medical artificial intelligence model, "ZEO Med 2," which outperformed OpenAI’s GPT-5.2 and Google’s Gemini on a benchmark based on South Korea’s national physician licensing examination.
Built by a team led by Professor Cha Won-cheol and Professor Son Myung-hee of Samsung Medical Center’s Department of Emergency Medicine in collaboration with ZeroOneAI, the project was supported by an advanced GPU infrastructure initiative under the Ministry of Science and ICT. Samsung Medical Center announced the benchmark results on July 29.
Strong Performance on Korean Physician Exam, Mixed Results Elsewhere
On the "KorMedMCQA" benchmark—a dataset derived from national licensing examinations for physicians, nurses, pharmacists, and dentists in South Korea—ZEO Med 2 answered 425 out of 435 physician-level questions correctly, achieving an accuracy rate of 97.70%. Under identical evaluation conditions conducted by ZeroOneAI, GPT-5.2 answered 424 questions correctly (97.47%), while Google Gemini 3.1 Flash-Lite scored 96.09%.
ZEO Med 2 also led among open-weight models, whose weights are publicly released for external development. It outperformed its own base model, Google Gemma 4 31B (97.01%), as well as Qwen3.6-27B (93.56%).
However, the model did not surpass rival systems across all medical evaluations. On the MedQA-USMLE benchmark, which draws from the United States Medical Licensing Examination, GPT-5.2 scored 95.99%, higher than ZEO Med 2’s 94.82%. Additionally, Gemini 3.1 Flash-Lite scored higher than ZEO Med 2 on Korean exam sets for nurses, pharmacists, and dentists, as well as on the combined average across the four professions.
The evaluation utilized a zero-shot approach, where no sample questions or answers were provided in advance. Each test question was presented to the model five times, with the most frequently selected response counted as the final answer.
To demonstrate the clinical reasoning required by the benchmark, Samsung Medical Center highlighted an item from the 90th National Medical Licensing Examination in 2026. The clinical scenario involved a 20-month-old female patient brought to the emergency department with a high fever who experienced a seizure, failed to regain consciousness, and suffered a recurrent seizure. ZEO Med 2 correctly selected lorazepam, a benzodiazepine anticonvulsant, as the appropriate medication.
Despite high scores on standardized tests, researchers noted that exam performance alone does not guarantee real-world clinical competence. In actual medical practice, clinicians must synthesize complex, multi-modal data—including medical history, lab results, diagnostic imaging, and active prescriptions—where errors carry critical safety risks.

Expanding Toward Voice, Imaging, and Clinical Utility
The model details, evaluation code, and weights for ZEO Med 2 have been published on the open AI platform Hugging Face. Fine-tuned on domain-specific medical data from Google’s Gemma 4 31B, the open-weight release allows external researchers and developers to utilize the system under specified conditions.
Moving forward, Samsung Medical Center plans to expand ZEO Med 2 beyond text processing into a multimodal model capable of analyzing audio and diagnostic imagery. The team is also establishing clinical safety mechanisms to oversee inputs, outputs, and automated actions, alongside model compression efforts to enable deployment on lighter computing hardware.
"The critical goal now is not simply incremental gains on test questions, but building a reliable model that spans voice and imaging, connects seamlessly to real-world services, and remains lightweight enough to deploy anywhere," said Professor Cha Won-cheol. "We will continue refining it into a clinically specialized AI that makes a tangible contribution in medical settings."
