Reference Hallucination, Citation Reliability, and Readability of Large Language Models in Anatomy-Related Question Answering
Clinical Anatomy, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Basım Tarihi: 2026
- Doi Numarası: 10.1002/ca.70187
- Dergi Adı: Clinical Anatomy
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Aerospace Database, EMBASE, MEDLINE, Natural Science Collection (ProQuest), Biological Science Database (ProQuest), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest), Materials Science & Engineering Collection (ProQuest), Technology Collection (ProQuest)
- Anahtar Kelimeler: anatomy education, citation accuracy, large language models, readability, reference hallucination
- Hacettepe Üniversitesi Adresli: Evet
Özet
Large language models (LLMs) are increasingly used in medical education and academic writing. However, concerns remain regarding reference hallucination, citation, and the reliability of LLM-generated content. This study aimed to evaluate the performance of ChatGPT 5.2, Gemini 3 Pro, and DeepSeek V3.2 in generating anatomy-related responses by assessing bibliographic reference accuracy, citation content consistency, and the readability of LLM-generated content. A total of 120 open-ended anatomy questions covering six anatomical categories (neuroanatomy, musculoskeletal, respiratory and circulatory, gastrointestinal, urogenital and endocrine, head and neck) were submitted to each model. Individual citation components, including author names, article titles, journal names, publication details, and PMIDs, were verified against indexed sources. Citation content consistency was evaluated using a three-point Likert scale. Readability was assessed using the Flesch Reading Ease score, Flesch–Kincaid Grade Level, Coleman–Liau, and Simple Measure of Gobbledygook indices. A total of 1800 references were analyzed. ChatGPT 5.2 demonstrated the lowest hallucination rate (23.2%), whereas Gemini 3 Pro and DeepSeek V3.2 exhibited substantially higher hallucination rates (45.8% and 47.5%, respectively). DeepSeek V3.2 achieved the highest accuracy for several individual bibliographic components, including author names, article titles, volumes, issues, pages, and journal names. PMID accuracy remained limited across all models, ranging from 25.1% to 57.6%. Citation content consistency differed significantly among the models (p < 0.001), with ChatGPT 5.2 demonstrating the highest proportion of fully supported citations (67.2%), compared with Gemini 3 Pro (42.5%) and DeepSeek V3.2 (41.0%). Citation accuracy differed significantly across most anatomical subcategories, with the greatest intermodel discrepancy observed in head and neck anatomy. Readability analyses indicated that the generated responses generally required college-level reading proficiency. Although LLMs can generate plausible anatomy-related responses, substantial limitations remain regarding reference accuracy, hallucination, and citation reliability. Human verification remains essential before incorporating LLM-generated references into academic or educational materials.