Artificial intelligence-based patient information in rotator cuff injuries: A cross-sectional comparative study of ChatGPT and DeepSeek models


Bozgeyik B., Öğümsöğütlü E., Huri G.

Medicine (United States), cilt.105, sa.27, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 105 Sayı: 27
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1097/md.0000000000049658
  • Dergi Adı: Medicine (United States)
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, BIOSIS, CINAHL, EMBASE, MEDLINE, Directory of Open Access Journals
  • Anahtar Kelimeler: artificial intelligence, ChatGPT, DeepSeek, rotator cuff injury
  • Hacettepe Üniversitesi Adresli: Evet

Özet

This study aimed to comparatively evaluate the performance of the chat generative pretrained transformer (ChatGPT) and DeepSeek artificial intelligence (AI) models in patient information about rotator cuff injuries. This cross-sectional comparative study was conducted in May 28, 2025 using ChatGPT-4o (OpenAI) and DeepSeek V3 (DeepSeek Inc.) models. Sixteen frequently asked questions related to rotator cuff injuries were posed to both the AI models. The responses were then independently assessed by 2 experienced orthopedic surgeons using the Journal of the American Medical Association (JAMA), response rating system, DISCERN, and 4-point Likert scales. In addition, the readability of the responses was analyzed using the Flesch-Kincaid Readability Score (FRES) and Flesch-Kincaid Grade Level. The primary outcome was overall information quality, secondary outcomes included JAMA benchmark adherence and readability metrics. None of the models met JAMA criteria. In terms of response rating system, there was no statistically significant difference between the 2 models (P >.05). DeepSeek demonstrated higher DISCERN scores compared to ChatGPT (50.12 vs 47.03), with a mean difference of 3.09 (95% CI: 1.58 to 4.61; P = .001). While there was no significant difference in accuracy, clarity, and consistency criteria between the 2 models in the 4-point Likert evaluation (P >.05), DeepSeek scored significantly higher than ChatGPT in the completeness criterion, with a mean difference of 0.75 (95% CI: 0.46 to 1.04; P = .001). In terms of readability, both models showed similar performance (FRES, P >.05; Flesch-Kincaid Grade Level, P >.05). Both AI models deliver satisfactory and clinically relevant information for rotator cuff injury patient education. Although DeepSeek was superior to ChatGPT in terms of completeness of patient information regarding rotator cuff injuries, the results were similar for the other criteria. The responses from both the AI tools were considered promising. However, they require improvements in terms of adherence to scientific standards, transparency, citations, and readability.