Diagnostic Performance of ChatGPT: The Case of Pulmonary Embolism Diagnosis in Burkina Faso

Authors

  • Somé ZM 1. Unité de Formation et de Recherche en Sciences de la Santé/Université Joseph KI ZERBO/ Burkina Faso
  • Yaméogo RA 1. Unité de Formation et de Recherche en Sciences de la Santé/Université Joseph KI ZERBO/ Burkina Faso
  • Somé Z 1. Unité de Formation et de Recherche en Sciences de la Santé/Université Joseph KI ZERBO/ Burkina Faso
  • Kagambéga-Zio L 2. Service de cardiologie du CHU Yalgado OUEDRAOGO/ Burkina Faso
  • Somé N 1. Unité de Formation et de Recherche en Sciences de la Santé/Université Joseph KI ZERBO/ Burkina Faso
  • Bénon-Kaboré L 2. Service de cardiologie du CHU Yalgado OUEDRAOGO/ Burkina Faso
  • Kaboré E 2. Service de cardiologie du CHU Yalgado OUEDRAOGO/ Burkina Faso
  • Ouédraogo S 3. Service de Cardiologie du CHU Régional de Ouahigouya/ Burkina Faso
  • Kologo KJ 1. Unité de Formation et de Recherche en Sciences de la Santé/Université Joseph KI ZERBO/ Burkina Faso
  • Tall-Thiam A 1. Unité de Formation et de Recherche en Sciences de la Santé/Université Joseph KI ZERBO/ Burkina Faso
  • Millogo GRC 1. Unité de Formation et de Recherche en Sciences de la Santé/Université Joseph KI ZERBO/ Burkina Faso
  • Yaméogo NV 1. Unité de Formation et de Recherche en Sciences de la Santé/Université Joseph KI ZERBO/ Burkina Faso
  • Méda N 1. Unité de Formation et de Recherche en Sciences de la Santé/Université Joseph KI ZERBO/ Burkina Faso
  • Zabsonré P 1. Unité de Formation et de Recherche en Sciences de la Santé/Université Joseph KI ZERBO/ Burkina Faso

DOI:

https://doi.org/10.67260/hra.v4i7.7871

Keywords:

ChatGPT, AI, pulmonary embolism, performance, Burkina Faso

Abstract

Introduction. Generative artificial intelligence is a global health issue, but ChatGPT’s ability to diagnose medical conditions remains uncertain. This study aimed to evaluate ChatGPT’s diagnostic performance for pulmonary embolism (PE) in suspected cases in Burkina Faso. Patients and Methods. This was a diagnostic accuracy study comparing ChatGPT’s diagnostic suggestions with the results of chest CT angiography or pulmonary scintigraphy. Results. A total of 140 suspected cases of PE were identified, 70% of which were confirmed. Chest pain (p < 0.001) and its recent onset (p = 0.009) were significantly associated with diagnostic confirmation. ChatGPT performed well on the medical history (specificity 81–86%, PPV 89–93%) and laboratory data (sensitivity 79–89%), but poorly on the physical examination. The paid version outperformed the free version, with good agreement between the two models (Cohen’s k: 0.6–0.8). Conclusion. ChatGPT can assist medical staff in diagnosing PE in patients with acute chest pain, provided there is a subscription, a well-formulated prompt, and access to laboratory and imaging tests. Its integration could reduce the overprescription of CT angiography.

References

1. Akoglu H. User’s guide to correlation coefficients. Turk J Emerg Med. 2018;18(3):91‑3.

2. Hajian-Tilaki K. Sample size estimation in diagnostic test studies of biomedical informatics. J Biomed Inform. 2014;48:193‑204.

3. Hajian-Tilaki KO, Hanley JA, Joseph L, Collet JP. A Comparison of Parametric and Nonparametric Approaches to ROC Analysis of Quantitative Diagnostic Tests. Med Decis Making. 1997;17(1):94‑102.

4. Obuchowski NA. Sample size calculations in studies of test accuracy. Stat Methods Med Res. 1998;7(4):371‑92.

5. Seghda TAA, Dan Naibé T, Dabiré YE, Nacanabo MW, Damoué Seghda S, Dah DC, et al. Performance du score de probabilité à 4 niveaux 4PEPS pour le diagnostic de l’embolie pulmonaire dans une population d’Afrique subsaharienne : données du Registre des Embolies Pulmonaires du Centre Hospitalier Universitaire de Bogodogo, Burkina Faso. Ann Cardiol Angéiologie. 2024;73(5):101798.

6. Introducing ChatGPT [Internet]. 2024 [cité 5 mars 2025]. Disponible sur: https://openai.com/index/chatgpt/

7. Traoré M, Konaté M, Sidibé FM, Koné AC, NDiaye M, Diawara Y, et al. [CT Scan Angiography In The Pulmonary Embolism’s Diagosis At Radiology And Nuclear Medicine Department In Hôpital Du Point « G »]. Mali Med. 2019;34(1):7‑12.

8. Mvondo SM, Koral EB, Kanku JP, Kouassi A, Kouadjo J, Bengono R, et al. Détection, Diagnostic et Recherche de Signes de Gravité de l’Embolie Pulmonaire : une Étude Angioscanographique sur un An. Health Sci Dis. 2020 ; 21(1).

9. Tambe J, Moifo B, Fongang E, Guegang E, Juimo AG. Acute pulmonary embolism in the era of multi-detector CT: a reality in sub-Saharan Africa. BMC Med Imaging. 2012;12:31.

10. Kologo KJ, Millogo GRC, Kambiré Y, Adoko H, Kagambéga LJ, Thiam/Tall A, et al. Insuffisance cardiaque et anémie dans le service de cardiologie du Centre Hospitalier Universitaire Yalgado OUEDRAOGO: Aspects épidemiologiques, thérapeutiques et pronostiques: Heart failure and anemia in the cardiology department of the Yalgado Ouedraogo university teaching Hospital: epidemiology, management and prognosis. Health Sci Dis. 2022; 23(2).

11. Konstantinides SV, Meyer G, Becattini C, Bueno H, Geersing GJ, Harjola VP, et al. 2019 ESC Guidelines for the diagnosis and management of acute pulmonary embolism developed in collaboration with the European Respiratory Society (ERS): The Task Force for the diagnosis and management of acute pulmonary embolism of the European Society of Cardiology (ESC). Eur Heart J. 2020;41(4):543‑603.

12. Righini M, Aujesky D, Roy PM, Cornuz J, De Moerloose P, Bounameaux H, et al. Clinical Usefulness of D-Dimer Depending on Clinical Probability and Cutoff Value in Outpatients With Suspected Pulmonary Embolism. Arch Intern Med. 2004;164(22):2483.

13. Thiam A, Kinda G, Tindano C, Adoko H, Bouda C, Millogo GRC and al. Venous thromboembolic disease in Burkina Faso : results of the prospective registry REMAVET (registry of Venous Thromboembolic). 2017 ; 6(1) : 1.

14. Le Gal G, Righini M, Roy PM, Sanchez O, Aujesky D, Bounameaux H, et al. Prediction of Pulmonary Embolism in the Emergency Department: The Revised Geneva Score. Ann Intern Med. 2006;144(3):165.

15. Chang MC. Use of artificial intelligence in the field of pain medicine. World J Clin Cases. 2024;12(2):236‑9.

16. Penaloza A, Verschuren F, Meyer G, Quentin-Georget S, Soulie C, Thys F, et al. Comparison of the Unstructured Clinician Gestalt, the Wells Score, and the Revised Geneva Score to Estimate Pretest Probability for Suspected Pulmonary Embolism. Ann Emerg Med. 2013;62(2):117-124.e2.

17. Wells P, Anderson D, Rodger M, Ginsberg J, Kearon C, Gent M, et al. Derivation of a Simple Clinical Model to Categorize Patients Probability of Pulmonary Embolism: Increasing the Models Utility with the SimpliRED D-dimer. Thromb Haemost. 2000;83(03):416‑20.

18. Le Gal G, Righini M, Roy PM, Sanchez O, Aujesky D, Bounameaux H, et al. Prediction of pulmonary embolism in the emergency department: the revised Geneva score. Ann Intern Med. 2006;144(3):165‑71.

19. Pernod G, Caterino J, Maignan M, Tissier C, Kassis J, Lazarchick J, et al. D-Dimer Use and Pulmonary Embolism Diagnosis in Emergency Units: Why Is There Such a Difference in Pulmonary Embolism Prevalence between the United States of America and Countries Outside USA? Fukumoto Y, éditeur. PLOS ONE. 2017;12(1):e0169268.

20. Shukla R, Mishra AK, Banerjee N, Verma A. The Comparison of ChatGPT 3.5, Microsoft Bing, and Google Gemini for Diagnosing Cases of Neuro-Ophthalmology. Cureus. 2024;16(4):e58232.

21. Gosak L, Pruinelli L, Topaz M, Štiglic G. The ChatGPT effect and transforming nursing education with generative AI: Discussion paper. Nurse Educ Pract. 2024;75:103888.

22. Zhu L, Mou W, Lai Y, Chen J, Lin S, Xu L, et al. Step into the era of large multimodal models: a pilot study on ChatGPT-4V (ision)’s ability to interpret radiological images. Int J Surg Lond Engl. 2024;110(7):4096‑102.

23. Lecler A, Duron L, Soyer P. Revolutionizing radiology with GPT-based models: Current applications, future possibilities and limitations of ChatGPT. Diagn Interv Imaging. 2023;104(6):269‑74.

24. Harskamp RE, De Clercq L. Performance of ChatGPT as an AI-assisted decision support tool in medicine: a proof-of-concept study for interpreting symptoms and management of common cardiac conditions (AMS℡HEART-2). Acta Cardiol. 2024;79(3):358‑66.

25. Harada Y, Suzuki T, Harada T, Sakamoto T, Ishizuka K, Miyagami T, et al. Performance evaluation of ChatGPT in detecting diagnostic errors and their contributing factors: an analysis of 545 case reports of diagnostic errors. BMJ Open Qual. 2024;13(2):e002654.

26. Laohawetwanit T, Namboonlue C, Apornvirat S. Accuracy of GPT-4 in histopathological image detection and classification of colorectal adenomas. J Clin Pathol. 2024;jcp-2023-209304.

27. Hershenhouse JS, Mokhtar D, Eppler MB, Rodler S, Storino Ramacciotti L, Ganjavi C, et al. Accuracy, readability, and understandability of large language models for prostate cancer information to the public. Prostate Cancer Prostatic Dis. 2024;

28. Pan Y, Jiao FY. Application of artificial intelligence in the diagnosis and treatment of Kawasaki disease. World J Clin Cases. 2024;12(23):5304‑7.

29. Vaira LA, Lechien JR, Abbate V, Allevi F, Audino G, Beltramini GA, et al. Accuracy of ChatGPT-Generated Information on Head and Neck and Oromaxillofacial Surgery: A Multicenter Collaborative Analysis. Otolaryngol--Head Neck Surg Off J Am Acad Otolaryngol-Head Neck Surg. 2024;170(6):1492‑503.

30. Dabbas WF, Odeibat YM, Alhazaimeh M, Hiasat MY, Alomari AA, Marji A, et al. Accuracy of ChatGPT in Neurolocalization. Cureus. 2024;16(4):e59143.

31. Suárez A, Díaz-Flores García V, Algar J, Gómez Sánchez M, Llorente de Pedro M, Freire Y. Unveiling the ChatGPT phenomenon: Evaluating the consistency and accuracy of endodontic question answers. Int Endod J. 2024;57(1):108‑13.

32. Liu X, Wu J, Shao A, Shen W, Ye P, Wang Y, et al. Uncovering Language Disparity of ChatGPT on Retinal Vascular Disease Classification: Cross-Sectional Study. J Med Internet Res. 2024;26:e51926.

33. Ming S, Yao X, Guo X, Guo Q, Xie K, Chen D, et al. Performance of ChatGPT in Ophthalmic Registration and Clinical Diagnosis: Cross-Sectional Study. J Med Internet Res. 2024;26:e60226.

34. Delsoz M, Madadi Y, Raja H, Munir WM, Tamm B, Mehravaran S, et al. Performance of ChatGPT in diagnosis of corneal eye diseases. Cornea. 2024;43(5):664‑70.

35. Lechien JR, Naunheim MR, Maniaci A, Radulesco T, Saibene AM, Chiesa-Estomba CM, et al. Performance and Consistency of ChatGPT-4 Versus Otolaryngologists: A Clinical Case Series. Otolaryngol--Head Neck Surg Off J Am Acad Otolaryngol-Head Neck Surg. 2024;170(6):1519‑26.

36. Stoneham S, Livesey A, Cooper H, Mitchell C. ChatGPT versus clinician: challenging the diagnostic capabilities of artificial intelligence in dermatology. Clin Exp Dermatol. 2024;49(7):707‑10.

37. Albaladejo A, Lorleac’h A, Allain JS. [The spring of artificial intelligence: AI vs. expert for internal medicine cases]. Rev Med Interne. 2024;45(7):409‑14.

38. Johnson D, Goodman R, Patrinely J, Stone C, Zimmerman E, Donald R, et al. Assessing the Accuracy and Reliability of AI-Generated Medical Responses: An Evaluation of the Chat-GPT Model [Internet]. Research Square; 2023 [cité 5 mars 2025]. Disponible sur: https://www.researchsquare.com/article/rs-2566942/v1

39. Kachman MM, Brennan I, Oskvarek JJ, Waseem T, Pines JM. How artificial intelligence could transform emergency care. Am J Emerg Med. 2024;81:40‑6.

Published

07/01/2026

How to Cite

Somé ZM, Yaméogo RA, Somé Z, Kagambéga-Zio L, Somé N, Bénon-Kaboré L, … Zabsonré P. (2026). Diagnostic Performance of ChatGPT: The Case of Pulmonary Embolism Diagnosis in Burkina Faso . HEALTH RESEARCH IN AFRICA, 4(7), 5–16. https://doi.org/10.67260/hra.v4i7.7871

Issue

Section

Research Articles

Most read articles by the same author(s)

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.