Archive \ Volume.17 2026 Issue 3

Where Does the Answer Come From? A Scoping Review of Evidence Tracing, Hallucination, and Safeguards in Generative Medicines Information

, , ,
  1. Department of Generative AI and Medicines Information, Faculty of Pharmacy, Hanoi University of Pharmacy, Hanoi, Vietnam.
  2. Department of Evidence Tracing and Hallucination Detection, Faculty of Pharmacy, Can Tho University of Medicine and Pharmacy, Can Tho, Vietnam.
  3. Department of AI Safeguards in Pharmacy, Faculty of Pharmacy, Hue University, Hue, Vietnam.

Abstract

Generative artificial intelligence can produce medicines information rapidly, yet fluent answers can conceal uncertainty about where claims originated, whether cited sources exist, and whether retrieved evidence supports the wording presented. This scoping review mapped how generative medicines information systems attach evidence to claims, define and measure hallucination and citation error, implement retrieval and provenance, evaluate claim–source alignment, and apply human verification, conflict handling, and escalation safeguards. A protocol-led scoping review used the Population–Concept–Context framework and followed JBI and PRISMA-ScR principles. Peer-reviewed studies of generative systems producing medication or medicines information were eligible when they reported extractable evidence-tracing, citation, hallucination, retrieval, or safeguard methods. Records were screened, charted, appraised, and synthesized through evidence-tracing and safeguard taxonomies. The evidence base was methodologically heterogeneous and concentrated in retrospective, cross-sectional, or simulated evaluations. Systems attached evidence through model-generated references, search-linked citations, constrained document retrieval, or retrieval-augmented generation, but citation presence did not reliably establish bibliographic validity, relevance, or claim-level support. Hallucination definitions varied across fabricated references, incorrect citation elements, unsupported factual statements, and source–claim mismatch. Retrieval reduced some unsupported generation yet introduced failures involving document selection, temporal validity, partial entailment, and evidence conflict. Human review was frequently recommended, but escalation thresholds, reviewer workload, audit trails, and prospective workflow performance were rarely evaluated. Evidence tracing in generative medicines information remains fragmented. Trustworthy evaluation requires separate assessment of source existence, source quality, retrieval completeness, claim-level support, contradiction, and human adjudication. Current safeguards alone do not establish clinical safety or deployment readiness.


Downloads: 20
Views: 64

How to cite:
Vancouver
Huy NT, Minh PQ, Bich LT, Nam TV. Where Does the Answer Come From? A Scoping Review of Evidence Tracing, Hallucination, and Safeguards in Generative Medicines Information. Arch Pharm Pract. 2026;17(3):84-93. https://doi.org/10.51847/MDBmOwWqrB
APA
Huy, N. T., Minh, P. Q., Bich, L. T., & Nam, T. V. (2026). Where Does the Answer Come From? A Scoping Review of Evidence Tracing, Hallucination, and Safeguards in Generative Medicines Information. Archives of Pharmacy Practice, 17(3), 84-93. https://doi.org/10.51847/MDBmOwWqrB

Download Citation
References
  1. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med. 2023;29(8):1930-40. doi:10.1038/s41591-023-02448-8
  2. Moor M, Banerjee O, Abad ZSH, Krumholz HM, Leskovec J, Topol EJ, et al. Foundation models for generalist medical artificial intelligence. Nature. 2023;616(7956):259-65. doi:10.1038/s41586-023-05881-4
  3. Lee P, Bubeck S, Petro J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N Engl J Med. 2023;388(13):1233-9. doi:10.1056/NEJMsr2214184
  4. Howell MD. Generative artificial intelligence, patient safety and healthcare quality: A review. BMJ Qual Saf. 2024;33(11):748-54. doi:10.1136/bmjqs-2023-016690
  5. Tricco AC, Lillie E, Zarin W, O’Brien KK, Colquhoun H, Levac D, et al. PRISMA extension for scoping reviews (PRISMA-ScR): Checklist and explanation. Ann Intern Med. 2018;169(7):467-73. doi:10.7326/M18-0850
  6. Peters MDJ, Marnie C, Tricco AC, Pollock D, Munn Z, Alexander L, et al. Updated methodological guidance for the conduct of scoping reviews. JBI Evid Synth. 2020;18(10):2119-26. doi:10.11124/JBIES-20-00167
  7. Munn Z, Peters MDJ, Stern C, Tufanaru C, McArthur A, Aromataris E. Systematic review or scoping review? Guidance for authors when choosing between a systematic or scoping review approach. BMC Med Res Methodol. 2018;18(1):143. doi:10.1186/s12874-018-0611-x
  8. Ong JCL, Chen MH, Ng N, Elangovan K, Tan NYT, Jin L, et al. A scoping review on generative AI and large language models in mitigating medication related harm. NPJ Digit Med. 2025;8(1):182. doi:10.1038/s41746-025-01565-7
  9. Peters MDJ, Godfrey C, McInerney P, Khalil H, Larsen P, Marnie C, et al. Best practice guidance and reporting items for the development of scoping review protocols. JBI Evid Synth. 2022;20(4):953-68. doi:10.11124/JBIES-21-00242
  10. Pollock D, Davies EL, Peters MDJ, Tricco AC, Alexander L, McInerney P, et al. Undertaking a scoping review: A practical guide for nursing and midwifery students, clinicians, researchers, and academics. J Adv Nurs. 2021;77(4):2102-13. doi:10.1111/jan.14743
  11. Bramer WM, de Jonge GB, Rethlefsen ML, Mast F, Kleijnen J. A systematic approach to searching: An efficient and complete method to develop literature searches. J Med Libr Assoc. 2018;106(4):531-41. doi:10.5195/jmla.2018.283
  12. Rethlefsen ML, Kirtley S, Waffenschmidt S, Ayala AP, Moher D, Page MJ, et al. PRISMA-S: An extension to the PRISMA Statement for reporting literature searches in systematic reviews. Syst Rev. 2021;10(1):39. doi:10.1186/s13643-020-01542-z
  13. Pollock D, Peters MDJ, Khalil H, McInerney P, Alexander L, Tricco AC, et al. Recommendations for the extraction, analysis, and presentation of results in scoping reviews. JBI Evid Synth. 2023;21(3):520-32. doi:10.11124/JBIES-22-00123
  14. Aljamaan F, Temsah MH, Altamimi I, Al-Eyadhy A, Jamal A, Alhasan K, et al. Reference Hallucination Score for medical artificial intelligence chatbots: Development and usability study. JMIR Med Inform. 2024;12. doi:10.2196/54345
  15. Walters WH, Wilder EI. Fabrication and errors in the bibliographic citations generated by ChatGPT. Sci Rep. 2023;13(1):14045. doi:10.1038/s41598-023-41032-5
  16. Sebo P. How accurate are the references generated by ChatGPT in internal medicine? Intern Emerg Med. 2024;19(1):247-9. doi:10.1007/s11739-023-03484-5
  17. Morath B, Chiriac U, Jaszkowski E, Deiß C, Nürnberg H, Hörth K, et al. Performance and risks of ChatGPT used in drug information: An exploratory real-world analysis. Eur J Hosp Pharm. 2024;31(6):491-7. doi:10.1136/ejhpharm-2023-003750
  18. Huang X, Estau D, Liu X, Yu Y, Qin J, Li Z. Evaluating the performance of ChatGPT in clinical pharmacy: A comparative study of ChatGPT and clinical pharmacists. Br J Clin Pharmacol. 2024;90(1):232-8. doi:10.1111/bcp.15896
  19. Roosan D, Padua P, Khan R, Khan H, Verzosa C, Wu Y. Effectiveness of ChatGPT in clinical pharmacy and the role of artificial intelligence in medication therapy management. J Am Pharm Assoc (2003). 2024;64(2):422-8.e8. doi:10.1016/j.japh.2023.11.023
  20. Chelli M, Descamps J, Lavoué V, Trojani C, Azar M, Deckert M, et al. Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: Comparative analysis. J Med Internet Res. 2024;26. doi:10.2196/53164
  21. Taş N, Erden Y, Temel MH, Bağcıer F. A reproducible statistical evaluation framework for large-sample assessment of AI-generated medical references: Cross-platform application of the Reference Hallucination Score. BMC Med Res Methodol. 2026. Online ahead of print. doi:10.1186/s12874-026-02906-0
  22. Wu K, Wu E, Wei K, Zhang A, Casasola A, Nguyen T, et al. An automated framework for assessing how well large language models cite relevant medical references. Nat Commun. 2025;16(1):3615. doi:10.1038/s41467-025-58551-6
  23. Amugongo LM, Mascheroni P, Brooks S, Doering S, Seidel J. Retrieval augmented generation for large language models in healthcare: A systematic review. PLOS Digit Health. 2025;4(6). doi:10.1371/journal.pdig.0000877
  24. Wang C, Chen Y. Evaluating large language models for evidence-based clinical question answering. Patterns (N Y). 2026;7(5):101519. doi:10.1016/j.patter.2026.101519
  25. de Jesus DR, de Souza Júnior AP, de Albergaria ET, Pagano AS, de Oliveira IJR, Dias CDS, et al. Enhanced LLM-supported instructions for medication use through retrieval-augmented generation. Comput Biol Med. 2025;198(Pt A):111135. doi:10.1016/j.compbiomed.2025.111135
  26. Frazer A, Yoon E, Durant KM. Artificial intelligence for drug information: Evaluating its use in real-world pharmacy practice. Am J Health Syst Pharm. 2026;83(12):669-76. doi:10.1093/ajhp/zxaf350
  27. van Nuland M, Lobbezoo AFH, van de Garde EMW, Herbrink M, van Heijl I, Bognár T, et al. Assessing accuracy of ChatGPT in response to questions from day to day pharmaceutical care in hospitals. Explor Res Clin Soc Pharm. 2024;15:100464. doi:10.1016/j.rcsop.2024.100464
  28. Triplett S, Ness-Engle GL, Behnen EM. A comparison of drug information question responses by a drug information center and by ChatGPT. Am J Health Syst Pharm. 2025;82(8):448-60. doi:10.1093/ajhp/zxae316
  29. Sridharan K, Sivaramakrishnan G. Unlocking the potential of advanced large language models in medication review and reconciliation: A proof-of-concept investigation. Explor Res Clin Soc Pharm. 2024;15:100492. doi:10.1016/j.rcsop.2024.100492
  30. Albogami Y, Alfakhri A, Alaqil A, Alkoraishi A, Alshammari H, Elsharawy Y, et al. Safety and quality of AI chatbots for drug-related inquiries: A real-world comparison with licensed pharmacists. Digit Health. 2024;10:20552076241253523. doi:10.1177/20552076241253523
  31. Krichevsky B, Engeli S, Bode-Böger SM, Koop F, Schulze Westhoff M, Schröder S, et al. Human vs. artificial intelligence: Physicians outperform ChatGPT in real-world pharmacotherapy counselling. Br J Clin Pharmacol. 2026;92(3):891-902. doi:10.1002/bcp.70321
  32. Yang R, Wong MYH, Li H, Li X, Zhu W, Liao J, et al. Retrieval-augmented generation in medicine: A scoping review of technical implementations, clinical applications, and ethical considerations. Cell Rep Med. 2026. Online ahead of print. Article 102927. doi:10.1016/j.xcrm.2026.102927
  33. Asgari E, Montaña-Brown N, Dubois M, Khalil S, Balloch J, Au Yeung J, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. NPJ Digit Med. 2025;8(1):274. doi:10.1038/s41746-025-01670-7
  34. Chen F, Li Y, Chen Y, Bian Z, Duo L, Zhou Q, et al. Strategies for the analysis and elimination of hallucinations in artificial intelligence generated medical knowledge. J Evid Based Med. 2025;18(3). doi:10.1111/jebm.70075
  35. Vaccaro M, Almaatouq A, Malone T. When combinations of humans and AI are useful: A systematic review and meta-analysis. Nat Hum Behav. 2024;8(12):2293-303. doi:10.1038/s41562-024-02024-1
  36. Topol EJ. High-performance medicine: The convergence of human and artificial intelligence. Nat Med. 2019;25(1):44-56. doi:10.1038/s41591-018-0300-7
  37. Gilbert S, Harvey H, Melvin T, Vollebregt E, Wicks P. Large language model AI chatbots require approval as medical devices. Nat Med. 2023;29(10):2396-8. doi:10.1038/s41591-023-02412-6

 

 

 


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.