Archive \ Volume.17 2026 Issue 3

Can Large Language Models Be Trusted with Pharmacy Work? A Systematic Review of Accuracy, Safety, and Clinical Use

, , ,
  1. Department of LLM Trust and Pharmacy Work, Faculty of Pharmacy, University of Freiburg, Freiburg, Germany.
  2. Department of AI Accuracy and Clinical Safety, Faculty of Pharmacy, Technical University of Munich, Munich, Germany.
  3. Department of LLM Evaluation in Pharmacy Practice, Faculty of Pharmacy, University of Kiel, Kiel, Germany.

Abstract

Large language models (LLMs) can generate medication information, counselling content, interaction assessments, and pharmacotherapy suggestions, but answer quality alone does not establish safe or useful pharmacy practice. To determine how accurately and safely LLM-based systems perform pharmacy-relevant work, what clinical-use evidence exists, and which task-, model-, comparator-, and setting-specific boundaries constrain claims of trustworthiness. A protocol-driven systematic review was conducted in accordance with PRISMA 2020. PubMed/MEDLINE, publisher and Crossref-indexed records, and backward and forward citation searches were examined through 3 August 2026. Eligible studies were peer-reviewed Q1 journal articles evaluating an LLM or LLM-enabled system on a medication-work task. Two reviewers independently screened records, extracted study and model characteristics, harmonized outcomes, and appraised risk of bias, reproducibility, and applicability. Findings were synthesized narratively by pharmacy task and evaluation design because outcomes were unsuitable for meta-analysis. Forty-six records were identified, four duplicates were removed, 42 records were screened, 32 full-text reports were assessed, and 20 studies were included. Evidence was dominated by retrospective, cross-sectional, simulated-case, and output-rating designs. Performance varied with task complexity, model and version, prompting, grounding, reference standard, assessor, and scoring definition. Retrieval augmentation and customization improved selected output measures, but hallucination, omission, inconsistency, weak referencing, and unsafe or insufficiently qualified recommendations remained. Real-world questions or patient-derived data rarely represented prospective workflow implementation, and no included study established improved patient outcomes or safe autonomous pharmacy deployment. LLM trustworthiness in pharmacy is conditional rather than general. Current evidence supports bounded, task-specific evaluation under professional verification, not unsupervised clinical delegation. Prospective, externally validated, harm-sensitive studies are required.


Downloads: 30
Views: 63

How to cite:
Vancouver
Fischer D, Meier L, Koch S, Braun T. Can Large Language Models Be Trusted with Pharmacy Work? A Systematic Review of Accuracy, Safety, and Clinical Use. Arch Pharm Pract. 2026;17(3):10-8. https://doi.org/10.51847/WHAwvGP2Pm
APA
Fischer, D., Meier, L., Koch, S., & Braun, T. (2026). Can Large Language Models Be Trusted with Pharmacy Work? A Systematic Review of Accuracy, Safety, and Clinical Use. Archives of Pharmacy Practice, 17(3), 10-18. https://doi.org/10.51847/WHAwvGP2Pm

Download Citation
References
  1. Ong JCL, Chen MH, Ng N, Elangovan K, Tan NYT, Jin L, et al. A scoping review on generative AI and large language models in mitigating medication related harm. NPJ Digit Med. 2025;8(1):182. doi:10.1038/s41746-025-01565-7
  2. Bates DW, Levine D, Syrowatka A, Kuznetsova M, Craig KJT, Rui A, et al. The potential of artificial intelligence to improve patient safety: A scoping review. NPJ Digit Med. 2021;4(1):54. doi:10.1038/s41746-021-00423-6
  3. Choudhury A, Asan O. Role of artificial intelligence in patient safety outcomes: Systematic literature review. JMIR Med Inform. 2020;8(7). doi:10.2196/18599
  4. Syrowatka A, Song W, Amato MG, Foer D, Edrees H, Co Z, et al. Key use cases for artificial intelligence to reduce the frequency of adverse drug events: A scoping review. Lancet Digit Health. 2022;4(2). doi:10.1016/S2589-7500(21)00229-6
  5. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ. 2021;372. doi:10.1136/bmj.n71
  6. Rethlefsen ML, Kirtley S, Waffenschmidt S, Ayala AP, Moher D, Page MJ, et al. PRISMA-S: An extension to the PRISMA statement for reporting literature searches in systematic reviews. Syst Rev. 2021;10(1):39. doi:10.1186/s13643-020-01542-z
  7. Moons KGM, Damen JAA, Kaul T, Hooft L, Andaur Navarro CL, Dhiman P, et al. PROBAST+AI: An updated quality, risk-of-bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388. doi:10.1136/bmj-2024-082505
  8. Vasey B, Nagendran M, Campbell B, Clifton DA, Collins GS, Denaxas S, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med. 2022;28(5):924-33. doi:10.1038/s41591-022-01772-9
  9. Singhal K, Azizi S, Tu T, Mahdavi SS, Wei J, Chung HW, et al. Large language models encode clinical knowledge. Nature. 2023;620(7972):172-80. doi:10.1038/s41586-023-06291-2
  10. Lee P, Bubeck S, Petro J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N Engl J Med. 2023;388(13):1233-9. doi:10.1056/NEJMsr2214184
  11. Liu S, Wright AP, Patterson BL, Wanderer JP, Turer RW, Nelson SD, et al. Using AI-generated suggestions from ChatGPT to optimize clinical decision support. J Am Med Inform Assoc. 2023;30(7):1237-45. doi:10.1093/jamia/ocad072
  12. Haug CJ, Drazen JM. Artificial intelligence and machine learning in clinical medicine, 2023. N Engl J Med. 2023;388(13):1201-8. doi:10.1056/NEJMra2302038
  13. Campbell M, McKenzie JE, Sowden A, Katikireddi SV, Brennan SE, Ellis S, et al. Synthesis without meta-analysis (SWiM) in systematic reviews: Reporting guideline. BMJ. 2020;368. doi:10.1136/bmj.l6890
  14. Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385. doi:10.1136/bmj-2023-078378
  15. Nagendran M, Chen Y, Lovejoy CA, Gordon AC, Komorowski M, Harvey H, et al. Artificial intelligence versus clinicians: Systematic review of design, reporting standards, and claims of deep learning studies. BMJ. 2020;368. doi:10.1136/bmj.m689
  16. Haibe-Kains B, Adam GA, Hosny A, Khodakarami F, Massive Analysis Quality Control (MAQC) Society Board of Directors, Waldron L, et al. Transparency and reproducibility in artificial intelligence. Nature. 2020;586(7829). doi:10.1038/s41586-020-2766-y
  17. Huang X, Estau D, Liu X, Yu Y, Qin J, Li Z. Evaluating the performance of ChatGPT in clinical pharmacy: A comparative study of ChatGPT and clinical pharmacists. Br J Clin Pharmacol. 2024;90(1):232-8. doi:10.1111/bcp.15896
  18. Roosan D, Padua P, Khan R, Khan H, Verzosa C, Wu Y. Effectiveness of ChatGPT in clinical pharmacy and the role of artificial intelligence in medication therapy management. J Am Pharm Assoc (2003). 2024;64(2):422-8. doi:10.1016/j.japh.2023.11.023
  19. Morath B, Chiriac U, Jaszkowski E, Deiss C, Nurnberg H, Horth K, et al. Performance and risks of ChatGPT used in drug information: An exploratory real-world analysis. Eur J Hosp Pharm. 2024;31(6):491-7. doi:10.1136/ejhpharm-2023-003750
  20. Al-Dujaili Z, Al Omari S, Pillai J, Al Faraj A. Assessing the accuracy and consistency of ChatGPT in clinical pharmacy management: A preliminary analysis with clinical pharmacy experts worldwide. Res Social Adm Pharm. 2023;19(12):1590-4. doi:10.1016/j.sapharm.2023.08.012
  21. Grossman S, Zerilli T, Nathan JP. Appropriateness of ChatGPT as a resource for medication-related questions. Br J Clin Pharmacol. 2024;90(10):2691-5. doi:10.1111/bcp.16212
  22. van Nuland M, Erdogan A, Acar C, Contrucci R, Hilbrants S, Maanach L, et al. Performance of ChatGPT on factual knowledge questions regarding clinical pharmacy. J Clin Pharmacol. 2024;64(9):1095-100. doi:10.1002/jcph.2443
  23. Montastruc F, Storck W, de Canecaude C, Victor L, Li J, Cesbron C, et al. Will artificial intelligence chatbots replace clinical pharmacologists? An exploratory study in clinical practice. Eur J Clin Pharmacol. 2023;79(10):1375-84. doi:10.1007/s00228-023-03547-8
  24. Li L, Du P, Huang X, Zhao H, Ni M, Yan M, et al. Comparative analysis of generative artificial intelligence systems in solving clinical pharmacy problems: Mixed methods study. JMIR Med Inform. 2025;13. doi:10.2196/76128
  25. Kiyomiya K, Aomori T, Ohtani H. Medication counseling for OTC drugs using customized ChatGPT-4: Comparison with ChatGPT-3.5 and ChatGPT-4o. Digit Health. 2025;11:20552076251323810. doi:10.1177/20552076251323810
  26. Shin E, Hartman M, Ramanathan M. Performance of the ChatGPT large language model for decision support in community pharmacy. Br J Clin Pharmacol. 2024;90(12):3320-33. doi:10.1111/bcp.16215
  27. Albogami Y, Alfakhri A, Alaqil A, Alkoraishi A, Alshammari H, Elsharawy Y, et al. Safety and quality of AI chatbots for drug-related inquiries: A real-world comparison with licensed pharmacists. Digit Health. 2024;10:20552076241253523. doi:10.1177/20552076241253523
  28. Hsu HY, Hsu KC, Hou SY, Wu CL, Hsieh YW, Cheng YD. Examining real-world medication consultations and drug-herb interactions: ChatGPT performance evaluation. JMIR Med Educ. 2023;9. doi:10.2196/48433
  29. de Jesus DDR, de Souza Junior AP, de Albergaria ET, Pagano AS, de Oliveira IJR, Dos Santos Dias C, et al. Enhanced LLM-supported instructions for medication use through retrieval-augmented generation. Comput Biol Med. 2025;198(Pt A):111135. doi:10.1016/j.compbiomed.2025.111135
  30. Andrikyan W, Sametinger SM, Kosfeld F, Jung-Poppe L, Fromm MF, Maas R, et al. Artificial intelligence-powered chatbots in search engines: A cross-sectional study on the quality and risks of drug information for patients. BMJ Qual Saf. 2025;34(2):100-9. doi:10.1136/bmjqs-2024-017476
  31. Krichevsky B, Engeli S, Bode-Boger SM, Koop F, Schulze Westhoff M, Schroder S, et al. Human vs artificial intelligence: Physicians outperform ChatGPT in real-world pharmacotherapy counselling. Br J Clin Pharmacol. 2026;92(3):891-902. doi:10.1002/bcp.70321
  32. Ong JCL, Jin L, Elangovan K, Lim GYS, Lim DYZ, Sng GGR, et al. Large language model as clinical decision support system augments medication safety in 16 clinical specialties. Cell Rep Med. 2025;6(10):102323. doi:10.1016/j.xcrm.2025.102323
  33. Radha Krishnan RP, Hung EH, Ashford M, Edillo CE, Gardner C, Hatrick HB, et al. Evaluating ChatGPT performance in predicting drug-drug interactions from hospitalized patient data. Br J Clin Pharmacol. 2024;90(12):3361-6. doi:10.1111/bcp.16275
  34. Buzancic I, Belec D, Drzaic M, Kummer I, Brkic J, Fialova D, et al. Clinical decision-making in benzodiazepine deprescribing by healthcare providers vs AI-assisted approach. Br J Clin Pharmacol. 2024;90(3):662-74. doi:10.1111/bcp.15963
  35. Bischof T, al Jalali V, Zeitlinger M, Jorda A, Hana M, Singeorzan KN, et al. Chat GPT vs clinical decision support systems in the analysis of drug-drug interactions. Clin Pharmacol Ther. 2025;117(4):1142-7. doi:10.1002/cpt.3585
  36. Chase A, Most A, Sikora A, Smith SE, Devlin JW, Xu S, et al. Evaluation of large language models' ability to identify clinically relevant drug-drug interactions and generate high-quality clinical pharmacotherapy recommendations. Am J Health Syst Pharm. 2026;83(13). doi:10.1093/ajhp/zxaf168
  37. Cabitza F, Rasoini R, Gensini GF. Unintended consequences of machine learning in medicine. JAMA. 2017;318(6):517-8. doi:10.1001/jama.2017.7797
  38. Ghassemi M, Oakden-Rayner L, Beam AL. The false hope of current approaches to explainable artificial intelligence in health care. Lancet Digit Health. 2021;3(11). doi:10.1016/S2589-7500(21)00208-9
  39. Wiens J, Saria S, Sendak M, Ghassemi M, Liu VX, Doshi-Velez F, et al. Do no harm: A roadmap for responsible machine learning for health care. Nat Med. 2019;25(9):1337-40. doi:10.1038/s41591-019-0548-6

 

 

 


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.