TY - JOUR T1 - Can Large Language Models Be Trusted with Pharmacy Work? A Systematic Review of Accuracy, Safety, and Clinical Use A1 - Daniel Fischer A1 - Laura Meier A1 - Stefan Koch A1 - Thomas Braun JF - Archives of Pharmacy Practice JO - Arch Pharm Pract SN - 2320-5210 Y1 - 2026 VL - 17 IS - 3 DO - 10.51847/WHAwvGP2Pm SP - 10 EP - 18 N2 - Large language models (LLMs) can generate medication information, counselling content, interaction assessments, and pharmacotherapy suggestions, but answer quality alone does not establish safe or useful pharmacy practice. To determine how accurately and safely LLM-based systems perform pharmacy-relevant work, what clinical-use evidence exists, and which task-, model-, comparator-, and setting-specific boundaries constrain claims of trustworthiness. A protocol-driven systematic review was conducted in accordance with PRISMA 2020. PubMed/MEDLINE, publisher and Crossref-indexed records, and backward and forward citation searches were examined through 3 August 2026. Eligible studies were peer-reviewed Q1 journal articles evaluating an LLM or LLM-enabled system on a medication-work task. Two reviewers independently screened records, extracted study and model characteristics, harmonized outcomes, and appraised risk of bias, reproducibility, and applicability. Findings were synthesized narratively by pharmacy task and evaluation design because outcomes were unsuitable for meta-analysis. Forty-six records were identified, four duplicates were removed, 42 records were screened, 32 full-text reports were assessed, and 20 studies were included. Evidence was dominated by retrospective, cross-sectional, simulated-case, and output-rating designs. Performance varied with task complexity, model and version, prompting, grounding, reference standard, assessor, and scoring definition. Retrieval augmentation and customization improved selected output measures, but hallucination, omission, inconsistency, weak referencing, and unsafe or insufficiently qualified recommendations remained. Real-world questions or patient-derived data rarely represented prospective workflow implementation, and no included study established improved patient outcomes or safe autonomous pharmacy deployment. LLM trustworthiness in pharmacy is conditional rather than general. Current evidence supports bounded, task-specific evaluation under professional verification, not unsupervised clinical delegation. Prospective, externally validated, harm-sensitive studies are required. UR - https://archivepp.com/article/can-large-language-models-be-trusted-with-pharmacy-work-a-systematic-review-of-accuracy-safety-an-1asvqsvjootmkrg ER -