Abstract / Summary
ABSTRACT Introduction Large language models (LLM) and generative Artificial Intelligence are increasingly providing health information. We have applied a novel method to evaluate the harm reduction information generated by popular LLMs. We focused on the filtration of oral medications for injection and the availability of local harm reduction services. Methods We designed a scoring rubric to evaluate the accuracy of the information provided by each LLM and applied it to four unique, commonly used LLMs. Results Gemini 2.5 Flash (87%) and ChatGPT‐4o (77%) scored higher than DeepSeek R1 Distill Llama 3.3 70B (22%) and Llama 4 (18%). Overall, both Gemini and ChatGPT‐4o gave comprehensive and accurate instructions for the use of micron filters for drug injection, while neither DeepSeek R1 nor Llama 4 provided instructions on their use. Gemini 2.5 and ChatGPT‐4o both scored highly for harm reduction general information and situational questions, while DeepSeek R1 and Llama 4 provided limited to no responses. Gemini 2.5 and ChatGPT‐4o both provided accurate, location‐specific information on local harm reduction services. Specific models also generated misinformation, further affirming the need for in‐person, nuanced discussions with harm reduction specialists. Discussion and Conclusions In this study, we presented a novel method for evaluating LLMs' potential effectiveness for harm reduction across common techniques and situations. Most of the tested LLMs provided reasonable information in a non‐judgmental manner, but they also provided misinformation, and some wouldn't answer questions due to built‐in safety parameters. The developed scoring rubric can be easily adapted to other harm reduction topics for ongoing evaluation.