Abstract / Summary
Objectives. LLMs are increasingly used to support academic writing, but their tendency to hallucinate citations raises concerns about the integrity of AI-assisted scholarly work. This study provides a timely snapshot of citation reliability in LLM-generated academic reviews by testing whether citation fabrication and accuracy differ across four LLMs and review topic specificity. Method. ChatGPT-5.4 Plus, ChatGPT-5.3 Free, Gemini-3, and Claude Sonnet-4.6 were tasked with generating 16 literature reviews on mental health, divided into general topics (broad areas with substantial evidence bases) and specialised topics (narrower areas with smaller evidence bases). Results. Only 6.1% of the 825 LLM-generated citations were fabricated, with fabrication absent for ChatGPT-5.4 Plus (0%), very low for Claude Sonnet-4.6 (0.5%) and ChatGPT-5.3 Free (5.5%), and highest for Google Gemini-3 (20.2%). A similar model-level pattern emerged for citation accuracy (real citations containing no bibliographic errors). Review topic specificity was not associated with citation fabrication or accuracy. Conclusion. Findings suggest possible improvements in citation reliability relative to estimates from prior studies testing older model iterations, while also showing that current performance is strongly model-dependent. These findings highlight the need for rigorous human verification of LLM-generated references and stronger safeguards to protect research integrity as LLMs are integrated into research.