Abstract / Summary
Abstract We report a real, executed comparison of two common LLM adaptation strategies — retrieval-augmented generation (RAG) and LoRA fine-tuning — against zero-shot baselines, for psychiatric diagnosis on the same 24-case held-out vignette set (DSM-Bench-24) used throughout our companion papers. RAG was evaluated on two local models (llama3.2:1b, llama3.2:3b, via Ollama); LoRA fine-tuning was evaluated on a third (Qwen2.5-1.5B-Instruct, via CPU PyTorch+PEFT, after MLX proved infeasible on this machine’s macOS version and Ollama does not support training). Neither adaptation method reliably improves category-level diagnostic accuracy: RAG improved the 1B model (75.0%→79.2%) but degraded the 3B model (83.3%→75.0%); LoRA fine-tuning left Qwen’s accuracy unchanged (70.8%→70.8%, same count, different cases). The striking and consistent finding is on the safety-relevant axis: highrisk recall (on the 2 ground-truth high-risk cases) dropped to 0/2 under every single adaptation condition we tested — both RAG configurations and the fine-tuned model alike — regardless of what happened to accuracy. For fine-tuning we trace this to a specific, disclosed methodological flaw: our training data labeled the Borderline Personality Disorder training example’s risk tier as moderate rather than high, teaching the model exactly the wrong association. For RAG the mechanism is less clear, since the retrieved DSM criteria for BPD do mention self-harm; we present this as an open question rather than a resolved mechanism. We emphasize that the general phenomenon — adaptation decoupling accuracy from safety — is already established: RAG has been shown to reduce general LLM safety even with safe retrieved content (An et al., 2025), a 34-model clinical benchmark has shown RAG improving accuracy while high-risk error remains elevated (Wind et al., 2026), and an entire subfield documents benign fine-tuning degrading safety alignment (Hsu et al., 2024; Zhang et al., 2026). Our contribution is narrower: a small-scale, mechanism-transparent demonstration of this known pattern specifically for psychiatric risk-tier calibration, not its discovery