Abstract / Summary
Facial emotion-recognition tasks are widely used to investigate social-cognitive differences across mental health conditions, but theirscientific value depends on the faces used to measure performance. Established photographic databases are difficult and expensive toexpand, and many provide little ethnic diversity. AI-generated faces offer a more representative, flexible, and scalable alternative,provided they capture the same clinically relevant individual differences as photographs of human faces. We compared 161photorealistic AI-generated faces with 72 images from FACES, an established human-photograph set. Prolific participants contributed83,206 valid trials; 77,238 trials covering the six shared emotions entered the primary comparison (accuracy \textit{n} = 1,047; responsetime \textit{n} = 1,046). AI-diverse expressions were recognised more accurately (89.12\% vs.\ 81.72\%) and faster (2.29 vs.\ 2.64 s) thanFACES expressions, while AI-diverse and AI-White performance did not differ significantly. The AI task retained the individual differences needed for research: lower accuracy on both sources was associated most consistently with psychotic-like experiences and obsessive-compulsive symptoms, with magnitudes aligned with previous literature. AI-diverse and FACES accuracy correlated at \textit{r} = .56, and corrected split-half reliability was .70 and .60, respectively. Women were more accurate than men; participant ethnicity was unrelated to performance. Higher AI accuracy indicates that the expressions were clear; it does not show general psychometric superiority, possibly because very easy items can compress differences among high performers. The combination of clear expressions, comparable symptom sensitivity, and broader representation shows that generative AI can modernise a core social-cognition research tool while making stimulus design more controllable, expandable, and shareable.