Abstract / Summary
AbstractArabMindBench is a proposed multidialectal and culturally grounded benchmark for evaluating mental health artificial intelligence systems in Arabic. The benchmark focuses on clinical reasoning, uncertainty, missing-information detection, longitudinal interpretation, and safety-aware escalation across dialectal and culturally contextualized mental health narratives.The initial pilot focuses on Egyptian Arabic and a defined Saudi Arabic variety. Cases are authored natively in each dialect rather than translated, and cases representing the same clinical construct are linked through a shared concept identifier and independently reviewed for clinical equivalence.The benchmark evaluates five core tasks: risk detection, differential reasoning, missing-information identification, longitudinal reasoning, and safety/escalation. Abstention is evaluated as a cross-cutting capability using explicit information-sufficiency judgments. Independent expert annotations are retained, disagreement is preserved, and adjudication does not overwrite the original judgments. Safety-critical errors are reported separately as hard failures rather than being absorbed into a composite score.This record describes the research protocol and benchmark design prior to the completion of the pilot evaluation. No model-performance results are reported in this version. The planned development process begins with a 20-case dry run, followed by schema and annotation-guideline refinement and a planned 100–150-case feasibility pilot.Status: Pre-pilot research protocol.Clinical use: This benchmark is not a clinical decision-support system. Benchmark scores are not evidence of clinical safety and are not intended for diagnosis, treatment, or emergency decision-making.