As AI and large language models (LLMs) are increasingly deployed to address global mental health challenges, a critical shortage of high-quality datasets for training and evaluation remains. According to a recent paper on arxiv.org, researchers often generate synthetic clinical personas to simulate user data and test digital mental health support systems, but most validated personas rely on English-centric contexts. The paper, by Xu and Abdullah, investigates whether similar persona-based methods can be used to generate multilingual mental health datasets.
The Problem with English-Centric Personas
The authors modified nationality and language parameters in personas to generate clinical dialogues in Mandarin, Bengali, and Hindi. They then examined how different LLMs perform when evaluating the depression severity of these generated multilingual datasets against a baseline in English. The paper reports that just adding nationality and language parameters in personas might not be adequate, as it can introduce clinical inconsistency across languages.
Findings from the Study
The study found that LLM judge models often exhibit inaccuracies in assessing depression severity in non-English texts, with performance varying across different models. This exposes systemic limitations of applying English-centric personas to multilingual contexts. The paper emphasizes that the findings highlight the urgent need for culturally responsive data generation to ensure equitable mental health systems globally.
Implications for Global AI Deployment
These results serve as a cautionary example for enterprise technology leaders deploying AI in multilingual settings. Even when AI systems are intended for global use, localization efforts that merely swap language and nationality parameters risk producing unreliable outcomes. The paper underscores that culturally responsive data generation—not just translation or nationality adjustments—is essential for building equitable AI applications across different regions and cultures.