| Sumario: | Purpose: Childhood trauma can disrupt communication, yet early signs often go unrecognized in regions affected by ongoing war, where immediate physical needs take precedence. Vocal biomarkers—acoustic features linked to emotional and motor regulation—offer a promising, noninvasive means of detecting traumalinked speech disruptions. This study applied a hybrid framework to distinguish trauma exposure in Arabic-speaking children living amid active conflict. The aim was to support scalable, speech-based tools for early trauma identification in low-resource, humanitarian settings. Method: We analyzed 200 publicly available recordings of spontaneous speech from Arabic-speaking girls (ages 8-12 years): 100 trauma-exposed participants from Gaza (Palestinian) and 100 non-exposed controls from Jordan. Core acoustic features (fundamental frequency [F0], jitter, shimmer, harmonics-to-noise ratio [HNR], voice onset time [VOT], first formant, second formant) informed statistical testing and theory-driven composite indices. Exploratory features—including Mel-frequency cepstral coefficients and eGeMAPSv02 descriptors—were used to train binary classification models. Three classifiers (random forest, ridge regression, and logistic regression) were evaluated using nested cross-validation and bootstrap resampling. Composite indices were combined into a 0-10 Trauma Risk Score. Generalizability was assessed using an independent Lebanese cohort(n= 80), with the trained classifier applied using fixed exploratory features and preprocessing parameters. Core features were tested post hoc for cross-cohort stability. Results: Trauma-exposed children showed reduced F0 and HNR, elevated shimmer and jitter, and prolonged VOT (Cohen's d > 1.2). Binary classification models achieved strong performance (area under the curve [AUC] = .89-.92); logistic regression reached AUC = .996 under cross-validation. Composite indices (AUCs > .90) stratified 68% into Moderate/High Trauma Risk. Lebanese validation confirmed generalizability, with theory-driven features showing stable predictive patterns. Conclusions: Vocal biomarkers reliably distinguished trauma exposure in Arabicspeaking children using a simple logistic regression model. This strong performance highlights the potential of speech-based tools as scalable, noninvasive methods for early trauma detection. Further validation is needed to support their use in diverse humanitarian and conflict-affected settings.
|