SentinelMind: A Four-channel Deep Learning Platform for Behavioral Depression Screening through Integrated Text, Speech, Facial, and Video Signal Processing
Abstract
Mental health disorders, particularly clinical depression, represent a major global public health crisis that is often underdiagnosed due to the subjective nature of traditional diagnostic practices. While automated screening tools have emerged, they commonly rely on a single behavioral modality, such as written text or speech prosody, failing to capture the full spectrum of observable patient indicators. This paper presents SentinelMind, an integrated, multi-modal screening platform that simultaneously processes four distinct behavioral channels: written text, spoken audio, static facial images, and temporal video signals. The platform utilizes specialized deep learning encoders, incorporating a fine-tuned RoBERTa transformer for semantic language analysis, a Wav2Vec2 model for speech emotion mapping, and a DeepFace visual engine paired with an OpenCV detection backend to extract static and dynamic facial expression distributions. Individual modality predictions are dynamically combined through an adaptive Cross-Modal Gated Attention Fusion mechanism that adjusts weighting based on signal confidence. Deployed as a lightweight Flask web application, SentinelMind operates strictly on commodity hardware under a 3.2 GB peak RAM footprint. Experimental results demonstrate that the four-channel fused configuration achieves an overall binary classification F1-score of 92.0 %, representing a significant performance gain over bimodal (88.3 %) and individual channel baselines. The system delivers an explainable and responsive screening utility suitable for clinical workflows.
References
World Health Organization, “Depressive disorder (depression),” WHO Fact Sheets, 2024.
A. Satiani, J. Niedermier, B. Satiani and D. P. Svendsen, Projected workforce of psychiatrists in the United States: A population analysis,” Psychiatric Services, vol. 69, no. 6, pp. 710–713, 2018.
K. Kroenke, R. L. Spitzer and J. B. W. Williams, “The PHQ-9: Validity of a brief depression severity measure,” Journal of General Internal Medicine, vol. 16, pp. 606–613, 2001.
A. T. Beck, R. A. Steer and G. K. Brown, “Beck depression inventory-II,” Psychological Assessment, 1996.
Y. Li, S. Kumbale, Y. Chen, T. Surana, E. S. Chng and C. Guan, “Automated depression detection from text and audio: A systematic review,” in IEEE Journal of Biomedical and Health Informatics, vol. 29, no. 10, pp. 7498–7513, Oct. 2025.
American Psychiatric Association, Diagnostic and Statistical Manual of Mental Disorders, 5th ed. Arlington, VA, USA: American Psychiatric Association, 2013.
J. F. Cohn et al., “Detecting depression from facial actions and vocal prosody,” 2009 3rd International Conference on Affective Computing and Intelligent Interaction and Workshops, Amsterdam, Netherlands, 2009, pp. 1–7.
A. Pampouchidou et al., “Automatic assessment of depression based on visual cues: A systematic review,” in IEEE Transactions on Affective Computing, vol. 10, no. 4, pp. 445–470, Oct.-Dec. 2019.
M. Trotzek, S. Koitka and C. M. Friedrich, “Utilizing neural networks and linguistic metadata for early detection of depression indications in text sequences,” in IEEE Transactions on Knowledge and Data Engineering, vol. 32, no. 3, pp. 588–601, Mar. 2020.
J. Devlin, M.-W. Chang, K. Lee and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers,” Proceedings of NAACL-HLT 2019, 2018, pp. 4171–4186.
Y. Liu et al., “RoBERTa: A robustly optimized BERT pretraining approach,” arXiv, 2019.
S. Ji, T. Zhang, L. Ansari, J. Fu, P. Tiwari and E. Cambria, “MentalBERT: Publicly available pretrained language models for mental healthcare,” Proceedings of the Thirteenth Language Resources and Evaluation Conference, Jun. 2022, pp. 7184–7190.
B. Hadzic et al., “Enhancing early depression detection with AI: A comparative use of NLP models,” SICE Journal of Control, Measurement, and System Integration, vol. 17, no. 1, pp. 135–143, 2024.
X. Ma, H. Yang, Q. Chen, D. Huang and Y. Wang, “DepAudioNet: An efficient deep model for audio based depression classification,” in Proceedings of the 6th International Workshop on Audio/Visual Emotion Challenge, Oct. 2016, pp. 35–42.
A. Baevski, H. Zhou, A. Mohamed and M. Auli, “wav2vec 2.0: Self-supervised learning of speech representations,” arXiv, 2020.
S. Chen et al., “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal on Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, Oct. 2022.
W. Wu, C. Zhang and P. C. Woodland, “Self-supervised representations in speech-based depression detection,” 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 2023, pp. 1–5.
J. Ye et al., “Multi-modal depression detection based on emotional audio and evaluation text,” Journal of Affective Disorders, vol. 295, pp. 904–913, Dec. 2021.
G. Lam, H. Dongyan and W. Lin, “Context-aware deep learning for multi-modal depression detection,” 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK, 2019, pp. 3946–3950.
M. A. Uddin, J. B. Joolee and Y. -K. Lee, “Depression level prediction using deep spatiotemporal features and multilayer Bi-LTSM,” in IEEE Transactions on Affective Computing, vol. 13, no. 2, pp. 864–870, Apr.-Jun. 2022.
A. Mollahosseini, B. Hasani and M. H. Mahoor, “AffectNet: A database for facial expression, valence, and arousal computing in the wild,” in IEEE Transactions on Affective Computing, vol. 10, no. 1, pp. 18–31, Jan.–Mar. 2019.
K. He, X. Zhang, S. Ren and J. Sun, “Deep residual learning for image recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, pp. 770-778.