Enhancing Virtual Assistants Through Context-aware Multimodal Emotion Recognition using Visual, Textual, and Speech Intelligence

Authors

  • Harshitha G. S
  • Prathibhavani P. M

Keywords:

Context-aware systems, DeBERTa-v3, Emotion recognition, Human-computer interaction, Multimodal learning, Swin transformer, Virtual assistant, Wav2Vec2

Abstract

The growing adoption of virtual assistants has transformed the way users interact with digital systems. However, most existing assistants primarily rely on textual commands and often overlook the emotional state of the user, resulting in interactions that may feel impersonal or contextually inadequate. Human emotions are naturally expressed through multiple channels, including facial expressions, spoken language, and textual communication. Motivated by this observation, this work proposes the Adaptive Context-Aware Multimodal Emotion-Aware Virtual Assistant (ACME-VA), a framework designed to recognize and interpret emotions from facial, textual, and speech inputs simultaneously. The proposed system employs a Swin Transformer-based model for facial emotion recognition, DeBERTa-v3 for textual emotion analysis, and Wav2Vec2 for speech emotion recognition. To improve the reliability of emotion prediction, a context-aware reasoning module is incorporated to analyze modality confidence, maintain emotional history, and identify temporal changes in user emotions. The emotional representations obtained from different modalities are combined using a cross-modal attention mechanism to generate a comprehensive understanding of the user’s emotional state. Based on the final prediction, the assistant generates responses that are both contextually relevant and emotionally adaptive.

References

S. G. Rajesh, S. V. Madangarli, G. S. Pisharady, and R. Subrahmanyam, “Enhancement of virtual assistants through multimodal AI for emotion recognition,” IEEE Access, vol. 13, pp. 102159–102179, Jun. 9, 2025.

J. Goodfellow, D. Erhan, P. L. Carrier, A. Courville, M. Mirza, B. Hamner, W. Cukierski, Y. Tang, D. Thaler, D. H. Lee, and Y. Zhou, “Challenges in representation learning: A report on three machine learning contests,” Neural Networks, vol. 64, pp. 59–63, Apr. 2015.

A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30, 2017.

D. Demszky, E. Movshovitz-Attias, D. Ko, A. Cowen, G. Nemade, and S. Ravi, “GoEmotions: A dataset of fine-grained emotions,” in Proc. 58th Annu. Meeting Assoc. Comput. Linguistics (ACL), 2020, pp. 4040–4054.

J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. 2019 Conf. North Amer. Chapter Assoc. Comput. Linguistics: Human Language Technologies, vol. 1, 2019, pp. 4171–4186.

Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “RoBERTa: A robustly optimized BERT pretraining approach,” arXiv preprint arXiv:1907.11692, Jul. 26, 2019.

P. He, X. Liu, J. Gao, and W. Chen, “DeBERTa: Decoding-enhanced BERT with disentangled attention,” arXiv preprint arXiv:2006.03654, Jun. 5, 2020.

Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin Transformer: Hierarchical vision transformer using shifted windows,” In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 9992–10002.

A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” In Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 12449–12460.

S. R. Livingstone and F. A. Russo, “The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English,” PLoS ONE, vol. 13, no. 5, Art. no. e0196391, May 16, 2018.

T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423–443, Feb. 2019.

A. Z. Zeng, M. Pantic, G. I. Roisman, and T. S. Huang, “A survey of affect recognition methods: Audio, visual, and spontaneous expressions,” Proceedings of the 9th International Conference on Multimodal Interfaces, 2007, pp. 126–133.

Mollahosseini, B. Hasani, and M. H. Mahoor, “AffectNet: A database for facial expression, valence, and arousal computing in the wild,” In Proceedings of the 9th International Conference on Multimodal Interfaces, 2007, vol. 10, no. 1, pp. 18–31, Jan.–Mar. 2019.

Y. Zhang, S. Sun, M. Galley, Y.-C. Chen, C. Brockett, X. Gao, J. Gao, J. Liu, and W. B. Dolan, “DialoGPT: Large-scale generative pre-training for conversational response generation,” In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, 2020, pp. 270–278.

W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou, “MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers,” In Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 5776–5788.

C. Busso, M. Bulut, C. C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” Language Resources and Evaluation, vol. 42, no. 4, pp. 335–359, Dec. 2008.

S. Poria, N. Majumder, D. Hazarika, E. Cambria, A. Gelbukh, and A. Hussain, “Multimodal sentiment analysis: Addressing key issues and setting up the baselines,” IEEE Intelligent Systems, vol. 33, no. 6, pp. 17–25, Nov.–Dec. 2018.

Y. Kim, “Convolutional neural networks for sentence classification,” in Proc. 2014 Conf. Empirical Methods in Natural Language Processing (EMNLP), 2014, pp. 1746–1751.

Published

2026-08-13