JARVIS: A Speaker-Authenticated Voice Command System Using GMM-SVM Fusion and NLP
Keywords:
Audio feature extraction, Biometric security, Emotion detection, Gaussian mixture model, Intelligent voice assistant, Intent recognition, Machine learningAbstract
This study presents the development of JARVIS (Just A Rather Very Intelligent System), a secure voice assistant designed for hands-free interaction and reliable user authentication. The system combines Support Vector Machine (SVM) and Gaussian Mixture Model (GMM) techniques to verify speakers with improved accuracy and protection against unauthorized access. To improve audio clarity, preprocessing methods such as noise reduction, silence removal, pre-emphasis filtering, and normalization are applied. A detailed feature set including MFCC, delta coefficients, chroma features, and spectral contrast is used to capture unique voice patterns. For command understanding, the system adopts a three-level intent recognition approach using TF-IDF-based Logistic Regression, keyword detection, and fuzzy matching. JARVIS is implemented as an interactive Streamlit application with features like waveform visualization, MFCC heatmaps, pitch tracking, emotion analysis, and command history monitoring. Experimental evaluation shows that the system delivers reliable authentication, quick response time, and effective multilingual support for English, Hindi, Kannada, Tamil, and Telugu, making it suitable for secure smart automation and voice-controlled applications.
References
S. More, J. Helonde and P. G. Burade, “Noise reduction strategies for optimising speech signal clarity: A comprehensive review,” 2025 International Conference on Sustainable Communication Networks and Application (ICSCN), Theni, India, 2025, pp. 480–485.
H. Jafarzadeh Asl, M. Ghazvini Nejad, A. Edraki, M. Asgharian and V. Partovi Nia, “Tiny noise-robust voice activity detector for voice assistants,” 2025 IEEE 35th International Workshop on Machine Learning for Signal Processing (MLSP), Istanbul, Turkiye, 2025, pp. 1–6.
F. Hoque, R. Karim, R. Kuri, K. N. Hasan and A. R. M. Mahamudul Hasan Rana, “Intent classification in chatbots with human-AI collaborated large-scale dataset,” 2025 International Conference on Electrical, Computer and Communication Engineering (ECCE), Chittagong, Bangladesh, 2025, pp. 1–7.
Z. Shi et al., “Tool learning in the wild: Empowering language models as automatic tool agents,” Proceedings of the ACM on Web Conference 2025, pp. 2222–2237, Apr. 2025.
J. Xie, S. Syu, and H.-y. Lee, “Non-instructional fine-tuning: Enabling instruction-following capabilities in pre-trained language models without instruction-following data,” arXiv, Aug. 2024.
J. Lin, M. Ge, W. Wang, H. Li, and M. Feng, “Selective HuBERT: Self-supervised pre-training for target speaker in clean and mixture speech,” IEEE signal processing letters, vol. 31, pp. 1014–1018, 2024.
J. Park et al., “Conformer-based on-device streaming speech recognition with KD compression and two-pass architecture,” 2022 IEEE Spoken Language Technology Workshop (SLT), Doha, Qatar, 2023, pp. 92–99.
S. Woo, “Whisper, A breakthrough in speech-recognition AI,” ENERZAi, May 27, 2025.
S. Kamble, A. Hande, S. Ghatul, A. P. Bangar, and A. A. Khatri, “Jarvis AI: An intelligent personal voice assistant using Python and artificial intelligence,” International Journal of Advanced Research in Science, Communication and Technology, vol. 5, no. 1, pp. 410–414, Nov. 2025.
N. Jee, S. Kumar, R. R. Patel, R. Mandal, R. K. Singh and H. Vardhan, “Advancements in voice assistants: A study of speech recognition and emotional intelligence,” 2024 13th International Conference on System Modeling & Advancement in Research Trends (SMART), Moradabad, India, 2024, pp. 283–286.
R. A. Malik, C. Setianingsih and M. Nasrun, “Speaker recognition for device controlling using MFCC and GMM algorithm,” 2020 2nd International Conference on Electrical, Control and Instrumentation Engineering (ICECIE), Kuala Lumpur, Malaysia, 2020, pp. 1–6.
D. Pawade, A. Sakhapara, R. Ashtekar, D. Bakhai and S. Tyagi, “Voice based authentication using Mel-frequency Cepstral coefficients and Gaussian Mixture Model,” 2022 IEEE Bombay Section Signature Conference (IBSSC), Mumbai, India, 2022, pp. 1–6.