Deep Learning Approaches for Real-Time Emotion Recognition
-
DOI:
https://doi.org/10.67228/30713315/IJAIDT-2025PI3V1ZPublished 01-04-2025
Emotion Recognition, Deep Learning, CNN, RNN, LSTM, Affective Computing, Real-Time Systems, Facial Expression Analysis, Speech Processing, Multimodal Learning Issue
Section
ArticlesHow to Cite
[1]V. Iyer, “Deep Learning Approaches for Real-Time Emotion Recognition”, IJAIDT, vol. 8, no. 1, pp. 01–12, Jan. 2025, doi: 10.67228/30713315/IJAIDT-2025PI3V1Z.Abstract
Real-time emotion recognition has become an important area in affective computing due to advances in deep learning and the need for intelligent human-computer interaction. This paper reviews pre-2019 deep learning approaches for emotion recognition using modalities like facial expressions, speech, and multimodal data. Traditional methods using handcrafted features (LBP, HOG, MFCC) had limited performance, while deep learning models such as CNNs, RNNs, LSTMs, and DBNs improved accuracy through automatic feature extraction and temporal learning. CNNs performed well in facial recognition tasks, while RNNs and LSTMs were effective for speech analysis. Multimodal systems showed higher accuracy but faced challenges like synchronization and complexity. Despite improvements, real-time deployment remains difficult due to high computational demands, cultural variability, and environmental noise. Overall, deep learning has significantly advanced emotion recognition, but challenges like efficiency, generalization, and interpretability still need to be addressed.
References
[1] Paul Ekman, “An argument for basic emotions,” Cognition and Emotion, 1992.
[2] Timo Ojala et al., “Multiresolution gray-scale and rotation invariant texture classification with local binary patterns,” IEEE TPAMI, 2002.
[3] Navneet Dalal and Bill Triggs, “Histograms of oriented gradients for human detection,” CVPR, 2005.
[4] Gajula, S. (2024). Cybersecurity risk prediction using graph neural networks. Journal of Information Systems Engineering and Management.
[5] Steven B. Davis and Paul Mermelstein, “Comparison of parametric representations for monosyllabic word recognition,” IEEE TASSP, 1980.
[6] Geoffrey Hinton et al., “A fast learning algorithm for deep belief nets,” Neural Computation, 2006.
[7] Gajula, S. (2023). A review of anomaly identification in finance frauds using machine learning system. International Journal of Current Engineering and Technology, 13(6), 568–575. https://ijcet.evegenis.org/index.php/ijcet/article/view/820
[8] Yann LeCun et al., “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, 1998.
[9] Alex Krizhevsky et al., “ImageNet classification with deep convolutional neural networks,” NIPS, 2012.
[10] Sepp Hochreiter and Jürgen Schmidhuber, “Long short-term memory,” Neural Computation, 1997.
[11] Jeffrey Elman, “Finding structure in time,” Cognitive Science, 1990.
[12] Björn Schuller et al., “Automatic recognition of emotion-related user states in spontaneous children’s speech,” IEEE TASLP, 2011.
[13] Gajula, S. (2024). Adaptive zero trust architecture for securing financial microservices. Computer Fraud & Security, 2024(12), 643–655. https://doi.org/10.52710/cfs.845
[14] Maja Pantic and Leon J. M. Rothkrantz, “Toward an affect-sensitive multimodal human–computer interaction,” Proceedings of the IEEE, 2003.
[15] Zeng Zhihong et al., “A survey of affect recognition methods: Audio, visual, and spontaneous expressions,” IEEE TPAMI, 2009.
[16] Sayan Saha et al., “A study on emotion recognition from speech using deep learning,” 2018 International Conference, 2018.
[17] Shizhe Chen et al., “Multimodal emotion recognition with deep learning,” ACM Multimedia, 2017.
[18] Soujanya Poria et al., “Multimodal sentiment analysis: Addressing key issues and setting up the baselines,” IEEE Intelligent Systems, 2017.
Downloads
How to Cite
[1]V. Iyer, “Deep Learning Approaches for Real-Time Emotion Recognition”, IJAIDT, vol. 8, no. 1, pp. 01–12, Jan. 2025, doi: 10.67228/30713315/IJAIDT-2025PI3V1Z.