Real-Time Video Analytics Using Deep Learning Algorithms
-
DOI:
https://doi.org/10.67228/30713315/IJAIDT-2024PI4T8K1Published 05-05-2024
Real-Time Video Analytics, Deep Learning, Convolutional Neural Networks, Object Detection, Action Recognition, Edge Computing, Computer Vision, Video Surveillance Issue
Section
ArticlesHow to Cite
[1]E. Williams and R. M. Aslan, “Real-Time Video Analytics Using Deep Learning Algorithms”, IJAIDT, vol. 7, no. 1, pp. 01–13, May 2024, doi: 10.67228/30713315/IJAIDT-2024PI4T8K1.Abstract
Real-time video analytics has advanced rapidly with deep learning and high-performance computing, enabling applications in surveillance, smart cities, healthcare, and industry. Traditional computer vision methods struggle with complex scenes and real-time processing, while deep learning models like CNNs, RNNs, and transformers can learn spatial and temporal features directly from data. This paper reviews deep learning-based approaches for tasks such as object detection, tracking, action recognition, and event detection. It proposes a framework including data acquisition, preprocessing, model training, and deployment for low-latency, high-accuracy systems. Results show that deep learning methods outperform traditional techniques in accuracy and robustness, even under challenging conditions. The study also discusses challenges like computational cost, scalability, and privacy, and highlights future directions such as lightweight models, edge computing, self-supervised learning, and explainable AI.
References
[1] R. Cucchiara, C. Grana, M. Piccardi, and A. Prati, “Detecting moving objects, ghosts, and shadows in video streams,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 25, no. 10, pp. 1337–1342, 2003.
[2] A. Elgammal, D. Harwood, and L. Davis, “Non-parametric model for background subtraction,” in Proc. European Conference on Computer Vision (ECCV), 2000, pp. 751–767.
[3] T. Bouwmans, F. Porikli, B. Höferlin, and A. Vacavant, Background Modeling and Foreground Detection for Video Surveillance, CRC Press, 2014.
[4] R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp. 580–587.
[5] R. Girshick, “Fast R-CNN,” in Proc. IEEE International Conference on Computer Vision (ICCV), 2015, pp. 1440–1448.
[6] S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 6, pp. 1137–1149, 2017.
[7] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 779–788.
[8] W. Liu et al., “SSD: Single shot multibox detector,” in Proc. European Conference on Computer Vision (ECCV), 2016, pp. 21–37.
[9] G. Jocher et al., “YOLOv5,” GitHub repository, 2020. [Online]. Available: https://github.com/ultralytics/yolov5
[10] D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3D convolutional networks,” in Proc. IEEE International Conference on Computer Vision (ICCV), 2015, pp. 4489–4497.
[11] J. Donahue et al., “Long-term recurrent convolutional networks for visual recognition and description,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 2625–2634.
[12] K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in Advances in Neural Information Processing Systems (NeurIPS), 2014, pp. 568–576.
[13] A. Vaswani et al., “Attention is all you need,” in Advances in Neural Information Processing Systems (NeurIPS), 2017, pp. 5998–6008.
[14] G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” in Proc. International Conference on Machine Learning (ICML), 2021, pp. 813–824.
[15] S. Teerapittayanon, B. McDanel, and H. T. Kung, “Distributed deep neural networks over the cloud, the edge and end devices,” in Proc. IEEE International Conference on Distributed Computing Systems (ICDCS), 2017, pp. 328–339.
[16] Gajula, S. (2023). A review of anomaly identification in finance frauds using machine learning system. International Journal of Current Engineering and Technology, 13(6), 568–575. https://ijcet.evegenis.org/index.php/ijcet/article/view/820
Downloads
How to Cite
[1]E. Williams and R. M. Aslan, “Real-Time Video Analytics Using Deep Learning Algorithms”, IJAIDT, vol. 7, no. 1, pp. 01–13, May 2024, doi: 10.67228/30713315/IJAIDT-2024PI4T8K1.