Hybrid Deep Learning Frameworks for Robotic Object Recognition

  • Authors

    • N. Seshagiri Information Technology Pioneer, National Informatics Centre, India Author
    • H. N. Mahabala Computer Scientist, Tata Institute of Fundamental Research, India Author

    DOI:

    https://doi.org/10.67228/30715725/IJIARE-2025PI4T1I

    Published 03-05-2025

  • Robotic Object Recognition, Hybrid Deep Learning, Computer Vision, Convolutional Neural Networks, Vision Transformers, Sensor Fusion, Intelligent Robotics, Autonomous Systems, Object Detection, Artificial Intelligence

    Issue

    Section

    Articles

    How to Cite

    [1]
    N. Seshagiri and H. N. Mahabala, “Hybrid Deep Learning Frameworks for Robotic Object Recognition”, IJIARE, vol. 8, no. 1, pp. 01–17, Mar. 2025, doi: 10.67228/30715725/IJIARE-2025PI4T1I.
  • Abstract

    Robotic object recognition is a fundamental capability that enables autonomous robots to interact intelligently with dynamic environments. Traditional vision-based methods, such as SIFT, SURF, HOG, and template matching, perform well under controlled conditions but struggle with variations in lighting, viewpoint, occlusion, and complex backgrounds. Recent advances in deep learning have significantly improved recognition accuracy by automatically learning features from raw image data. However, individual deep learning models often face challenges related to computational cost, inference speed, and limited generalization. This paper proposes a Hybrid Deep Learning Framework that integrates Convolutional Neural Networks (CNNs), Vision Transformers (ViTs), attention mechanisms, and multimodal sensor fusion (RGB, depth, and LiDAR) to enhance recognition accuracy and efficiency. The framework combines local and global feature extraction, adaptive feature fusion, intelligent object recognition, and robotic decision-making for real-time perception and task execution. It supports applications in industrial automation, warehouse logistics, autonomous mobile robots, healthcare, agriculture, and service robotics while improving robustness, scalability, and computational efficiency.

  • References

    [1] D. G. Lowe, "Distinctive image features from scale-invariant keypoints," International Journal of Computer Vision, vol. 60, no. 2, pp. 91–110, Nov. 2004.

    [2] H. Bay, T. Tuytelaars, and L. Van Gool, "SURF: Speeded Up Robust Features," in Proc. European Conf. Computer Vision (ECCV), Graz, Austria, 2006, pp. 404–417.

    [3] N. Dalal and B. Triggs, "Histograms of oriented gradients for human detection," in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), San Diego, CA, USA, 2005, pp. 886–893.

    [4] T. Ojala, M. Pietikäinen, and T. Mäenpää, "Multiresolution gray-scale and rotation invariant texture classification with local binary patterns," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 24, no. 7, pp. 971–987, Jul. 2002.

    [5] C. Harris and M. Stephens, "A combined corner and edge detector," in Proc. Alvey Vision Conference, Manchester, U.K., 1988, pp. 147–151.

    [6] A. Krizhevsky, I. Sutskever, and G. E. Hinton, "ImageNet classification with deep convolutional neural networks," in Advances in Neural Information Processing Systems (NeurIPS), vol. 25, 2012, pp. 1097–1105.

    [7] K. Simonyan and A. Zisserman, "Very deep convolutional networks for large-scale image recognition," in Proc. International Conf. Learning Representations (ICLR), 2015.

    [8] K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, pp. 770–778.

    [9] C. Szegedy et al., "Going deeper with convolutions," in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 2015, pp. 1–9.

    [10] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, "You Only Look Once: Unified, real-time object detection," in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, pp. 779–788.

    [11] S. Ren, K. He, R. Girshick, and J. Sun, "Faster R-CNN: Towards real-time object detection with region proposal networks," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 6, pp. 1137–1149, Jun. 2017.

    [12] M. Tan and Q. Le, "EfficientNet: Rethinking model scaling for convolutional neural networks," in Proc. International Conf. Machine Learning (ICML), 2019, pp. 6105–6114.

    [13] A. Vaswani et al., "Attention is all you need," in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017, pp. 5998–6008.

    [14] A. Dosovitskiy et al., "An image is worth 16×16 words: Transformers for image recognition at scale," in Proc. International Conf. Learning Representations (ICLR), 2021.

    [15] S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, "CBAM: Convolutional Block Attention Module," in Proc. European Conf. Computer Vision (ECCV), Munich, Germany, 2018, pp. 3–19.

    [16] Gajula, S. (2024). Cybersecurity risk prediction using graph neural networks. Journal of Information Systems Engineering and Management.

    [17] Gajula, S. (2024). Adaptive zero trust architecture for securing financial microservices. Computer Fraud & Security, 2024(12), 643–655. https://doi.org/10.52710/cfs.845.

  • Downloads

Similar Articles

21-30 of 84

You may also start an advanced similarity search for this article.