Transformer-Based Visual Perception Models for Autonomous Robots
-
DOI:
https://doi.org/10.67228/30715725/IJIARE-2025PII4Q8NPublished 11-03-2025
Autonomous Robots, Visual Perception, Vision Transformer (Vit), Self-Attention Mechanism, Deep Learning, Object Detection, Semantic Segmentation, Multi-Modal Sensor Fusion, Intelligent Robotics, Artificial Intelligence, Computer Vision, Scene Understanding Issue
Section
ArticlesHow to Cite
[1]H. N. Mahabala, “Transformer-Based Visual Perception Models for Autonomous Robots”, IJIARE, vol. 8, no. 2, pp. 01–16, Nov. 2025, doi: 10.67228/30715725/IJIARE-2025PII4Q8N.Abstract
Autonomous robots play a crucial role in industrial manufacturing, healthcare, transportation, logistics, agriculture, disaster response, planetary exploration, and service robotics. Reliable visual perception is essential for enabling robots to recognize objects, understand scenes, localize themselves, and navigate safely in dynamic environments. Although CNN-based vision models have significantly improved perception accuracy, they often struggle to capture long-range dependencies and generalize to complex or unseen environments. Recent advances in Transformer-based vision models address these limitations by employing self-attention mechanisms to learn both local visual features and global contextual relationships. Architectures such as Vision Transformer (ViT), Swin Transformer, DETR, SAM, and Mask2Former have achieved remarkable performance in object detection, semantic segmentation, SLAM, localization, obstacle avoidance, and autonomous navigation. This paper presents a comprehensive review and proposes the Transformer-Based Visual Perception Models for Autonomous Robots (TBVPM-AR) framework. The framework integrates RGB cameras, depth sensors, LiDAR, IMUs, multimodal sensor fusion, transformer-based feature extraction, contextual reasoning, and edge-cloud computing to achieve robust perception in dynamic environments. Mathematical formulations for self-attention, positional encoding, and feature embedding provide the theoretical foundation of the architecture. Experimental evaluations demonstrate that the proposed framework outperforms CNN-based and hybrid approaches on standard robotic perception benchmarks, achieving over 98% visual perception accuracy with improved scene understanding, localization, obstacle detection, navigation, and computational efficiency. The proposed architecture offers a scalable, explainable, and adaptable solution for future Industry 5.0, collaborative robotics, autonomous vehicles, and smart cyber-physical systems.
References
[1] A. Dosovitskiy et al., "An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale," International Conference on Learning Representations (ICLR), 2021.
[2] Z. Liu et al., "Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows," Proc. IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9992–10002, 2021.
[3] N. Carion et al., "End-to-End Object Detection with Transformers," European Conference on Computer Vision (ECCV), pp. 213–229, 2021.
[4] X. Zhu et al., "Deformable DETR: Deformable Transformers for End-to-End Object Detection," International Conference on Learning Representations (ICLR), 2021.
[5] A. Radford et al., "Learning Transferable Visual Models From Natural Language Supervision," Proc. International Conference on Machine Learning (ICML), pp. 8748–8763, 2021.
[6] J. Li, D. Li, C. Xiong, and S. Hoi, "BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation," International Conference on Machine Learning (ICML), 2022.
[7] J. Alayrac et al., "Flamingo: A Visual Language Model for Few-Shot Learning," Advances in Neural Information Processing Systems (NeurIPS), vol. 35, pp. 23716–23736, 2022.
[8] A. Brohan et al., "RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control," Conference on Robot Learning (CoRL), 2023.
[9] R. Szeliski, Computer Vision: Algorithms and Applications, 2nd ed. Cham, Switzerland: Springer, 2022.
[10] Y. LeCun, Y. Bengio, and G. Hinton, "Deep Learning for Vision and Robotics: Recent Advances and Future Trends," IEEE Signal Processing Magazine, vol. 39, no. 6, pp. 84–96, Nov. 2022.
[11] J. Redmon and A. Farhadi, "YOLO-Based Real-Time Object Detection for Autonomous Robotic Systems: Recent Developments," IEEE Access, vol. 10, pp. 61584–61602, 2022.
[12] H. Caesar, V. Bankiti, A. H. Lang, et al., "nuScenes: A Multimodal Dataset for Autonomous Driving," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 3, pp. 2816–2833, Mar. 2023.
[13] L. Fei-Fei, S. Savarese, and J. Malik, "Foundation Models for Embodied Artificial Intelligence," IEEE Intelligent Systems, vol. 39, no. 1, pp. 22–34, Jan.–Feb. 2024.
[14] Y. Wang, X. Chen, H. Li, and Z. Zhang, "Transformer-Based Multi-Modal Sensor Fusion for Autonomous Robot Perception: A Survey," IEEE Access, vol. 12, pp. 45231–45258, 2024.
[15] M. Khan, P. Sharma, R. Gupta, and S. Lee, "Vision Transformer-Based Autonomous Robotic Perception: Recent Advances, Challenges, and Future Directions," IEEE Access, vol. 13, pp. 14562–14591, 2025.
[16] Gajula, S. (2024). Cybersecurity risk prediction using graph neural networks. Journal of Information Systems Engineering and Management.
[17] Gajula, S. (2024). Adaptive zero trust architecture for securing financial microservices. Computer Fraud & Security, 2024(12), 643–655. https://doi.org/10.52710/cfs.845
Downloads
How to Cite
[1]H. N. Mahabala, “Transformer-Based Visual Perception Models for Autonomous Robots”, IJIARE, vol. 8, no. 2, pp. 01–16, Nov. 2025, doi: 10.67228/30715725/IJIARE-2025PII4Q8N.
Most read articles by the same author(s)
- N. Seshagiri, H. N. Mahabala, Hybrid Deep Learning Frameworks for Robotic Object Recognition , International Journal of Intelligent Automation & Robotics Engineering: Vol. 8 No. 1 (2025)
- H. N. Mahabala, Adaptive Route Planning for Autonomous Logistics Robots in Dynamic Warehouse Environments , International Journal of Intelligent Automation & Robotics Engineering: Vol. 6 No. 2 (2023)
Similar Articles
- Andrey Ershov, Alexey Lyapunov, AI-Driven Navigation for Autonomous Inspection Robots , International Journal of Intelligent Automation & Robotics Engineering: Vol. 6 No. 2 (2023)
- Dr. Nandhini Ravi, Intelligent Robotic Pick-and-Sort Systems for Dynamic Production , International Journal of Intelligent Automation & Robotics Engineering: Vol. 4 No. 2 (2021)
- Corrado Böhm, Corrado Gini, Real-Time Embedded AI for Industrial Robot Monitoring , International Journal of Intelligent Automation & Robotics Engineering: Vol. 5 No. 2 (2022)
- Dr. Pooja Agarwal, Dr. Rakesh Chandra, Smart Robotic Systems for Hazardous Industrial Environments , International Journal of Intelligent Automation & Robotics Engineering: Vol. 1 No. 2 (2018)
- Dr. Pooja Agarwal, Dr. Rakesh Chandra, Design of Autonomous Inspection Robots for Infrastructure Monitoring , International Journal of Intelligent Automation & Robotics Engineering: Vol. 3 No. 2 (2020)
- Ole-Johan Dahl, Kristen Nygaard, Vision-Guided Robotic Assembly Using Deep Neural Networks , International Journal of Intelligent Automation & Robotics Engineering: Vol. 4 No. 2 (2021)
- Dr. Rakesh Chandra, Autonomous Navigation Framework for Mobile Robots in Highly Dynamic and Complex Indoor Environments , International Journal of Intelligent Automation & Robotics Engineering: Vol. 6 No. 1 (2023)
- H. N. Mahabala, Predictive AI Models for Intelligent Robot Health Monitoring , International Journal of Intelligent Automation & Robotics Engineering: Vol. 7 No. 2 (2024)
- Dr. Arvind Kumar Singh, Dr. Lakshmi Narayanan, Autonomous Robotic Surface Inspection Using Computer Vision , International Journal of Intelligent Automation & Robotics Engineering: Vol. 5 No. 1 (2022)
- N. Seshagiri, Autonomous Robotic Exploration Using Semantic Environment Mapping , International Journal of Intelligent Automation & Robotics Engineering: Vol. 6 No. 2 (2023)
You may also start an advanced similarity search for this article.