Transformer-Based Visual Perception Models for Autonomous Robots
-
DOI:
https://doi.org/10.67228/30715725/IJIARE-2025PII4Q8NPublished 11-03-2025
Autonomous Robots, Visual Perception, Vision Transformer (Vit), Self-Attention Mechanism, Deep Learning, Object Detection, Semantic Segmentation, Multi-Modal Sensor Fusion, Intelligent Robotics, Artificial Intelligence, Computer Vision, Scene Understanding Issue
Section
ArticlesHow to Cite
[1]H. N. Mahabala, “Transformer-Based Visual Perception Models for Autonomous Robots”, IJIARE, vol. 8, no. 2, pp. 01–16, Nov. 2025, doi: 10.67228/30715725/IJIARE-2025PII4Q8N.Abstract
Autonomous robots play a crucial role in industrial manufacturing, healthcare, transportation, logistics, agriculture, disaster response, planetary exploration, and service robotics. Reliable visual perception is essential for enabling robots to recognize objects, understand scenes, localize themselves, and navigate safely in dynamic environments. Although CNN-based vision models have significantly improved perception accuracy, they often struggle to capture long-range dependencies and generalize to complex or unseen environments. Recent advances in Transformer-based vision models address these limitations by employing self-attention mechanisms to learn both local visual features and global contextual relationships. Architectures such as Vision Transformer (ViT), Swin Transformer, DETR, SAM, and Mask2Former have achieved remarkable performance in object detection, semantic segmentation, SLAM, localization, obstacle avoidance, and autonomous navigation. This paper presents a comprehensive review and proposes the Transformer-Based Visual Perception Models for Autonomous Robots (TBVPM-AR) framework. The framework integrates RGB cameras, depth sensors, LiDAR, IMUs, multimodal sensor fusion, transformer-based feature extraction, contextual reasoning, and edge-cloud computing to achieve robust perception in dynamic environments. Mathematical formulations for self-attention, positional encoding, and feature embedding provide the theoretical foundation of the architecture. Experimental evaluations demonstrate that the proposed framework outperforms CNN-based and hybrid approaches on standard robotic perception benchmarks, achieving over 98% visual perception accuracy with improved scene understanding, localization, obstacle detection, navigation, and computational efficiency. The proposed architecture offers a scalable, explainable, and adaptable solution for future Industry 5.0, collaborative robotics, autonomous vehicles, and smart cyber-physical systems.
References
[1] A. Dosovitskiy et al., "An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale," International Conference on Learning Representations (ICLR), 2021.
[2] Z. Liu et al., "Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows," Proc. IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9992–10002, 2021.
[3] N. Carion et al., "End-to-End Object Detection with Transformers," European Conference on Computer Vision (ECCV), pp. 213–229, 2021.
[4] X. Zhu et al., "Deformable DETR: Deformable Transformers for End-to-End Object Detection," International Conference on Learning Representations (ICLR), 2021.
[5] A. Radford et al., "Learning Transferable Visual Models From Natural Language Supervision," Proc. International Conference on Machine Learning (ICML), pp. 8748–8763, 2021.
[6] J. Li, D. Li, C. Xiong, and S. Hoi, "BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation," International Conference on Machine Learning (ICML), 2022.
[7] J. Alayrac et al., "Flamingo: A Visual Language Model for Few-Shot Learning," Advances in Neural Information Processing Systems (NeurIPS), vol. 35, pp. 23716–23736, 2022.
[8] A. Brohan et al., "RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control," Conference on Robot Learning (CoRL), 2023.
[9] R. Szeliski, Computer Vision: Algorithms and Applications, 2nd ed. Cham, Switzerland: Springer, 2022.
[10] Y. LeCun, Y. Bengio, and G. Hinton, "Deep Learning for Vision and Robotics: Recent Advances and Future Trends," IEEE Signal Processing Magazine, vol. 39, no. 6, pp. 84–96, Nov. 2022.
[11] J. Redmon and A. Farhadi, "YOLO-Based Real-Time Object Detection for Autonomous Robotic Systems: Recent Developments," IEEE Access, vol. 10, pp. 61584–61602, 2022.
[12] H. Caesar, V. Bankiti, A. H. Lang, et al., "nuScenes: A Multimodal Dataset for Autonomous Driving," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 3, pp. 2816–2833, Mar. 2023.
[13] L. Fei-Fei, S. Savarese, and J. Malik, "Foundation Models for Embodied Artificial Intelligence," IEEE Intelligent Systems, vol. 39, no. 1, pp. 22–34, Jan.–Feb. 2024.
[14] Y. Wang, X. Chen, H. Li, and Z. Zhang, "Transformer-Based Multi-Modal Sensor Fusion for Autonomous Robot Perception: A Survey," IEEE Access, vol. 12, pp. 45231–45258, 2024.
[15] M. Khan, P. Sharma, R. Gupta, and S. Lee, "Vision Transformer-Based Autonomous Robotic Perception: Recent Advances, Challenges, and Future Directions," IEEE Access, vol. 13, pp. 14562–14591, 2025.
[16] Gajula, S. (2024). Cybersecurity risk prediction using graph neural networks. Journal of Information Systems Engineering and Management.
[17] Gajula, S. (2024). Adaptive zero trust architecture for securing financial microservices. Computer Fraud & Security, 2024(12), 643–655. https://doi.org/10.52710/cfs.845
Downloads
How to Cite
[1]H. N. Mahabala, “Transformer-Based Visual Perception Models for Autonomous Robots”, IJIARE, vol. 8, no. 2, pp. 01–16, Nov. 2025, doi: 10.67228/30715725/IJIARE-2025PII4Q8N.
Most read articles by the same author(s)
- N. Seshagiri, H. N. Mahabala, Hybrid Deep Learning Frameworks for Robotic Object Recognition , International Journal of Intelligent Automation & Robotics Engineering: Vol. 8 No. 1 (2025)
- H. N. Mahabala, Adaptive Route Planning for Autonomous Logistics Robots in Dynamic Warehouse Environments , International Journal of Intelligent Automation & Robotics Engineering: Vol. 6 No. 2 (2023)
Similar Articles
- Mr. Vikram Sethi, Edge Intelligence for Real-Time Industrial Automation Systems , International Journal of Intelligent Automation & Robotics Engineering: Vol. 8 No. 1 (2025)
- Louis Pouzin, Jacques Arsac, Agent-Based Machine Learning Frameworks for Autonomous Predictive Decision Systems , International Journal of Intelligent Automation & Robotics Engineering: Vol. 4 No. 1 (2021)
- Alexey Lyapunov, AI-Based Dynamic Task Allocation in Multi-Robot Systems , International Journal of Intelligent Automation & Robotics Engineering: Vol. 8 No. 1 (2025)
- Michael Rabin, Amir Pnueli, Autonomous Factory Automation through Cyber-Physical Production Systems , International Journal of Intelligent Automation & Robotics Engineering: Vol. 6 No. 2 (2023)
- Ole Secher Olesen, Digital Twin-Enabled Autonomous Robotic Maintenance Frameworks , International Journal of Intelligent Automation & Robotics Engineering: Vol. 7 No. 1 (2024)
- Narendra Karmarkar, Federated Learning Architectures for Distributed Robotic Intelligence , International Journal of Intelligent Automation & Robotics Engineering: Vol. 8 No. 1 (2025)
- Alan Bundy, Karen Spärck Jones, Edge Computing Architectures for Intelligent Embedded Robotic Systems , International Journal of Intelligent Automation & Robotics Engineering: Vol. 7 No. 2 (2024)
- Dr. Linda Martinez, Dr. Mark Richardson, Digital Twin-Based Performance Optimization of Industrial Robots , International Journal of Intelligent Automation & Robotics Engineering: Vol. 4 No. 2 (2021)
- Dr. Priya Natarajan, Dr. Suresh Babu Reddy, AI-Powered Motion Prediction Models for Mobile Robots , International Journal of Intelligent Automation & Robotics Engineering: Vol. 2 No. 1 (2019)
- N. Seshagiri, Adaptive Embedded Control Architectures for Intelligent Mechatronic Systems , International Journal of Intelligent Automation & Robotics Engineering: Vol. 5 No. 2 (2022)
You may also start an advanced similarity search for this article.