Transformer-Based Visual Perception Models for Autonomous Robots
-
DOI:
https://doi.org/10.67228/30715725/IJIARE-2025PII4Q8NPublished 11-03-2025
Autonomous Robots, Visual Perception, Vision Transformer (Vit), Self-Attention Mechanism, Deep Learning, Object Detection, Semantic Segmentation, Multi-Modal Sensor Fusion, Intelligent Robotics, Artificial Intelligence, Computer Vision, Scene Understanding Issue
Section
ArticlesHow to Cite
[1]H. N. Mahabala, “Transformer-Based Visual Perception Models for Autonomous Robots”, IJIARE, vol. 8, no. 2, pp. 01–16, Nov. 2025, doi: 10.67228/30715725/IJIARE-2025PII4Q8N.Abstract
Autonomous robots play a crucial role in industrial manufacturing, healthcare, transportation, logistics, agriculture, disaster response, planetary exploration, and service robotics. Reliable visual perception is essential for enabling robots to recognize objects, understand scenes, localize themselves, and navigate safely in dynamic environments. Although CNN-based vision models have significantly improved perception accuracy, they often struggle to capture long-range dependencies and generalize to complex or unseen environments. Recent advances in Transformer-based vision models address these limitations by employing self-attention mechanisms to learn both local visual features and global contextual relationships. Architectures such as Vision Transformer (ViT), Swin Transformer, DETR, SAM, and Mask2Former have achieved remarkable performance in object detection, semantic segmentation, SLAM, localization, obstacle avoidance, and autonomous navigation. This paper presents a comprehensive review and proposes the Transformer-Based Visual Perception Models for Autonomous Robots (TBVPM-AR) framework. The framework integrates RGB cameras, depth sensors, LiDAR, IMUs, multimodal sensor fusion, transformer-based feature extraction, contextual reasoning, and edge-cloud computing to achieve robust perception in dynamic environments. Mathematical formulations for self-attention, positional encoding, and feature embedding provide the theoretical foundation of the architecture. Experimental evaluations demonstrate that the proposed framework outperforms CNN-based and hybrid approaches on standard robotic perception benchmarks, achieving over 98% visual perception accuracy with improved scene understanding, localization, obstacle detection, navigation, and computational efficiency. The proposed architecture offers a scalable, explainable, and adaptable solution for future Industry 5.0, collaborative robotics, autonomous vehicles, and smart cyber-physical systems.
References
[1] A. Dosovitskiy et al., "An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale," International Conference on Learning Representations (ICLR), 2021.
[2] Z. Liu et al., "Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows," Proc. IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9992–10002, 2021.
[3] N. Carion et al., "End-to-End Object Detection with Transformers," European Conference on Computer Vision (ECCV), pp. 213–229, 2021.
[4] X. Zhu et al., "Deformable DETR: Deformable Transformers for End-to-End Object Detection," International Conference on Learning Representations (ICLR), 2021.
[5] A. Radford et al., "Learning Transferable Visual Models From Natural Language Supervision," Proc. International Conference on Machine Learning (ICML), pp. 8748–8763, 2021.
[6] J. Li, D. Li, C. Xiong, and S. Hoi, "BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation," International Conference on Machine Learning (ICML), 2022.
[7] J. Alayrac et al., "Flamingo: A Visual Language Model for Few-Shot Learning," Advances in Neural Information Processing Systems (NeurIPS), vol. 35, pp. 23716–23736, 2022.
[8] A. Brohan et al., "RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control," Conference on Robot Learning (CoRL), 2023.
[9] R. Szeliski, Computer Vision: Algorithms and Applications, 2nd ed. Cham, Switzerland: Springer, 2022.
[10] Y. LeCun, Y. Bengio, and G. Hinton, "Deep Learning for Vision and Robotics: Recent Advances and Future Trends," IEEE Signal Processing Magazine, vol. 39, no. 6, pp. 84–96, Nov. 2022.
[11] J. Redmon and A. Farhadi, "YOLO-Based Real-Time Object Detection for Autonomous Robotic Systems: Recent Developments," IEEE Access, vol. 10, pp. 61584–61602, 2022.
[12] H. Caesar, V. Bankiti, A. H. Lang, et al., "nuScenes: A Multimodal Dataset for Autonomous Driving," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 3, pp. 2816–2833, Mar. 2023.
[13] L. Fei-Fei, S. Savarese, and J. Malik, "Foundation Models for Embodied Artificial Intelligence," IEEE Intelligent Systems, vol. 39, no. 1, pp. 22–34, Jan.–Feb. 2024.
[14] Y. Wang, X. Chen, H. Li, and Z. Zhang, "Transformer-Based Multi-Modal Sensor Fusion for Autonomous Robot Perception: A Survey," IEEE Access, vol. 12, pp. 45231–45258, 2024.
[15] M. Khan, P. Sharma, R. Gupta, and S. Lee, "Vision Transformer-Based Autonomous Robotic Perception: Recent Advances, Challenges, and Future Directions," IEEE Access, vol. 13, pp. 14562–14591, 2025.
[16] Gajula, S. (2024). Cybersecurity risk prediction using graph neural networks. Journal of Information Systems Engineering and Management.
[17] Gajula, S. (2024). Adaptive zero trust architecture for securing financial microservices. Computer Fraud & Security, 2024(12), 643–655. https://doi.org/10.52710/cfs.845
Downloads
How to Cite
[1]H. N. Mahabala, “Transformer-Based Visual Perception Models for Autonomous Robots”, IJIARE, vol. 8, no. 2, pp. 01–16, Nov. 2025, doi: 10.67228/30715725/IJIARE-2025PII4Q8N.
Most read articles by the same author(s)
- N. Seshagiri, H. N. Mahabala, Hybrid Deep Learning Frameworks for Robotic Object Recognition , International Journal of Intelligent Automation & Robotics Engineering: Vol. 8 No. 1 (2025)
- H. N. Mahabala, Adaptive Route Planning for Autonomous Logistics Robots in Dynamic Warehouse Environments , International Journal of Intelligent Automation & Robotics Engineering: Vol. 6 No. 2 (2023)
Similar Articles
- H. N. Mahabala, Adaptive Route Planning for Autonomous Logistics Robots in Dynamic Warehouse Environments , International Journal of Intelligent Automation & Robotics Engineering: Vol. 6 No. 2 (2023)
- Dr. Martinez Finigan, Development of a Cost-Efficient 6-DoF Service Robot , International Journal of Intelligent Automation & Robotics Engineering: Vol. 9 No. 1 (2026)
- Samuel O’ Brell, Kelvin Ling, Human–Robot Shared Workspace Safety Enhancement Using Predictive Control , International Journal of Intelligent Automation & Robotics Engineering: Vol. 9 No. 1 (2026)
- Dr. Rajesh Kumar Sharma, A Reconfigurable Industrial Robot Architecture for Smart Manufacturing , International Journal of Intelligent Automation & Robotics Engineering: Vol. 2 No. 1 (2019)
- David Thompson, Robotic Process Optimization Using Evolutionary Algorithms , International Journal of Intelligent Automation & Robotics Engineering: Vol. 3 No. 1 (2020)
- Dr. Meena Krishnan, Dr. Arvind Kumar Singh, AI-Based Embedded Controllers for Precision Motion Systems , International Journal of Intelligent Automation & Robotics Engineering: Vol. 5 No. 1 (2022)
- Dr. Suresh Babu Reddy, Dr. Anita Verma, Distributed Intelligence Frameworks for Cooperative Mobile Robotics , International Journal of Intelligent Automation & Robotics Engineering: Vol. 7 No. 1 (2024)
- Niklaus Wirth, AI-Enhanced Process Optimization in Automated Production Systems , International Journal of Intelligent Automation & Robotics Engineering: Vol. 7 No. 1 (2024)
- Mr. Marco Bianchi, Ms. Laura Conti, Digital Twin-Based Predictive Control for Intelligent Manufacturing , International Journal of Intelligent Automation & Robotics Engineering: Vol. 5 No. 1 (2022)
- Dr. K. Balasubramanian, Dr. Meena Krishnan, Explainable Deep Reinforcement Learning for Dynamic Credit Limit Adjustment , International Journal of Intelligent Automation & Robotics Engineering: Vol. 1 No. 1 (2018)
You may also start an advanced similarity search for this article.