Autonomous Robotics with Vision-Language AI for Industrial Inspection and Maintenance
-
DOI:
https://doi.org/10.67228/30716357/IJMRSE-2024PII8T1VPublished 09-04-2024
Autonomous Robotics, Vision-Language Ai, Industrial Inspection, Predictive Maintenance, Computer Vision, Intelligent Automation, Industry 4.0, Deep Learning Issue
Section
ArticlesHow to Cite
Autonomous Robotics with Vision-Language AI for Industrial Inspection and Maintenance. (2024). International Journal of Modern Research in Science & Engineering, 7(2), 01-18. https://doi.org/10.67228/30716357/IJMRSE-2024PII8T1VAbstract
Industrial environments are rapidly evolving toward Industry 4.0, where autonomous robots and AI enable intelligent inspection and maintenance with minimal human intervention. Traditional inspection methods rely on manual operations, fixed robotic programming, and isolated computer vision techniques that struggle in dynamic industrial environments. This work proposes FACTS, an autonomous industrial inspection and maintenance framework integrating robotics with vision-language AI. The framework combines multimodal sensing, advanced computer vision models (CNNs, Vision Transformers, and vision-language foundation models), and natural language understanding to interpret equipment conditions, detect defects, reason about maintenance requirements, and execute autonomous corrective actions. Vision-language feature fusion enhances contextual understanding, while mathematical optimization improves decision accuracy and computational efficiency. The proposed system supports applications in manufacturing, power plants, aerospace, and critical infrastructure by enabling defect detection, predictive maintenance, safety monitoring, and remote assistance. Experimental evaluation demonstrates significant improvements in inspection accuracy, fault classification, decision-making, maintenance prediction, reduced downtime, and lower human intervention, providing a foundation for next-generation intelligent industrial automation.
References
[1] M. Prunella, R. Scardigno, D. Buongiorno, and A. Brunetti, “Deep Learning for Automatic Vision-Based Recognition of Industrial Surface Defects: A Survey,” IEEE Access, vol. 11, pp. 43370–43423, 2023, doi: 10.1109/ACCESS.2023.3271748.
[2] S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 6, pp. 1137–1149, 2017, doi: 10.1109/TPAMI.2016.2577031.
[3] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You Only Look Once: Unified, Real-Time Object Detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 779–788, doi: 10.1109/CVPR.2016.91.
[4] K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778, doi: 10.1109/CVPR.2016.90.
[5] G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely Connected Convolutional Networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 4700–4708, doi: 10.1109/CVPR.2017.243.
[6] M. Tan and Q. Le, “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” in Proceedings of the International Conference on Machine Learning (ICML), 2019, pp. 6105–6114.
[7] A. Dosovitskiy et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale,” in International Conference on Learning Representations (ICLR), 2021.
[8] S. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. Shahbaz Khan, and M. Shah, “Transformers in Vision: A Survey,” ACM Computing Surveys, vol. 54, no. 10s, pp. 1–41, 2022, doi: 10.1145/3505244.
[9] J. Zhang, J. Huang, S. Jin, and S. Lu, “Vision-Language Models for Vision Tasks: A Survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 8, pp. 5625–5644, 2024, doi: 10.1109/TPAMI.2024.3369699.
[10] A. Radford et al., “Learning Transferable Visual Models From Natural Language Supervision,” in Proceedings of the International Conference on Machine Learning (ICML), 2021, pp. 8748–8763.
[11] Y. Zeng, X. Zhang, and H. Li, “Multi-Modal Learning for Vision-Language Understanding: A Survey,” IEEE Transactions on Artificial Intelligence, vol. 5, no. 2, pp. 1–18, 2024.
[12] A. Agarwal et al., “Robotic Defect Inspection with Visual and Tactile Perception for Large-Scale Components,” 2023.
[13] A. Kirillov et al., “Segment Anything,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023.
[14] B. Zitkovic, “Artificial Intelligence-Based Robotic Inspection Systems for Industry 4.0 Manufacturing Environments,” IEEE Access, vol. 12, pp. 1–15, 2024.
[15] Taluri, R. (2024). A Cloud-Native Reference Architecture for Data Engineering, Generative AI, and Decision Intelligence Using AWS and Amazon Bedrock. International Journal of AI, BigData, Computational and Management Studies, 5(1), 218-227. https://doi.org/10.63282/3050-9416.IJAIBDCMS-V5I1P122
Downloads
How to Cite
Autonomous Robotics with Vision-Language AI for Industrial Inspection and Maintenance. (2024). International Journal of Modern Research in Science & Engineering, 7(2), 01-18. https://doi.org/10.67228/30716357/IJMRSE-2024PII8T1V