Self-Supervised Learning Models for Autonomous Robotic Navigation

  • Authors

    • Donald Michie Professor of Machine Intelligence, University of Edinburgh, United Kingdom. Author
    • Roger Needham Professor, University of Cambridge, United Kingdom. Author

    DOI:

    https://doi.org/10.67228/30715725/IJIARE-2024PII5L3W

    Published 10-05-2024

  • Self-Supervised Learning, Autonomous Robotic Navigation, Deep Learning, Representation Learning, Contrastive Learning, Robot Perception, Reinforcement Learning, Computer Vision, LiDAR, Multi-Modal Learning, Intelligent Robotics

    Issue

    Section

    Articles

    How to Cite

    [1]
    D. Michie and R. Needham, “Self-Supervised Learning Models for Autonomous Robotic Navigation”, IJIARE, vol. 7, no. 2, pp. 01–16, Oct. 2024, doi: 10.67228/30715725/IJIARE-2024PII5L3W.
  • Abstract

    Autonomous robotic navigation has become a key capability for intelligent robots operating in dynamic environments such as warehouses, hospitals, smart cities, agriculture, and autonomous transportation. While supervised learning methods achieve strong navigation performance, they depend on large labeled datasets that are costly and time-consuming to obtain. Self-Supervised Learning (SSL) addresses this limitation by enabling robots to learn robust visual and spatial representations directly from unlabeled sensor data through self-generated learning objectives. This paper reviews recent advances in SSL techniques, including representation learning, contrastive learning, predictive learning, masked image modeling, and multimodal sensor fusion for autonomous navigation. It also examines the integration of data from RGB cameras, LiDAR, IMUs, GPS, depth sensors, and odometry to improve perception, localization, obstacle avoidance, and path planning in unknown environments. Finally, the paper discusses key challenges such as domain adaptation, computational efficiency, safety, and continual learning, highlighting SSL's potential to enable scalable, adaptive, and lifelong autonomous robotic navigation.

  • References

    [1] K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked Autoencoders Are Scalable Vision Learners,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 16000–16009, 2022.

    [2] M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging Properties in Self-Supervised Vision Transformers,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9650–9660, 2021.

    [3] J.-B. Grill, F. Strub, F. Altché, et al., “Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning,” Advances in Neural Information Processing Systems (NeurIPS), vol. 33, pp. 21271–21284, 2021.

    [4] X. Chen, S. Xie, and K. He, “An Empirical Study of Training Self-Supervised Vision Transformers,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9640–9649, 2021.

    [5] A. Dosovitskiy, L. Beyer, A. Kolesnikov, et al., “An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale,” International Conference on Learning Representations (ICLR), 2021.

    [6] Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A ConvNet for the 2020s,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11976–11986, 2022.

    [7] Y. LeCun, I. Misra, and J. Mairal, “Self-Supervised Learning: The Dark Matter of Intelligence,” Nature Reviews Physics, vol. 4, no. 9, pp. 1–15, 2022.

    [8] P. Goyal, M. Mahajan, A. Gupta, and I. Misra, “Scaling and Benchmarking Self-Supervised Visual Representation Learning,” Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pp. 6391–6402, 2021.

    [9] C. Cadena, L. Carlone, H. Carrillo, et al., “Past, Present, and Future of Simultaneous Localization and Mapping: Toward the Robust-Perception Age,” IEEE Transactions on Robotics, vol. 38, no. 3, pp. 1307–1332, 2022.

    [10] M. Bjerkeng, T. C. Thorrud, and T. A. Johansen, “Self-Supervised Visual Representation Learning for Autonomous Robotic Navigation,” IEEE Robotics and Automation Letters, vol. 7, no. 4, pp. 10835–10842, 2022.

    [11] H. Wang, Y. Wang, S. Liu, and X. Zhang, “Self-Supervised Multimodal Representation Learning for Autonomous Robotic Systems,” IEEE Transactions on Industrial Informatics, vol. 19, no. 5, pp. 6123–6134, 2023.

    [12] J. Guo, X. Li, Z. Wang, and Y. Chen, “Vision Transformer-Based Robotic Perception for Autonomous Navigation: A Survey,” IEEE Access, vol. 12, pp. 21345–21370, 2024.

    [13] X. Chen, L. Zhao, Y. Sun, and J. Wu, “Multimodal Self-Supervised Learning for Intelligent Robot Navigation,” IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 2, pp. 1489–1503, 2024.

  • Downloads

Similar Articles

11-20 of 84

You may also start an advanced similarity search for this article.