Deep Learning-driven Computer Vision for Reliable Autonomous Driving: Perception, Sensor Fusion, and Future Directions

Authors

  • Suraj R. Nalawade
  • Tapase H. O.
  • Madhura Yadav

Keywords:

3D perception, Autonomous driving, Computer vision, Deep learning, Edge AI, Object detection, Sensor fusion

Abstract

Computer vision is a central perception technology for autonomous driving because it converts visual observations of the road environment into information that can be used for navigation and decision support. A vehicle must recognize road users, estimate their locations, identify lanes and traffic signs, interpret traffic lights, and understand the surrounding scene while operating under strict latency and reliability constraints. This study presents a structured review of deep learning-driven computer vision techniques for autonomous driving, with emphasis on object detection, lane and road-boundary detection, semantic and instance segmentation, traffic-sign recognition, tracking, depth estimation, and three-dimensional perception. It also examines the role of convolutional neural networks, one-stage and two-stage object detectors, point-cloud learning, and sensor-fusion strategies that combine cameras, LiDAR, and radar. Rather than treating a perception model as an isolated classifier, the study considers the complete perception pipeline, including data preparation, inference, temporal tracking, uncertainty, computational deployment, and validation. The results synthesize the capabilities and limitations of the reviewed approaches and show why accuracy alone is insufficient for safety-critical autonomous driving. Particular attention is given to adverse weather, illumination changes, occlusion, dataset bias, domain shift, computational cost, and the need for interpretable and verifiable outputs. Future perspectives include edge AI, transformer-based perception, improved 3D scene understanding, simulation-based testing, explainable AI, continual adaptation, and multimodal fusion. The review concludes that dependable autonomous driving requires an integrated perception stack in which algorithmic accuracy, real-time execution, robustness, and system-level validation are considered together.

References

Y. Lecun, L. Bottou, Y. Bengio and P. Haffner, “Gradient-based learning applied to document recognition,” in Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, Nov. 1998.

A. Krizhevsky, I. Sutskever, G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” Communications of the ACM, May 2017, pp. 84–90.

K. He, X. Zhang, S. Ren and J. Sun, “Deep residual learning for image recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, pp. 770–778.

S. Ren, K. He, R. Girshick and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 6, pp. 1137–1149, Jun. 2017.

W. Liu et al., “SSD: Single shot multibox detector,” in European Conference on Computer Vision, Sept. 2017, pp. 21–37.

J. Redmon, S. Divvala, R. Girshick and A. Farhadi, “You only look once: Unified, real-time object detection,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, pp. 779–788.

K. He, G. Gkioxari, P. Dollár and R. Girshick, “Mask R-CNN,” 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 2017, pp. 2980–2988.

R. Q. Charles, H. Su, M. Kaichun and L. J. Guibas, “PointNet: Deep learning on point sets for 3D classification and segmentation,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017, pp. 77–85.

Y. Zhou and O. Tuzel, “VoxelNet: End-to-End learning for point cloud based 3D object detection,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018, pp. 4490–4499.

A. Geiger, P. Lenz and R. Urtasun, “Are we ready for autonomous driving? The KITTI vision benchmark suite,” 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA, 2012, pp. 3354–3361.

M. Cordts et al., “The cityscapes dataset for semantic urban scene understanding,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, pp. 3213–3223.

L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in European Conference on Computer Vision, 2018, pp. 833–351.

A. Bochkovskiy, C.-Y. Wang, H.-Y. M. Liao, “Yolov4: Optimal speed and accuracy of object detection,” arXiv, Apr. 2020.

E. Arnold, O. Y. Al-Jarrah, M. Dianati, S. Fallah, D. Oxtoby and A. Mouzakitis, “A survey on 3D object detection methods for autonomous driving applications,” in IEEE Transactions on Intelligent Transportation Systems, vol. 20, no. 10, pp. 3782–3795, Oct. 2019.

M. Bojarski et al., “End to end learning for self-driving cars,” arXiv, Apr. 2016.

R. Cipolla, Y. Gal and A. Kendall, “Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018, pp. 7482–7491.

A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “CARLA: An open urban driving simulator,” in Conference on Robot Learning, 2017, pp. 1–16.

S. Grigorescu, B. Trasnea, T. Cocias, and G. Macesanu, “A survey of deep learning techniques for autonomous driving,” Journal of Field Robotics, vol. 37, no. 3, pp. 362–386, Nov. 2019.

Z. Li et al., "BEVFormer: Learning Bird’s-Eye-view representation from LiDAR-Camera via spatiotemporal transformers,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 3, pp. 2020–2036, Mar. 2025.

Y. Liu, T. Wang, X. Zhang, and J. Sun, “PETR: Position embedding transformation for multi-view 3D object detection,” in European Conference on Computer Vision, 2022, pp. 531–548.

X. Bai et al., “TransFusion: Robust LiDAR-camera fusion for 3D object detection with transformers,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 2022, pp. 1080–1089.

Y. Li et al., “Bevdepth: Acquisition of reliable depth for multi-view 3d object detection,” in Proceedings of the AAAI Conference on Artificial Intelligence, 2023, pp. 1477–1485.

Y. Hu et al., “Planning-oriented autonomous driving,” 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 2023, pp. 17853–17862.

G. K. Erabati and H. Araujo, “Li3DeTr: A LiDAR based 3D detection transformer,” 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 2023, pp. 4239–4248.

Z. Li, S. Lan, J. M. Alvarez and Z. Wu, “BEVNeXt: Reviving dense BEV frameworks for 3D object detection,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2024, pp. 20113–20123.

X. Li, B. Fan, J. Tian and H. Fan, “GAFusion: Adaptive fusing LiDAR and camera with multiple guidance for 3D object detection,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2024, pp. 21209–21218.

Z. Lin et al., “RCBEVDet: Radar-camera fusion in Bird's Eye View for 3D object detection,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2024, pp. 14928–14937.

H. Yang et al., “UniPAD: A universal pre-training paradigm for autonomous driving,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2024, pp. 15238–15250.

C. Pan, B. Yaman, S. Velipasalar and L. Ren, “CLIP-BEVFormer: Enhancing multi-view image-based BEV detector with ground truth flow,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2024, pp. 15216–15225.

Published

2026-09-17

Issue

Section

Articles