CONFIDENCE-WEIGHTED MID-LEVEL LIDAR–STEREO FUSION WITH HYBRID PCA-K-MEANS CLUSTERING FOR MARITIME 3D OBJECT DETECTION
DOI:
https://doi.org/10.35631/JISTM.1144016Keywords:
Unmanned Surface Vessel, LiDAR–Stereo Fusion, Mid-Level Sensor Fusion, Principal Component Analysis, K-Means Clustering, Maritime 3D Object Detection, Guidance, Navigation and ControlAbstract
Autonomous Unmanned Surface Vessels (USVs) operating in nearshore environments require robust perception systems capable of handling challenging maritime conditions. However, LiDAR and stereo vision each suffer from complementary limitations: LiDAR produces sparse and reflection-sensitive measurements, while stereo vision provides dense depth estimates but performs poorly under occlusion and weak-texture conditions. While early- and late-fusion approaches have been widely explored, mid-level fusion strategies that preserve cross-modal feature interactions while adapting to sensor uncertainty remain underdeveloped for maritime perception. This paper introduces M-ACF Net (Maritime Adaptive Confidence Fusion Network), a confidence-adaptive mid-level fusion framework that dynamically integrates sparse LiDAR and dense stereo features based on pixel-level reliability. The fused representation is processed using PCA-K-Means clustering to generate orientation-invariant obstacle representations for Guidance, Navigation, and Control (GNC). The framework was implemented on an instrumented USV platform and evaluated on a multimodal dataset of 17,634 synchronised LiDAR–stereo–GNSS samples collected across three nearshore Malaysian sites under wave heights of 0.1–0.5 m. Across different wave conditions, M-ACF Net reduced depth-estimation RMSE by 44–60% compared with LiDAR-only measurements. The framework achieved 82.6% 3D Average Precision for ship detection under easy conditions and outperformed YOLOv8-Stereo, PointPillars, Pseudo-LiDAR, Frustum PointNet, and DSGN baselines, while matching DBSCAN's clustering quality at less than half its computation time and sustaining 18 frames per second on an embedded NVIDIA Jetson Orin NX. Field trials further validated real-time obstacle detection under operational conditions. These results demonstrate the effectiveness of confidence-adaptive mid-level fusion for real-time maritime obstacle perception.
Downloads
References
Aijazi, A., Laurent, M., Trassoudaine, L., & Checchin, P. (2020). Systematic evaluation and characterization of 3D solid state LiDAR sensors for autonomous ground vehicles. International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, XLIII-B1-2020, 199–203.
Balta, H., Velagic, J., Bosschaerts, W., De Cubber, G., & Siciliano, B. (2018). Fast statistical outlier removal based method for large 3D point clouds of outdoor environments. IFAC-PapersOnLine, 51(22), 348–353.
Benjamin, M. R., Schmidt, H., Newman, P. M., & Leonard, J. J. (2010). Nested autonomy for unmanned marine vehicles with MOOS-IvP. Journal of Field Robotics, 27(6), 834–875.
Bovcon, B., Muhovič, J., Vranac, D., Mozetič, D., Perš, J., & Kristan, M. (2021). MODS — A USV-oriented object detection and obstacle segmentation benchmark. IEEE Transactions on Intelligent Transportation Systems, 23(8), 13403.
Chen, X., Ma, H., Wan, J., Li, B., & Xia, T. (2017). Multi-view 3D object detection network for autonomous driving. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
Chen, Y., Liu, S., Shen, X., & Jia, J. (2020). DSGN: Deep stereo geometry network for 3D object detection. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12536–12545.
Chen, Y.-T., Shi, J., Ye, Z., Mertz, C., Kong, S., & Ramanan, D. (2021). Multimodal object detection via probabilistic ensembling. arXiv preprint.
Damodaran, B., Nair, B. B., & Akhilraj, N. (2023). Challenges of LiDAR and camera fusion for water surface detection in autonomous maritime navigation. Ocean Engineering, 284, 115263.
Er, M. J., Chen, J., Zhang, Y., & Gao, W. (2023). Research challenges, recent advances, and popular datasets in deep learning-based underwater marine object detection: A review. Sensors, 23(4), 1990.
Ester, M., Kriegel, H.-P., Sander, J., & Xu, X. (1996). A density-based algorithm for discovering clusters in large spatial databases with noise. Proceedings of the 2nd International Conference on Knowledge Discovery and Data Mining (KDD-96), 226–231.
Feng, Q., Cheng, S., & Meng, Y. (2020). A survey on autonomous surface vehicles: Control, perception, navigation, and safety. IEEE Access, 8, 221208–221233.
Hacinecipoglu, A. (2022). Point cloud-based human pose estimation using voxel grid downsampling. Applied Sciences, 12(9), 4559.
Irfan, A., Li, Y., Xinhua, E., & Sun, G. (2025). Land use and land cover classification with deep learning-based fusion of SAR and optical data. Remote Sensing, 17(7), 1298.
Jolliffe, I. T., & Cadima, J. (2016). Principal component analysis: A review and recent developments. Philosophical Transactions of the Royal Society A, 374(2065), 20150202. https://doi.org/10.1098/rsta.2015.0202
Ku, J., Mozifian, M., Lee, J., Harakeh, A., & Waslander, S. (2018). Joint 3D proposal generation and object detection from view aggregation. Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS).
Lang, A., Vora, S., Caesar, H., Zhou, L., Yang, J., & Beijbom, O. (2019). PointPillars: Fast encoders for object detection from point clouds. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
Liu, C., Xiong, S., Geng, Y., Cheng, S., Hu, F., Shao, B., Li, F., & Zhang, J. (2023). An embedded high-precision GNSS-visual-inertial multi-sensor fusion suite. NAVIGATION: Journal of the Institute of Navigation, 70(4).
Liu, Z., Zhang, Y., Yu, X., & Yuan, C. (2016). Unmanned surface vehicles: An overview of developments and challenges. Annual Reviews in Control, 41, 71–93.
MacQueen, J. (1967). Some methods for classification and analysis of multivariate observations. Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, 1, 281–297.
Norazaruddin, M. A. (2026). A mid-level LiDAR-stereo fusion framework for unmanned surface vessel geospatial awareness in near-shore environments [Unpublished doctoral thesis]. International Islamic University Malaysia.
Qi, C. R., Liu, W., Wu, C., Su, H., & Guibas, L. J. (2018). Frustum PointNets for 3D object detection from RGB-D data. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 918–927.
Ren, S., He, K., Girshick, R., & Sun, J. (2015). Faster R-CNN: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems, 28, 91–99.
Rousseeuw, P. J. (1987). Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20, 53–65.
Sindagi, V. A., Zhou, Y., & Tuzel, O. (2019). MVX-Net: Multimodal VoxelNet for 3D object detection. Proceedings of the International Conference on Robotics and Automation (ICRA).
Vora, S., Lang, A., Helou, B., & Beijbom, O. (2020). PointPainting: Sequential fusion for 3D object detection. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4603.
Weng, X., & Kitani, K. (2019). Monocular 3D object detection with pseudo-LiDAR point cloud. Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW).
Yan, Y., Mao, Y., & Li, B. (2018). SECOND: Sparsely embedded convolutional detection. Sensors, 18(10), 3337.
Zhang, R., Li, S., Ji, G., Zhao, X., Li, J., & Pan, M. (2021). Survey on deep learning-based marine object detection. Journal of Advanced Transportation, 2021, 1–18.
Zhou, B., He, Y., Qian, K., Ma, X., & Li, X. (2021). S4-SLAM: A real-time 3D LiDAR SLAM system for ground/water surface multi-scene outdoor applications. Autonomous Robots, 45(1), 1–18.
Zhou, Y., & Tuzel, O. (2018). VoxelNet: End-to-end learning for point cloud based 3D object detection. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
