点云
人工智能
计算机科学
计算机视觉
目标检测
RGB颜色模型
激光雷达
分割
水准点(测量)
特征(语言学)
模式识别(心理学)
遥感
地理
大地测量学
语言学
哲学
作者
Qingdong He,Zhengning Wang,Hao Zeng,Yi Zeng,Shuaicheng Liu,Shuaicheng Liu,Bing Zeng
出处
期刊:IEEE Transactions on Intelligent Transportation Systems
[Institute of Electrical and Electronics Engineers]
日期:2023-01-01
卷期号:24 (1): 152-162
被引量:7
标识
DOI:10.1109/tits.2022.3215766
摘要
3D object detection has become an emerging task in autonomous driving scenarios. Most of previous works process 3D point clouds using either projection-based or voxel-based models. However, both approaches contain some drawbacks. The voxel-based methods lack semantic information, while the projection-based methods suffer from numerous spatial information loss when projected to different views. In this paper, we propose the Stereo RGB and Deeper LIDAR (SRDL) framework which can utilize semantic and spatial information simultaneously such that the performance of network for 3D object detection can be improved naturally. Specifically, the network generates candidate boxes from stereo pairs and combines different region-wise features using a deep fusion scheme. The stereo strategy offers more information for prediction compared with prior works. Then, several local and global feature extractors are stacked in the segmentation module to capture richer deep semantic geometric features from point clouds. After aligning the interior points with fused features, the proposed network refines the prediction in a more accurate manner and encodes the whole box in a novel compact method. The decent experimental results on the challenging KITTI detection benchmark demonstrate the effectiveness of utilizing both stereo images and point clouds for 3D object detection.
科研通智能强力驱动
Strongly Powered by AbleSci AI