General and robust voxel feature learning with Transformer for 3D object detection 被引量：1

基于Transformer的通用和鲁棒体素特征学习的目标检测

下载PDF

导出

摘要 The self-attention networks and Transformer have dominated machine translation and natural language processing fields,and shown great potential in image vision tasks such as image classification and object detection.Inspired by the great progress of Transformer,we propose a novel general and robust voxel feature encoder for 3D object detection based on the traditional Transformer.We first investigate the permutation invariance of sequence data of the self-attention and apply it to point cloud processing.Then we construct a voxel feature layer based on the self-attention to adaptively learn local and robust context of a voxel according to the spatial relationship and context information exchanging between all points within the voxel.Lastly,we construct a general voxel feature learning framework with the voxel feature layer as the core for 3D object detection.The voxel feature with Transformer(VFT)can be plugged into any other voxel-based 3D object detection framework easily,and serves as the backbone for voxel feature extractor.Experiments results on the KITTI dataset demonstrate that our method achieves the state-of-the-art performance on 3D object detection. 自注意力网络和Transformer主导了机器翻译和自然语言处理领域,并在诸如图像分类和目标检测等图像视觉任务中显示出巨大潜力。受到Transformer在2D图像视觉任务中取得的巨大进步的启发,提出了一种基于传统Transformer的新颖和鲁棒的体素特征编码器。首先,探究自注意力对序列数据的排列不变性,并将其应用于点云数据处理。其次,基于自注意力构造体素特征层,根据体素内所有点之间的空间关系和上下文信息交换自适应地学习体素的局部和鲁棒上下文。最后,构建了以体素特征层为核心的通用3D目标检测框架。VFT(voxel feature learning with Transformer)是通用的体素特征提取器,可以嵌入任何其他基于体素方法的3D物体检测框架中。在KITTI数据集上进行的实验结果表明,本方法在3D目标检测方面表现出优越的性能。

作者 LI Yang GE Hongwei 李阳;葛洪伟(江南大学江苏省模式识别与计算智能实验室,江苏无锡214122;江南大学人工智能与计算机学院,江苏无锡214122)

机构地区 Jiangsu Provincial Engineering Laboratory of Pattern Recognition and Computational Intelligence School of Artificial Intelligence and Computer Science

出处《Journal of Measurement Science and Instrumentation》 CAS CSCD 2022年第1期51-60,共10页 测试科学与仪器(英文版)

基金 National Natural Science Foundation of China(No.61806006) Innovation Program for Graduate of Jiangsu Province(No.KYLX160-781) University Superior Discipline Construction Project of Jiangsu Province。

关键词 3D object detection self-attention networks voxel feature with Transformer(VFT) point cloud encoder-decoder 3D目标检测自注意力网络基于Transformer的体素特征学习点云编码解码器

分类号 TP3 [自动化与计算机技术—计算机科学与技术]