MAELi——基于遮蔽自编码器的大规模LiDAR点云特征学习 (MAELi $\unicode{x2013}$ Masked Autoencoder for Large-Scale LiDAR Point Clouds) - 专知论文

会员服务 ·

0

LIDAR · 点云 · 掩码 · 自编码器 · 3D ·

2023 年 3 月 17 日

MAELi $\unicode{x2013}$ Masked Autoencoder for Large-Scale LiDAR Point Clouds

翻译：MAELi——基于遮蔽自编码器的大规模LiDAR点云特征学习

Georg Krispel,David Schinagl,Christian Fruhwirth-Reisinger,Horst Possegger,Horst Bischof

from arxiv, 17 pages

We demonstrate how the often overlooked inherent properties of large-scale LiDAR point clouds can be effectively utilized for self-supervised representation learning. In pursuit of this goal, we design a highly data-efficient feature pre-training backbone that considerably reduces the need for tedious 3D annotations to train state-of-the-art object detectors. We propose Masked AutoEncoder for LiDAR point clouds (MAELi) that intuitively leverages the sparsity of LiDAR point clouds in both the encoder and decoder during reconstruction. Our approach results in more expressive and useful features, which can be directly applied to downstream perception tasks, such as 3D object detection for autonomous driving. In a novel reconstruction schema, MAELi distinguishes between free and occluded space and employs a new masking strategy that targets the LiDAR's inherent spherical projection. To demonstrate the potential of MAELi, we pre-train one of the most widely-used 3D backbones in an end-to-end manner and show the effectiveness of our unsupervised pre-trained features on various 3D object detection architectures. Our method achieves significant performance improvements when only a small fraction of labeled frames is available for fine-tuning object detectors. For instance, with ~800 labeled frames, MAELi features enhance a SECOND model by +10.79APH/LEVEL 2 on Waymo Vehicles.

翻译：我们展示了如何有效利用大规模LiDAR点云的常被忽略的内在特性进行自监督特征学习。为了实现这个目标，我们设计了一个高效的数据预训练骨干网络，大大减少了训练最先进的目标检测器所需的琐碎三维标注。我们提出了适用于LiDAR点云的遮蔽自编码器（MAELi），在编码器和解码器在重构过程中直观地利用了LiDAR点云的稀疏性。我们的方法产生了更具表现力和实用性的特征，可以直接应用于下游感知任务，例如自动驾驶中的三维物体检测。在一个新颖的重构模式中，MAELi区分了自由和遮挡空间，并采用了针对LiDAR固有的球面投影的新的掩蔽策略。为了证明MAELi的潜力，我们端到端地预训练了其中最广泛使用的三维骨干网络，展示了我们的无监督预训练特征在各种三维物体检测架构上的有效性。当只有少量标记帧可用于微调目标检测器时，我们的方法实现了显著的性能提升。例如，用约800个标记帧，MAELi特征在 Waymo Vehicles 上将一个SECOND模型提高了 +10.79APH/LEVEL 2。

0

相关内容

LIDAR

【CVPR2023】Mask3D:通过学习掩码3D先验对2D视觉transformer进行预训练

【CVPR2023】Mask3D:通过学习掩码3D先验对2D视觉transformer进行预训练

专知会员服务

24+阅读 · 2023年4月9日

【CVPR2022】CAT-Det:用于多模态三维物体检测的对比增强Transformer

【CVPR2022】CAT-Det:用于多模态三维物体检测的对比增强Transformer

专知会员服务

19+阅读 · 2022年4月7日

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

专知会员服务

46+阅读 · 2021年11月24日

【CVPR2021】空间一致性表示学习

专知会员服务

63+阅读 · 2021年3月12日

【CVPR2021】用Transformers无监督预训练进行目标检测

【CVPR2021】用Transformers无监督预训练进行目标检测

专知会员服务

58+阅读 · 2021年3月3日

【CVPR2020-中科院计算所】弱监督语义分割的自监督等价注意力机制，Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

【CVPR2020-中科院计算所】弱监督语义分割的自监督等价注意力机制，Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

专知会员服务

76+阅读 · 2020年4月10日

基于动态时空图CNNs的交通流预测，Dynamic Spatio-temporal Graph-based CNNs for Traffic Flow Prediction

基于动态时空图CNNs的交通流预测，Dynamic Spatio-temporal Graph-based CNNs for Traffic Flow Prediction

专知会员服务

136+阅读 · 2020年3月8日

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

专知会员服务

28+阅读 · 2020年2月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

VideoMAE：简单高效的视频自监督预训练新范式｜NeurIPS 2022

VideoMAE：简单高效的视频自监督预训练新范式｜NeurIPS 2022

新智元

0+阅读 · 2022年11月28日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【泡泡一分钟】用于RGBD语义分割的三维图神经网络(ICCV2017-546)

【泡泡一分钟】用于RGBD语义分割的三维图神经网络(ICCV2017-546)

泡泡机器人SLAM

22+阅读 · 2018年12月4日

【泡泡一分钟】学习紧密的几何特征（ICCV2017-17）

【泡泡一分钟】学习紧密的几何特征（ICCV2017-17）

泡泡机器人SLAM

20+阅读 · 2018年5月8日

【论文推荐】最新七篇图像检索相关论文—草图、Tie-Aware、场景图解析、叠加跨注意力机制、深度哈希、人群估计

【论文推荐】最新七篇图像检索相关论文—草图、Tie-Aware、场景图解析、叠加跨注意力机制、深度哈希、人群估计

专知

10+阅读 · 2018年4月22日

【泡泡一分钟】基于多视图卷积网络的草图三维重建技术(3dv-66)

【泡泡一分钟】基于多视图卷积网络的草图三维重建技术(3dv-66)

泡泡机器人SLAM

11+阅读 · 2018年3月31日

【泡泡一分钟】将3D全卷积网络应用于车辆激光点云处理（IROS-11）

【泡泡一分钟】将3D全卷积网络应用于车辆激光点云处理（IROS-11）

泡泡机器人SLAM

13+阅读 · 2018年3月23日

【泡泡一分钟】Matterport3D: 从室内RGBD数据集中训练 (3dv-22)

【泡泡一分钟】Matterport3D: 从室内RGBD数据集中训练 (3dv-22)

泡泡机器人SLAM

16+阅读 · 2017年12月31日

基于深度学习的金丝猴面部特性的检测与识别算法研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于张量投票的车载LiDAR数据的目标识别

国家自然科学基金

1+阅读 · 2015年12月31日

基于非易失内存设备的数据读写性能优化方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于2-D空间离散数据的质量与产出的预测方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于局部纹理特征的图像细节超分辨率技术研究

国家自然科学基金

1+阅读 · 2013年12月31日

碳材料改性的金/聚合物催化剂及其低温氧化醇的性能研究

国家自然科学基金

0+阅读 · 2013年12月31日

云环境中基于三维世界模型的图像表示与压缩

国家自然科学基金

2+阅读 · 2013年12月31日

基于复杂网络的中文文本语义相似度研究

国家自然科学基金

3+阅读 · 2012年12月31日

极化合成孔径雷达(SAR)图像地物并行分割分类研究与应用

国家自然科学基金

1+阅读 · 2012年12月31日

基于图像空间视觉相似性的质量评价方法

国家自然科学基金

0+阅读 · 2008年12月31日

Self-supervised dense representation learning for live-cell microscopy with time arrow prediction

Arxiv

0+阅读 · 2023年5月9日

Self-supervised Learning for Pre-Training 3D Point Clouds: A Survey

Arxiv

5+阅读 · 2023年5月8日

Equiangular Basis Vectors

Arxiv

0+阅读 · 2023年5月8日

Capacity Achieving Codes for an Erasure Queue-Channel

Arxiv

0+阅读 · 2023年5月7日

A vector quantized masked autoencoder for audiovisual speech emotion recognition

Arxiv

0+阅读 · 2023年5月5日

Contrastive Learning for Low-light Raw Denoising

Arxiv

0+阅读 · 2023年5月5日

Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks

Arxiv

0+阅读 · 2023年5月4日

Federated Ensemble-Directed Offline Reinforcement Learning

Arxiv

0+阅读 · 2023年5月4日

Masked Autoencoders Are Scalable Vision Learners

Arxiv

27+阅读 · 2021年11月11日

SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition

Arxiv

12+阅读 · 2021年5月30日

VIP会员

文章信息

相关主题

相关VIP内容

【CVPR2023】Mask3D:通过学习掩码3D先验对2D视觉transformer进行预训练

【CVPR2023】Mask3D:通过学习掩码3D先验对2D视觉transformer进行预训练

专知会员服务

24+阅读 · 2023年4月9日

【CVPR2022】CAT-Det:用于多模态三维物体检测的对比增强Transformer

【CVPR2022】CAT-Det:用于多模态三维物体检测的对比增强Transformer

专知会员服务

19+阅读 · 2022年4月7日

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

【NeurIPS2021】用于文本图表示学习的 GNN 嵌套 Transformer 模型：GraphFormers

专知会员服务

46+阅读 · 2021年11月24日

【CVPR2021】空间一致性表示学习

专知会员服务

63+阅读 · 2021年3月12日

【CVPR2021】用Transformers无监督预训练进行目标检测

【CVPR2021】用Transformers无监督预训练进行目标检测

专知会员服务

58+阅读 · 2021年3月3日

【CVPR2020-中科院计算所】弱监督语义分割的自监督等价注意力机制，Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

【CVPR2020-中科院计算所】弱监督语义分割的自监督等价注意力机制，Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

专知会员服务

76+阅读 · 2020年4月10日

基于动态时空图CNNs的交通流预测，Dynamic Spatio-temporal Graph-based CNNs for Traffic Flow Prediction

基于动态时空图CNNs的交通流预测，Dynamic Spatio-temporal Graph-based CNNs for Traffic Flow Prediction

专知会员服务

136+阅读 · 2020年3月8日

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

专知会员服务

28+阅读 · 2020年2月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

热门VIP内容

开通专知VIP会员享更多权益服务

【NTU博士论文】深度神经网络的参数高效推理与训练

人工智能：实时战斗适应

【NeurIPS2025】MIDAS：一种基于错配的用于失衡多模态学习的数据增强策略

从感知到认知：多模态大语言模型中视觉-语言交互推理综述

相关资讯

VideoMAE：简单高效的视频自监督预训练新范式｜NeurIPS 2022

VideoMAE：简单高效的视频自监督预训练新范式｜NeurIPS 2022

新智元

0+阅读 · 2022年11月28日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【泡泡一分钟】用于RGBD语义分割的三维图神经网络(ICCV2017-546)

【泡泡一分钟】用于RGBD语义分割的三维图神经网络(ICCV2017-546)

泡泡机器人SLAM

22+阅读 · 2018年12月4日

【泡泡一分钟】学习紧密的几何特征（ICCV2017-17）

【泡泡一分钟】学习紧密的几何特征（ICCV2017-17）

泡泡机器人SLAM

20+阅读 · 2018年5月8日

【论文推荐】最新七篇图像检索相关论文—草图、Tie-Aware、场景图解析、叠加跨注意力机制、深度哈希、人群估计

【论文推荐】最新七篇图像检索相关论文—草图、Tie-Aware、场景图解析、叠加跨注意力机制、深度哈希、人群估计

专知

10+阅读 · 2018年4月22日

【泡泡一分钟】基于多视图卷积网络的草图三维重建技术(3dv-66)

【泡泡一分钟】基于多视图卷积网络的草图三维重建技术(3dv-66)

泡泡机器人SLAM

11+阅读 · 2018年3月31日

【泡泡一分钟】将3D全卷积网络应用于车辆激光点云处理（IROS-11）

【泡泡一分钟】将3D全卷积网络应用于车辆激光点云处理（IROS-11）

泡泡机器人SLAM

13+阅读 · 2018年3月23日

【泡泡一分钟】Matterport3D: 从室内RGBD数据集中训练 (3dv-22)

【泡泡一分钟】Matterport3D: 从室内RGBD数据集中训练 (3dv-22)

泡泡机器人SLAM

16+阅读 · 2017年12月31日

相关论文

Self-supervised dense representation learning for live-cell microscopy with time arrow prediction

Arxiv

0+阅读 · 2023年5月9日

Self-supervised Learning for Pre-Training 3D Point Clouds: A Survey

Arxiv

5+阅读 · 2023年5月8日

Equiangular Basis Vectors

Arxiv

0+阅读 · 2023年5月8日

Capacity Achieving Codes for an Erasure Queue-Channel

Arxiv

0+阅读 · 2023年5月7日

A vector quantized masked autoencoder for audiovisual speech emotion recognition

Arxiv

0+阅读 · 2023年5月5日

Contrastive Learning for Low-light Raw Denoising

Arxiv

0+阅读 · 2023年5月5日

Hybrid Transducer and Attention based Encoder-Decoder Modeling for Speech-to-Text Tasks

Arxiv

0+阅读 · 2023年5月4日

Federated Ensemble-Directed Offline Reinforcement Learning

Arxiv

0+阅读 · 2023年5月4日

Masked Autoencoders Are Scalable Vision Learners

Arxiv

27+阅读 · 2021年11月11日

SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition

Arxiv

12+阅读 · 2021年5月30日

相关基金

基于深度学习的金丝猴面部特性的检测与识别算法研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于张量投票的车载LiDAR数据的目标识别

国家自然科学基金

1+阅读 · 2015年12月31日

基于非易失内存设备的数据读写性能优化方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

基于2-D空间离散数据的质量与产出的预测方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于局部纹理特征的图像细节超分辨率技术研究

国家自然科学基金

1+阅读 · 2013年12月31日

碳材料改性的金/聚合物催化剂及其低温氧化醇的性能研究

国家自然科学基金

0+阅读 · 2013年12月31日

云环境中基于三维世界模型的图像表示与压缩

国家自然科学基金

2+阅读 · 2013年12月31日

基于复杂网络的中文文本语义相似度研究

国家自然科学基金

3+阅读 · 2012年12月31日

极化合成孔径雷达(SAR)图像地物并行分割分类研究与应用

国家自然科学基金

1+阅读 · 2012年12月31日

基于图像空间视觉相似性的质量评价方法

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员