自动驾驶公共安全中的高效预训练单流点云目标跟踪

Efficient Pre-training Single-stream Point Cloud Object Tracking for Autonomous Driving Public Safety

  • 摘要: 自动驾驶中的公共安全问题受到社会的广泛关注,而基于点云数据的3维单目标跟踪技术在自动驾驶安全保障中发挥了关键作用。目前,大多数主流跟踪器利用Transformer捕获远程依赖关系的能力进行相似性匹配或增强目标特征。然而,这种应用忽视了Transformer在适应各种预训练模型方面的灵活性,而且增加了计算负担。为此,本文引入了一种基于单分支框架的预训练跟踪器(SPTracker)。这项工作的本质是将预训练参数高效微调应用到3维单目标跟踪领域,但面临独特的挑战和领域差距使其更为复杂。首先,当预先训练的主干网络迁移到下游任务时,通常要求上下游任务的网络结构具备兼容性。其次,与点云分析任务需要处理的复杂对象相比,动态的跟踪对象往往类别少,且几何结构更相似。为了克服这些限制,本文基于单流框架引入补丁嵌入网络,以统一Transformer的位置嵌入和输入方式,从而增强模型在上下游任务间的适应性。同时,通过对预训练网络参数高效微调,平衡了整个预训练网络的学习能力。最后,在密集鸟瞰图特征空间中设计了一个高效的目标定位网络,以更少的计算开销实现更好的性能。提出的SPTracker以52.6 帧/s的速度运行,模型的训练时间从约48 h减少至约16 h,在KITTI和nuScenes数据集上,SPTracker实现了较为先进的性能表现。

     

    Abstract: Public safety issues in autonomous driving receive widespread social attention. The technology of 3D single-object tracking based on point cloud data gains significant attention due to its crucial role in robotics and autonomous driving. Currently, most mainstream trackers leverage the ability of transformers to capture long-range dependencies for similarity matching or enhancing target features. However, this application overlooks the flexibility of transformers in adapting to various pre-trained models and increases the computational burden. To address this limitations, we introduce a pre-trained tracker based on a single-branch framework (SPTracker). The essence of this work lies in the efficient fine-tuning of pre-trained parameters applied to the domain of 3D single-object tracking. However, unique challenges and potential domain gaps make this application less straightforward than intuitively expected. First, when transferring a pre-trained backbone to downstream tasks, structural compatibility between upstream and downstream tasks is often required. Second, compared to the complex objects that point cloud analysis tasks must handle, dynamic tracked objects typically have fewer categories and greater geometric similarity. To overcome these limitations, based on the single-branch framework, we introduce a patch embedding network to unify the position embedding and input of the transformer, thereby enhancing the model’s adaptability between upstream and downstream tasks. Additionally, through efficient fine-tuning of the pre-trained network parameters, we balance the learning capacity of the entire model. Finally, we design an efficient object localization network in a dense bird’s-eye view feature space to achieve better performance with reduced computational overhead. The proposed SPTracker operates at a speed of 52.6 frames per second, reducing the model training time from approximately 48 h to around 16 h, while achieving advanced performance on the KITTI and nuScenes datasets.

     

/

返回文章
返回