基于多层次语义原型图卷积的动作识别方法

Multi-level Semantic Prototype Graph Convolutional Networks for Skeleton-based Action Recognition

  • 摘要: 图卷积网络在基于骨架的动作识别任务中表现优异,已成为动作识别领域的重要方向。然而,现有图卷积网络常局限于单一结构尺度构建图拓扑结构,且在空间建模过程中忽略了潜在的时间动态特性,导致其对复杂动作特征的表达能力受限。为此,本文提出一种基于多层次语义原型图卷积网络的动作识别方法。该方法引入多层次语义原型图卷积模块,利用人体结构先验划分4个语义粒度,并通过可学习的路由矩阵将物理节点映射至紧凑的原型空间,并行提取特征以捕捉骨架在不同尺度下的深层空间关联;同时设计时间上下文空间注意力模块,利用多分支时间卷积编码器提取局部运动特征并将其融入空间注意力机制,生成蕴含时序动态的图拓扑。在NTU RGB+D 60、NTU RGB+D 120、Kinetics-Skeleton 400及Northwestern-UCLA四个基准数据集上的实验结果表明,该方法的Top-1准确率分别达到93.6%、91.3%、50.3%和97.3%,在识别精度上优于当前主流方法。

     

    Abstract: Graph convolutional networks (GCN) have achieved remarkable performance in skeleton-based action recognition, emerging as a dominant research direction in the field. However, existing GCN are often limited to constructing graph topologies at a single physical joint scale, and tend to overlook potential temporal dynamics during spatial modeling, which restricts their ability to represent complex action features. To address these limitations, we propose an action recognition method based on a multi-level semantic prototype graph convolutional network. The method introduces a multi-level semantic prototype graph convolution module, which utilizes human structural priors to define four semantic granularities and maps physical nodes to a compact prototype space via a learnable routing matrix, and extracts features in parallel to capture deep spatial correlations of skeletons across different scales. Simultaneously, a temporal context spatial attention module is designed; this module utilizes a multi-branch temporal convolutional encoder to extract local motion features and integrates them into the spatial attention mechanism, generating graph topologies that encapsulate temporal dynamics. Experimental results on four benchmark datasets, NTU RGB+D 60, NTU RGB+D 120, Kinetics-Skeleton 400, and Northwestern-UCLA, demonstrate that the proposed method achieves Top-1 accuracies of 93.6%, 91.3%, 50.3% and 97.3% respectively, outperforming current state-of-the-art methods in recognition accuracy.

     

/

返回文章
返回