Abstract:
Graph convolutional networks (GCN) have achieved remarkable performance in skeleton-based action recognition, emerging as a dominant research direction in the field. However, existing GCN are often limited to constructing graph topologies at a single physical joint scale, and tend to overlook potential temporal dynamics during spatial modeling, which restricts their ability to represent complex action features. To address these limitations, we propose an action recognition method based on a multi-level semantic prototype graph convolutional network. The method introduces a multi-level semantic prototype graph convolution module, which utilizes human structural priors to define four semantic granularities and maps physical nodes to a compact prototype space via a learnable routing matrix, and extracts features in parallel to capture deep spatial correlations of skeletons across different scales. Simultaneously, a temporal context spatial attention module is designed; this module utilizes a multi-branch temporal convolutional encoder to extract local motion features and integrates them into the spatial attention mechanism, generating graph topologies that encapsulate temporal dynamics. Experimental results on four benchmark datasets, NTU RGB+D 60, NTU RGB+D 120, Kinetics-Skeleton 400, and Northwestern-UCLA, demonstrate that the proposed method achieves Top-1 accuracies of 93.6%, 91.3%, 50.3% and 97.3% respectively, outperforming current state-of-the-art methods in recognition accuracy.