• 基于MSS-YOLO-MPHands融合架构的装配手部动作识别研究

    Research on assembly action recognition based on MSS-YOLO-MPHands fusion architecture

    • 为解决工业智能化场景中因动态光照干扰、复杂背景、手部移动及局部遮挡所导致的手部动作识别精度不足问题,提出了一种MSS-YOLO-MPHands融合架构。该方案首先对YOLOv8进行改进,构建MSS-YOLO模型,以优化Mediapipe Hands的骨骼关键点检测效果;同时,采用MobileNetv4 UIB模块重构原有C2f结构,显著提升模型对工业场景中常见遮挡与光照变化的鲁棒性。为进一步增强检测性能,在主干网络中引入STA注意力机制与SPConv模块。最后,结合Bi-LSTM时序建模方法,构建端到端的动作识别框架。实验结果表明:MSS-YOLO-MPHands模型在数据集上的PCK达到93.85%,手部动作识别平均准确率为92.5%,验证了该方案在复杂工业环境下的有效性与可行性。

       

      Abstract: To address the issue of insufficient hand action recognition accuracy caused by dynamic illumination interference, complex backgrounds, hand movement, and partial occlusion in intelligent industrial scenarios, this paper proposes a fused architecture termed MSS-YOLO-MPHands. The scheme first improves YOLOv8 by constructing the MSS-YOLO model to optimize the skeletal keypoint detection performance of MediaPipe Hands. Meanwhile, the MobileNetV4-UIB module is adopted to reconstruct the original C2f structure, substantially enhancing the model's robustness against common occlusion and illumination variations in industrial scenes. To further boost detection performance, the STA (Spatial-Temporal Attention) mechanism and the SPConv module are introduced into the backbone network. Finally, a Bi-LSTM-based temporal modeling method is incorporated to build an end-to-end action recognition framework. Experimental results demonstrate that the MSS-YOLO-MPHands model achieves a PCK (Percentage of Correct Keypoints) of 93.85% and an average hand action recognition accuracy of 92.5%, validating the effectiveness and feasibility of the proposed scheme in complex industrial environments.

       

    /

    返回文章
    返回