• 基于并行卷积和双重注意力机制的说话人识别

    Speaker recognition based on parallel convolution and dual attention mechanism

    • 为解决说话人识别面对多种场景造成性能变差的问题,提出了一种基于并行卷积和双重注意力机制的识别方法。该方法基于ECAPA-TDNN模型结构,对网络组件和损失函数计算模块进行了改进。首先,残差模块的改进是引入“分组”的思想,通过在残差块构建多分支并行连接,获取多层次特征。其次,注意力模块的改进是利用通道注意力和空间注意力两种机制,对特征的不同位置进行注意力加权,便于模型自适应地选择和强调特征,捕获全局特征和局部关键信息。然后,使用Sub-center loss函数计算损失,应对多变化特征。最后,在大型的中文多场景数据集CN-Celeb上评估模型的有效性并选取数据集的六种单一场景测试说话人识别系统。实验结果表明:与ResNet34模型和ECAPA-TDNN模型相比,EER分别降低了6.03%和5.57%,minDCF分别降低了7.31%和 7.02%;6种单一场景测试结果的均值低于测试集结果,且在“drama”和“speech”场景下表现优异,EER最低仅有4.48%,minDCF最低为0.2322。说明该方法具有强大的优越性和适应性,能够针对不同场景进行有效识别,从而提高说话人识别准确率和鲁棒性。

       

      Abstract: In order to solve the problem of poor performance caused by various scenes in speaker recognition, a recognition method based on parallel convolution and dual attention mechanism is proposed. Based on the structure of ECAPA-TDNN model, this method improves the network components and loss function calculation module. Firstly, the improvement of residual module is to introduce the idea of “grouping”, and obtain multi-level features by constructing multi-branch parallel connection in residual block. Secondly, the improvement of the attention module is to use the two mechanisms of channel attention and spatial attention to weight the attention of different positions of features, which is convenient for the model to adaptively select and emphasize features and capture global features and local key information. Then, the Sub-center loss function is used to calculate the loss and deal with the multi-variable characteristics. Finally, the validity of the model is evaluated on a large Chinese multi-scene dataset CN-Celeb, and six single scenes of the dataset are selected to test the speaker recognition system. The experimental results show that compared with the ResNet34 model and ECAPA-TDNN model, the EER is reduced by 6.03% and 5.57%, and the minDCF is reduced by 7.31% and 7.02%, respectively. The average test results of six single scenarios are lower than the test set results, and perform well in “drama” and “speech” scenarios, with the lowest EER of 4.48% and the lowest minDCF of 0.2322. Experiments show that the method has strong superiority and adaptability, and can effectively identify different scenes, thus improving the accuracy and robustness of speaker recognition.

       

    /

    返回文章
    返回