Abstract:
In order to solve the problem of poor performance caused by various scenes in speaker recognition, a recognition method based on parallel convolution and dual attention mechanism is proposed. Based on the structure of ECAPA-TDNN model, this method improves the network components and loss function calculation module. Firstly, the improvement of residual module is to introduce the idea of “grouping”, and obtain multi-level features by constructing multi-branch parallel connection in residual block. Secondly, the improvement of the attention module is to use the two mechanisms of channel attention and spatial attention to weight the attention of different positions of features, which is convenient for the model to adaptively select and emphasize features and capture global features and local key information. Then, the Sub-center loss function is used to calculate the loss and deal with the multi-variable characteristics. Finally, the validity of the model is evaluated on a large Chinese multi-scene dataset CN-Celeb, and six single scenes of the dataset are selected to test the speaker recognition system. The experimental results show that compared with the ResNet34 model and ECAPA-TDNN model, the EER is reduced by 6.03% and 5.57%, and the minDCF is reduced by 7.31% and 7.02%, respectively. The average test results of six single scenarios are lower than the test set results, and perform well in “drama” and “speech” scenarios, with the lowest EER of 4.48% and the lowest minDCF of 0.2322. Experiments show that the method has strong superiority and adaptability, and can effectively identify different scenes, thus improving the accuracy and robustness of speaker recognition.