• 基于双通道多特征融合网络语音情感识别

    Dual stream multi-feature fusion speech emotion recognition

    • 单一语音特征无法充分表达语音情感,而多个声学特征简单拼接的融合方法容易造成情感信息丢失,且采用单个通道特征提取网络无法全面提取语音中的情感特征。针对上述问题,本文提出基于双通道特征融合网络语音情感识别,以梅尔频率倒谱系数、均方根、过零率和色度短时傅里叶变换这4种对情感种类贡献度较大的语音特征作为输入,采用双通道网络结构分别提取短时局部特征和全局上下文特征;将基于一维空洞卷积的局部特征提取网络和引入自注意力机制的双向长短时记忆全局特征提取网络并行,避免信息相互干扰;利用投票策略的集成学习方法实现各通道深层特征的充分融合,以获得语音中更深层的情感信息和更加精准的分类准确率。实验结果表明:基于双通道多特征融合网络模型在TESS、RAVDESS、SAVEE、CREMA-D数据集和混合数据集实现了99.89%、95.89%、96.61%、97.75%和95.13%的情感识别准确率,与同类型的多个语音情感识别模型相比性能优异,识别准确率高于其他模型。

       

      Abstract: Single acoustic features are insufficient to fully convey vocal emotions, and simple concatenation methods of fusing multiple acoustic features can lead to loss of emotional information, while single-stream feature extraction networks are inadequate for comprehensively extracting emotional features from speech. To address these issues, this paper proposes a dual- stream feature fusion network for speech emotion recognition. The system inputs four acoustic features (Mel-frequency cepstral coefficients, root mean square, zero-crossing rate, and chroma short-time Fourier transform) that contribute significantly to the distinction of emotional categories. It employs a dual-stream network architecture to extract both short-term local features and global contextual features. The network integrates a local feature extraction network based on one-dimensional dilated convolutions and a global feature extraction network utilizing a bi-directional long short-term memory network with a self-attention mechanism in parallel, thus preventing interference of information. An ensemble learning method with a voting strategy is utilized to achieve comprehensive fusion of deep features from each stream, extracting more profound emotional information and improving classification accuracy. Experimental results demonstrate that the dual-stream multi-feature fusion network model achieves emotion recognition accuracies of 99.89%, 95.89%, 96.61%, 97.75%, and 95.13% on the TESS, RAVDESS, SAVEE, CREMA-D datasets, and a mixed dataset, respectively. Compared to other similar speech emotion recognition models, it exhibits superior performance, with recognition accuracies surpassing those of alternative models.

       

    /

    返回文章
    返回