A Method upon Deep Learning for Speech Emotion Recognition

Nhat Truong Pham; Duc Ngoc Minh Dang; Sy Dung Nguyen

doi:10.25073/jaec.202044.311

ISSN (Online): 2588-123X
ISSN (Print):

About TDTU

Ton Duc Thang University (TDTU) is a public university with the main campus located in vibrant Ho Chi Minh City, Vietnam’s economic and educational hub. Founded in 1997, TDTU has developed into one of the largest and fastest growing universities in Vietnam with more than 22,000 students, enrolled in undergraduate and graduate programs ranging from science, engineering to business management, law, and humanities. To foster the country’s human resources and best serve the nation in the knowledge based economy of the 21st century, TDTU is combining vocational training with high-level research. The establishment of JAEC is one of TDTU’s efforts in this direction. More

Publication Information

Publisher

Ton Duc Thang University

Honorary Editor-in-Chief

Tran Trong Dao

Executive Editor

Nguyen Trung Thang

Chairman of the Editorial Board

Vice Chairman of the Editorial Board

Editorial Board

Hari Mohan Srivastava

Juan Carlos Burguillo Rial

Akhil Garg

Nguyen Pham Trung Hieu

Mahdi Shariati

Aleš Zamuda

Ngo Son Tung

User

Guide for Authors

View 'Guide for Authors' online

Submit Your Paper

In order to submit your paper, please login and navigate to the author page.

If you do not have an account, please consider registering one.

Track Your Paper

Track accepted paper
Once your article has been accepted you will receive an email from Author Services. This email contains a link to check the status of your articles.

Click here to track your accepted papers

Journal Content

Browse

Abstracting/Indexing

A Method upon Deep Learning for Speech Emotion Recognition

Nhat Truong Pham, Duc Ngoc Minh Dang, Sy Dung Nguyen

Abstract

Feature extraction and emotional classification are significant roles in speech emotion recognition. It is hard to extract and select the optimal features, researchers can not be sure what the features should be. With deep learning approaches, features could be extracted by using hierarchical abstraction layers, but it requires high computational resources and a large number of data. In this article, we choose static, differential, and acceleration coefficients of log Mel-spectrogram as inputs for the deep learning model. To avoid performance degradation, we also add a skip connection with dilated convolution network integration. All representatives are fed into a self-attention mechanism with bidirectional recurrent neural networks to learn long term global features and exploit context for each time step. Finally, we investigate contrastive center loss with softmax loss as loss function to improve the accuracy of emotion recognition. For validating robustness and effectiveness, we tested the proposed method on the Emo-DB and ERC2019 datasets. Experimental results show that the performance of the proposed method is strongly comparable with the existing state-of-the-art methods on the Emo-DB and ERC2019 with 88% and 67%, respectively.

This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium provided the original work is properly cited.

Full Text:

PDF

Time cited: 0

Download citation

DOI: http://dx.doi.org/10.25073/jaec.202044.311

Refbacks

There are currently no refbacks.

This work is licensed under a Creative Commons Attribution 4.0 International License.

Owner: Ton Duc Thang University. All rights reserved.
License No: 507/GP-BTTTT, issued: 18th November 2016
Contact address: 19, Nguyen Huu Tho Street, Tan Hung Ward, Ho Chi Minh City
Tel: +84-28 3775 5037 Fax: +84-28 3775 5055

This work is licensed under a Creative Commons Attribution 4.0 International License.

Username
Password