Relevance Units Machine based Dimensional and Continuous Speech Emotion Prediction This publication appears in: Multimedia Tools and Applications Authors: F. Wang, H. Sahli, J. Gao, D. Jiang and W. Verhelst Volume: 74 Issue: 22 Pages: 9983-10000 Publication Date: Oct. 2015
Abstract: Emotion plays a significant role in human-computer interaction. The continuing improvements in speech technology have let to many new and fascinating applications in human-computer interaction, context aware computing and computer mediated communication. Such applications require reliable online recognition of the user's affect. However most emotion recognition systems are based son speech via an isolated short sentence or word. We present a framework for online emotion recognition from speech. On the front-end, a voice activity detection algorithm is used to segment the input speech, and feature are estimated to model long-term properties. Then, dimensional and continuous emotion recognition is performed via a Relevance Units Machine (RUM). The advantages of the proposed systems are: (i) its computational efficiency in run-time (regression outputs can be produced continuously in pseudo real-time), (ii) RUM offers superior sparsity to the well-known Support Vector Regression (SVR) and Relevance Vector Machine for regression (RVR), and (iii) RUM's predictive performance is comparable to SVR and RVR.
|