. 2020 Sep 17;20(18):5328.

doi: 10.3390/s20185328.

FusionSense: Emotion Classification Using Feature Fusion of Multimodal Data and Deep Learning in a Brain-Inspired Spiking Neural Network

Clarence Tan¹, Gerardo Ceballos², Nikola Kasabov¹, Narayan Puthanmadam Subramaniyam^{3

4}

Affiliations

¹ Knowledge Engineering and Discovery Research Institute, Auckland University of Technology, Auckland 1010, New Zealand.
² School of Electrical Engineering, University of Los Andes, Merida 5101, Venezuela.
³ Faculty of Medicine and Health Technology and BioMediTech Institute, Tampere University, 33520 Tampere, Finland.
⁴ Department of Neuroscience and Biomedical Engineering, School of Science, Aalto University, 02150 Espoo, Finland.

PMID: 32957655
PMCID: PMC7571195
DOI: 10.3390/s20185328

FusionSense: Emotion Classification Using Feature Fusion of Multimodal Data and Deep Learning in a Brain-Inspired Spiking Neural Network

Clarence Tan et al. Sensors (Basel). 2020.

. 2020 Sep 17;20(18):5328.

doi: 10.3390/s20185328.

Authors

Clarence Tan¹, Gerardo Ceballos², Nikola Kasabov¹, Narayan Puthanmadam Subramaniyam^{3

4}

Affiliations

¹ Knowledge Engineering and Discovery Research Institute, Auckland University of Technology, Auckland 1010, New Zealand.
² School of Electrical Engineering, University of Los Andes, Merida 5101, Venezuela.
³ Faculty of Medicine and Health Technology and BioMediTech Institute, Tampere University, 33520 Tampere, Finland.
⁴ Department of Neuroscience and Biomedical Engineering, School of Science, Aalto University, 02150 Espoo, Finland.

PMID: 32957655
PMCID: PMC7571195
DOI: 10.3390/s20185328

Abstract

Using multimodal signals to solve the problem of emotion recognition is one of the emerging trends in affective computing. Several studies have utilized state of the art deep learning methods and combined physiological signals, such as the electrocardiogram (EEG), electroencephalogram (ECG), skin temperature, along with facial expressions, voice, posture to name a few, in order to classify emotions. Spiking neural networks (SNNs) represent the third generation of neural networks and employ biologically plausible models of neurons. SNNs have been shown to handle Spatio-temporal data, which is essentially the nature of the data encountered in emotion recognition problem, in an efficient manner. In this work, for the first time, we propose the application of SNNs in order to solve the emotion recognition problem with the multimodal dataset. Specifically, we use the NeuCube framework, which employs an evolving SNN architecture to classify emotional valence and evaluate the performance of our approach on the MAHNOB-HCI dataset. The multimodal data used in our work consists of facial expressions along with physiological signals such as ECG, skin temperature, skin conductance, respiration signal, mouth length, and pupil size. We perform classification under the Leave-One-Subject-Out (LOSO) cross-validation mode. Our results show that the proposed approach achieves an accuracy of 73.15% for classifying binary valence when applying feature-level fusion, which is comparable to other deep learning methods. We achieve this accuracy even without using EEG, which other deep learning methods have relied on to achieve this level of accuracy. In conclusion, we have demonstrated that the SNN can be successfully used for solving the emotion recognition problem with multimodal data and also provide directions for future research utilizing SNN for Affective computing. In addition to the good accuracy, the SNN recognition system is requires incrementally trainable on new data in an adaptive way. It only one pass training, which makes it suitable for practical and on-line applications. These features are not manifested in other methods for this problem.

Keywords: Evolving Spiking Neural Networks (eSNNs); NeuCube; Spatio-temporal data; facial emotion recognition; multimodal data.

PubMed Disclaimer

Conflict of interest statement

The authors declare no conflict of interest.

Figures

**Figure 1**
Example of face detection in Mahnob-HCI showing the feature points tracked along the video.

**Figure 2**
Facial landmarks detection.

**Figure 4**
Elicited signal features in the last 30 seconds of video.

**Figure 5**
Boxplot for features in Mahnob-HCI dataset for valence emotional dimension.

**Figure 6**
Proposed method for emotion valence classification using NeuCube.

**Figure 7**
Encoding Continuous feature values to five neurons spiking.

**Figure 8**
Input neurons location for facial and peripheral features classification. n1 means for the neuron coding the lowest values and n5 the highest ones for each feature. Note there are 3 layers of input neuron in the cube, located at $z = - 30$ (facial), $z = 0$ (peripheral), and $z = 30$ (facial).

**Figure 9**
Leaky integrate-and-fire model (LIFM) neuron model. Small circles at neuron inputs represent connection weights. Note that input 1 has a bigger weight and it produces a larger effect in PSP.

**Figure 10**
Hebbian Learning rule, connection (synaptic modification) vs difference between post- and pre-synaptic times.

**Figure 11**
Neuron activity pattern example when NeuCube is trained using each Separate data (low and high valence).

See this image and copyright information in PMC

Cited by

Integrating Spatial and Temporal Information for Violent Activity Detection from Video Using Deep Spiking Neural Networks.
Wang X, Yang J, Kasabov NK. Wang X, et al. Sensors (Basel). 2023 May 6;23(9):4532. doi: 10.3390/s23094532. Sensors (Basel). 2023. PMID: 37177737 Free PMC article.
Affective computing of multi-type urban public spaces to analyze emotional quality using ensemble learning-based classification of multi-sensor data.
Li R, Yuizono T, Li X. Li R, et al. PLoS One. 2022 Jun 3;17(6):e0269176. doi: 10.1371/journal.pone.0269176. eCollection 2022. PLoS One. 2022. PMID: 35657805 Free PMC article.
A novel signal channel attention network for multi-modal emotion recognition.
Du Z, Ye X, Zhao P. Du Z, et al. Front Neurorobot. 2024 Sep 11;18:1442080. doi: 10.3389/fnbot.2024.1442080. eCollection 2024. Front Neurorobot. 2024. PMID: 39323931 Free PMC article.
Domain adaptation spatial feature perception neural network for cross-subject EEG emotion recognition.
Lu W, Zhang X, Xia L, Ma H, Tan TP. Lu W, et al. Front Hum Neurosci. 2024 Dec 17;18:1471634. doi: 10.3389/fnhum.2024.1471634. eCollection 2024. Front Hum Neurosci. 2024. PMID: 39741785 Free PMC article.
CIT-EmotionNet: convolution interactive transformer network for EEG emotion recognition.
Lu W, Xia L, Tan TP, Ma H. Lu W, et al. PeerJ Comput Sci. 2024 Dec 23;10:e2610. doi: 10.7717/peerj-cs.2610. eCollection 2024. PeerJ Comput Sci. 2024. PMID: 39896395 Free PMC article.

See all "Cited by" articles

References

1. Calvo R.A., D’Mello S. Affect detection: An interdisciplinary review of models, methods, and their applications. IEEE Trans. Affect. Comput. 2010;1:18–37. doi: 10.1109/T-AFFC.2010.1. - DOI
1. Edwards J., Jackson H.J., Pattison P.E. Emotion recognition via facial expression and affective prosody in schizophrenia: A methodological review. Clin. Psychol. Rev. 2002;22:789–832. doi: 10.1016/S0272-7358(02)00130-7. - DOI - PubMed
1. Fong T., Nourbakhsh I., Dautenhahn K. A survey of socially interactive robots. Robot. Auton. Syst. 2003;42:143–166. doi: 10.1016/S0921-8890(02)00372-X. - DOI
1. Russell J.A. A circumplex model of affect. J. Personal. Soc. Psychol. 1980;39:1161. doi: 10.1037/h0077714. - DOI
1. Gunes H., Schuller B., Pantic M., Cowie R. Emotion representation, analysis and synthesis in continuous space: A survey; Proceedings of the Face and Gesture 2011; Santa Barbara, CA, USA. 21–25 March 2011; pp. 827–834.

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions

LinkOut - more resources

Full Text Sources
Research Materials
- NCI CPTC Antibody Characterization Program

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

FusionSense: Emotion Classification Using Feature Fusion of Multimodal Data and Deep Learning in a Brain-Inspired Spiking Neural Network

Affiliations

FusionSense: Emotion Classification Using Feature Fusion of Multimodal Data and Deep Learning in a Brain-Inspired Spiking Neural Network

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

Similar articles

Cited by

References

MeSH terms

LinkOut - more resources

Full Text Sources

Research Materials

Abstract

Conflict of interest statement

Figures

Similar articles

Cited by

References

MeSH terms

Related information

LinkOut - more resources

Full Text Sources

Research Materials