UPM Institutional Repository

Deep hierarchical weighted late fusion model for multimodal emotion recognition


Citation

Kamaruzaman, Nurul Nadhrah (2024) Deep hierarchical weighted late fusion model for multimodal emotion recognition. Doctoral thesis, Universiti Putra Malaysia.

Abstract

Emotion recognition gains its relevance in diverse areas such as healthcare, education, business marketing and entertainment. As time goes on, research in emotion recognition analysis has developed from the traditional single modal to complex multimodal analysis. Extraction of data from various modalities such as audio, text and video are required in order to gain a meaningful pattern, thus increasing the accuracy of the recognition. In this study, the focus was given on implementing the late fusion approach to integrate three modalities for multimodal emotion recognition. Nevertheless, the late fusion approach faces several challenges, including the loss of fine-grained temporal and spatial information, class imbalance across modalities and reliance on individual model training performance. To address these issues, this study presents the Deep Hierarchical Weighted Late Fusion (DHWLF) model, which is designed for an accurate multimodal emotion recognition. DHWLF employs deep hierarchical weighted late fusion method to effectively reduce the loss of fine-grained temporal and spatial information. This study also introduces individual emotion recognition models for audio, text, and video image modalities using deep learning techniques. Specifically, for the audio, this study introduces the SMOTE-2DCNN model, which combines the Synthetic Minority Oversampling Technique (SMOTE) method with a 2-Dimensional Convolutional Neural Network (2DCNN) to handle imbalanced class distribution and classify emotions accurately. For the text, this study implements the Bidirectional Encoder Representations from Transformers (BERT) language model known as TextBERT model to capture semantic and contextual information from textual data. Lastly, for the video, this study proposes the Fer2DCNN model, which utilizes the Cascade Haar Classifier to detect facial regions before feeding them into a 2DCNN for emotion classification. To evaluate the effectiveness of the DHWLF model, extensive experiments were conducted. The results demonstrated its superiority with the DHWLF model achieving 90% accuracy, outperforming both individual emotion recognition models and other state-of-the-art multimodal emotion recognition models from prior studies.


Download File

[img] Text
FSKTM 2024 23 - Full Text.pdf
Available under License Creative Commons Attribution Non-commercial No Derivatives.

Download (3MB)
[img] Text
FSKTM 2024 23.pdf
Restricted to Repository staff only
Available under License Creative Commons Attribution Non-commercial No Derivatives.

Download (4MB)

Additional Metadata

Item Type: Thesis (Doctoral)
Subject: Emotional intelligence
Subject: Artificial emotional intelligence
Subject: Multisensor data fusion
Call Number: FSKTM 2024 23
Chairman Supervisor: Nor Azura binti Husin
Divisions: Faculty of Computer Science and Information Technology
Keywords: Late fusion; Weighted late fusion; Multimodal; Emotion recognition; Deep learning
Sustainable Development Goals (SDGs): SDG 3: Good Health and Well-being, SDG 4: Quality Education, SDG 9: Industry, Innovation and Infrastructure
Depositing User: MS. HADIZAH NORDIN
Date Deposited: 13 Aug 2026 23:55
Last Modified: 13 Aug 2026 23:55
URI: http://psasir.upm.edu.my/id/eprint/127803
Statistic Details: View Download Statistic

Actions (login required)

View Item View Item