Citation
Kamaruzaman, Nurul Nadhrah
(2024)
Deep hierarchical weighted late fusion model for multimodal emotion recognition.
Doctoral thesis, Universiti Putra Malaysia.
Abstract
Emotion recognition gains its relevance in diverse areas such as healthcare, education,
business marketing and entertainment. As time goes on, research in emotion
recognition analysis has developed from the traditional single modal to complex
multimodal analysis. Extraction of data from various modalities such as audio, text
and video are required in order to gain a meaningful pattern, thus increasing the
accuracy of the recognition. In this study, the focus was given on implementing the
late fusion approach to integrate three modalities for multimodal emotion recognition.
Nevertheless, the late fusion approach faces several challenges, including the loss of
fine-grained temporal and spatial information, class imbalance across modalities and
reliance on individual model training performance. To address these issues, this study
presents the Deep Hierarchical Weighted Late Fusion (DHWLF) model, which is
designed for an accurate multimodal emotion recognition. DHWLF employs deep
hierarchical weighted late fusion method to effectively reduce the loss of fine-grained
temporal and spatial information. This study also introduces individual emotion
recognition models for audio, text, and video image modalities using deep learning techniques. Specifically, for the audio, this study introduces the SMOTE-2DCNN
model, which combines the Synthetic Minority Oversampling Technique (SMOTE)
method with a 2-Dimensional Convolutional Neural Network (2DCNN) to handle
imbalanced class distribution and classify emotions accurately. For the text, this study
implements the Bidirectional Encoder Representations from Transformers (BERT)
language model known as TextBERT model to capture semantic and contextual
information from textual data. Lastly, for the video, this study proposes the
Fer2DCNN model, which utilizes the Cascade Haar Classifier to detect facial regions
before feeding them into a 2DCNN for emotion classification. To evaluate the
effectiveness of the DHWLF model, extensive experiments were conducted. The
results demonstrated its superiority with the DHWLF model achieving 90% accuracy,
outperforming both individual emotion recognition models and other state-of-the-art
multimodal emotion recognition models from prior studies.
Download File
Additional Metadata
| Item Type: |
Thesis
(Doctoral)
|
| Subject: |
Emotional intelligence |
| Subject: |
Artificial emotional intelligence |
| Subject: |
Multisensor data fusion |
| Call Number: |
FSKTM 2024 23 |
| Chairman Supervisor: |
Nor Azura binti Husin |
| Divisions: |
Faculty of Computer Science and Information Technology |
| Keywords: |
Late fusion; Weighted late fusion; Multimodal; Emotion recognition; Deep learning |
| Sustainable Development Goals (SDGs): |
SDG 3: Good Health and Well-being, SDG 4: Quality Education, SDG 9: Industry, Innovation and Infrastructure |
| Depositing User: |
MS. HADIZAH NORDIN
|
| Date Deposited: |
13 Aug 2026 23:55 |
| Last Modified: |
13 Aug 2026 23:55 |
| URI: |
http://psasir.upm.edu.my/id/eprint/127803 |
| Statistic Details: |
View Download Statistic |
Actions (login required)
 |
View Item |