Adversarial Training and Cross-modal Feature Fusion in Multimodal Sentiment Analysis

Conference Publication ResearchOnline@JCU
Li, Junhuai;Lin, Chuang;Wang, Huaijun;Zhi, Yuxing;Chen, Jing;Huang, Tao
Abstract

Multimodal sentiment analysis recognizes emotions through text, audio, and visual modalities, but data incompleteness is a major challenge. Existing methods often focus on specific types of deficiencies and perform poorly when multiple types of noise are present simultaneously. To address this issue, we propose a noise-prompted adversarial training framework with a multimodal interaction model to enhance the model's robustness to missing modalities. The model first extracts common and unique features from each modality using a BERT text encoder and a shared-private encoder. Correlation measurements are then used to calculate the similarity between modalities, and a weighting mechanism is applied to the shared features. These features are deeply fused using a Transformer, and adversarial training combined with semantic reconstruction supervision helps the model learn a unified representation of noisy and clean data. Experimental results show that this method significantly improves the performance of multimodal sentiment analysis.

Journal

N/A

Publication Name

ICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings

Volume

N/A

ISBN/ISSN

979-8-3503-6874-1

Edition

N/A

Issue

N/A

Pages Count

5

Location

Hyderabad, India

Publisher

IEEE

Publisher Url

N/A

Publisher Location

Piscataway, NJ, USA

Publish Date

N/A

Url

N/A

Date

N/A

EISSN

N/A

DOI

10.1109/ICASSP49660.2025.10890023