Adversarial Training and Cross-modal Feature Fusion in Multimodal Sentiment Analysis
Conference Publication ResearchOnline@JCUMultimodal sentiment analysis recognizes emotions through text, audio, and visual modalities, but data incompleteness is a major challenge. Existing methods often focus on specific types of deficiencies and perform poorly when multiple types of noise are present simultaneously. To address this issue, we propose a noise-prompted adversarial training framework with a multimodal interaction model to enhance the model's robustness to missing modalities. The model first extracts common and unique features from each modality using a BERT text encoder and a shared-private encoder. Correlation measurements are then used to calculate the similarity between modalities, and a weighting mechanism is applied to the shared features. These features are deeply fused using a Transformer, and adversarial training combined with semantic reconstruction supervision helps the model learn a unified representation of noisy and clean data. Experimental results show that this method significantly improves the performance of multimodal sentiment analysis.
N/A
ICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings
N/A
979-8-3503-6874-1
N/A
N/A
5
Hyderabad, India
IEEE
N/A
Piscataway, NJ, USA
N/A
N/A
N/A
N/A
10.1109/ICASSP49660.2025.10890023
