Transformer-Based Architecture for Multi-Modal Sentiment Analysis in Social Media Streams
Zhang, X., Li, Y., Wang, H., Chen, J.. Transformer-Based Architecture for Multi-Modal Sentiment Analysis in Social Media Streams. Loshu Comput. Intell..
Vol.2, No.1. Jan 2025. https://doi.org/10.58921/ljci.2025.0101
Article
Recommended articles
Cited by 29
Metrics
Highlights
- This paper presents TransSentiment, a novel transformer-based architecture for multi-modal sentiment analysis combining text, image, and audio signals from social media content.
- Our model employs cross-modal attention to capture inter-modal dependencies and temporal dynamics across modalities.
- Evaluated on CMU-MOSI, MELD, and a newly collected TikTok-Sentiment dataset (47,000 samples), TransSentiment achieves state-of-the-art F1 scores of 89.3%, 91.7%, and 85.4% respectively, outperforming existing single-modal and fusion approaches by margins of 3.2–7.8%.
Abstract
This paper presents TransSentiment, a novel transformer-based architecture for multi-modal sentiment analysis combining text, image, and audio signals from social media content. Our model employs cross-modal attention to capture inter-modal dependencies and temporal dynamics across modalities. Evaluated on CMU-MOSI, MELD, and a newly collected TikTok-Sentiment dataset (47,000 samples), TransSentiment achieves state-of-the-art F1 scores of 89.3%, 91.7%, and 85.4% respectively, outperforming existing single-modal and fusion approaches by margins of 3.2–7.8%. Analysis reveals that audio features provide complementary information to text-image pairs, particularly for detecting sarcasm and irony.