Sarcasm Detection in Multilingual Text through Embedding-Enhanced Language Models: BERT Variants
Conference Publication ResearchOnline@JCUThe rise of social media has amplified sarcastic content, where literal and intended meanings often diverge. Sarcasm poses significant challenges for sentiment analysis due to its subtle linguistic nuances. Traditional models struggle with sarcasm detection because of their limited contextual understanding and reliance on handcrafted features. This study curates a large dataset of English headline news, translated into Urdu to create a multilingual dataset. BERT variants, such as mBERT and UrduBERT, are employed to capture context and learn relevant features using deep learning architectures. The study compares the performance of these BERT variants on the dataset and evaluates traditional algorithms for classifying sarcastic and non-sarcastic headlines. Model performance is rigorously assessed using precision, recall, and F1-score metrics. Results indicate that fine-tuned mBERT outperforms other BERT variants and traditional models in sarcasm detection by effectively capturing complex semantic patterns. The dataset and code are publicly available on GitHub.https://github.com/MislamSatti/Sarcasam-Detection-with-code
N/A
Proceedings of the IEEE International Multi Topic Conference Inmic
N/A
2835-8864
N/A
2024
6
Karachi, Pakistan
IEEE
N/A
Piscataway, NJ, USA
N/A
N/A
N/A
N/A
10.1109/INMIC64792.2024.11004369
