| dc.description.abstract |
In order to improve human-computer interaction and create emotionally intelligent systems, au tomatic emotion classification from speech is essential. Large pre-trained audio language models (ALMs) have shown good generalization on a variety of tasks, but little is known about how well they perform on low-resource and multilingual datasets, particularly when code-switched or bilingual speech is involved. In this study, we examine the ability of cutting-edge pre-trained models to categorize emotions from a bilingual speech dataset that includes both Tamil and English utterances. To ensure balanced representation in both languages, our carefully selected dataset contains labeled audio samples from core emotion classes, including happy, sad, angry, and neutral. Using sophisticated ALMs such as Wav2Vec2, HuBERT, and Whisper, we extract fixed-length embeddings. We then assess these representations using Convolutional Neural Net works (CNNs) and Fully Connected Networks (FCNs). We investigate a dual-branch CNN archi tecture supplemented with a contrastive loss component to address intra-class language variance and inter-class emotion similarity. Our findings shed light on embedding-level language-agnostic emotion representation and demonstrate the potential of ALMs in robust emotion recognition, even in multilingual contexts. |
en_US |