Please use this identifier to cite or link to this item: http://repository.iiitd.edu.in/xmlui/handle/123456789/2116
Full metadata record
DC FieldValueLanguage
dc.contributor.authorKumar, Niteen-
dc.contributor.authorAgarwal, Niket-
dc.contributor.authorBuduru, Arun Balaji (Advisor)-
dc.date.accessioned2026-09-09T13:31:38Z-
dc.date.available2026-09-09T13:31:38Z-
dc.date.issued2025-07-18-
dc.identifier.urihttp://repository.iiitd.edu.in/xmlui/handle/123456789/2116-
dc.description.abstractIn order to improve human-computer interaction and create emotionally intelligent systems, au tomatic emotion classification from speech is essential. Large pre-trained audio language models (ALMs) have shown good generalization on a variety of tasks, but little is known about how well they perform on low-resource and multilingual datasets, particularly when code-switched or bilingual speech is involved. In this study, we examine the ability of cutting-edge pre-trained models to categorize emotions from a bilingual speech dataset that includes both Tamil and English utterances. To ensure balanced representation in both languages, our carefully selected dataset contains labeled audio samples from core emotion classes, including happy, sad, angry, and neutral. Using sophisticated ALMs such as Wav2Vec2, HuBERT, and Whisper, we extract fixed-length embeddings. We then assess these representations using Convolutional Neural Net works (CNNs) and Fully Connected Networks (FCNs). We investigate a dual-branch CNN archi tecture supplemented with a contrastive loss component to address intra-class language variance and inter-class emotion similarity. Our findings shed light on embedding-level language-agnostic emotion representation and demonstrate the potential of ALMs in robust emotion recognition, even in multilingual contexts.en_US
dc.language.isoen_USen_US
dc.publisherIIIT-Delhien_US
dc.subjectEmotion detectionen_US
dc.subjectPretrained Modelsen_US
dc.subjectCNNen_US
dc.subjectAudio Featuresen_US
dc.subjectSpeechen_US
dc.titleSolving healthcare application of speech processingen_US
dc.typeOtheren_US
Appears in Collections:Year-2025

Files in This Item:
File Description SizeFormat 
btp_report_2025 - Niteen Kumar.pdf
  Restricted Access
213.08 kBAdobe PDFView/Open Request a copy


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.