Please use this identifier to cite or link to this item: http://repository.iiitd.edu.in/xmlui/handle/123456789/2122
Full metadata record
DC FieldValueLanguage
dc.contributor.authorGupta, Aniket-
dc.contributor.authorMohania, Mukesh (Advisor)-
dc.date.accessioned2026-09-10T08:49:39Z-
dc.date.available2026-09-10T08:49:39Z-
dc.date.issued2024-11-27-
dc.identifier.urihttp://repository.iiitd.edu.in/xmlui/handle/123456789/2122-
dc.description.abstractRecent advancements in text-to-speech (TTS) technology have made synthetic speech sound more natural and expressive. However, creating audiobook narration that accurately captures emotions and maintains context over multiple sentences is still a challenge. In this thesis, we present a new approach to audiobook speech synthesis using the VITS model, which is known for its stability and high quality, along with the Gigaspeech dataset. The Gigaspeech dataset is particularly useful for this task because it includes voice and text samples from real audiobooks, covering both dialogue and narration parts. A key contribution of our work is the emotional labeling of the Gigaspeech dataset, where each entry is tagged with specific emotions. Additionally, we propose using large language models (LLMs) to automatically determine the emotion and speaker ID for each sentence. This allows the TTS model to generate speech that not only fits the context but also conveys the appropriate emotions. We plan to achieve this as one of our future goal by fine-tuning the decoder of the VITS model using the emotionally labeled Gigaspeech dataset. Our experimental findings currently on smaller subsets of our data indicate that our text-to- speech system reliably generates natural audiobook narration, successfully capturing the unique speaking style of individual narrators.en_US
dc.language.isoen_USen_US
dc.publisherIIIT-Delhien_US
dc.subjectAudiobook speech synthesisen_US
dc.subjectText-to-speechen_US
dc.subjectVITSen_US
dc.subjectGigaspeech dataseten_US
dc.subjectModel fine-tuningen_US
dc.titleAudiobooks and emotion detectionen_US
dc.typeOtheren_US
Appears in Collections:Year-2024

Files in This Item:
File Description SizeFormat 
BTP Report - Aniket Gupta.pdf
  Restricted Access
1.47 MBAdobe PDFView/Open Request a copy


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.