| dc.description.abstract |
Recent advancements in text-to-speech (TTS) technology have made synthetic speech sound more natural and expressive. However, creating audiobook narration that accurately captures emotions and maintains context over multiple sentences is still a challenge. In this thesis, we present a new approach to audiobook speech synthesis using the VITS model, which is known for its stability and high quality, along with the Gigaspeech dataset. The Gigaspeech dataset is particularly useful for this task because it includes voice and text samples from real audiobooks, covering both dialogue and narration parts. A key contribution of our work is the emotional labeling of the Gigaspeech dataset, where each entry is tagged with specific emotions. Additionally, we propose using large language models (LLMs) to automatically determine the emotion and speaker ID for each sentence. This allows the TTS model to generate speech that not only fits the context but also conveys the appropriate emotions. We plan to achieve this as one of our future goal by fine-tuning the decoder of the VITS model using the emotionally labeled Gigaspeech dataset. Our experimental findings currently on smaller subsets of our data indicate that our text-to- speech system reliably generates natural audiobook narration, successfully capturing the unique speaking style of individual narrators. |
en_US |