IIIT-Delhi Institutional Repository

Speech-to-speech machine translation for Hindi dialects

Show simple item record

dc.contributor.author Sood, Sanmay
dc.contributor.author Rajput, Siddhart
dc.contributor.author Akhtar, Md. Shad (Advisor)
dc.date.accessioned 2026-09-05T09:46:45Z
dc.date.available 2026-09-05T09:46:45Z
dc.date.issued 2024-11-27
dc.identifier.uri http://repository.iiitd.edu.in/xmlui/handle/123456789/2098
dc.description.abstract The Indo-Aryan language family, spoken by millions across the Hindi Belt, is a vital part of India’s culture and heritage. However, many of its dialects, like Awadhi, Bhojpuri, Braj, and Magahi, are largely ignored in modern language technologies. These dialects lack proper resources, making it difficult to create tools that support them. As a result, they risk being left behind in a world that increasingly depends on digital communication dominated by languages like Modern Standard Hindi (MSH). Our research addresses this gap by focusing on Speech-to-Speech Machine Translation (S2ST) be tween Modern Standard Hindi (MSH) and these dialects. We aim to design a bidirectional S2ST system that can handle low-resource scenarios with limited parallel data. To achieve this, we plan to use a cascaded S2ST approach, combining Automatic Speech Recognition (ASR), Machine Transla tion (MT), and Text-to-Speech (TTS) modules. Additionally, we will explore direct S2ST methods, leveraging recent advancements in unit-to-unit translation using acoustic discrete units to enhance translation for low-resource languages. For our research, we utilized a dataset curated by Dr. Bhimrao Ambedkar University as part of the SpeeD-IA project. This dataset provides parallel speech and transcription data for four dialects—Awadhi, Bhojpuri, Braj, and Magahi. Additionally, we plan to record corresponding Hindi audios. It comprises 369 Hindi sentences translated into the respective dialects and spoken by 10 speakers per dialect, resulting in approximately 8 hours of audio data. However, due to repeated and shuffled sentences, the dataset required extensive manual alignment and mapping to ensure accuracy. Our work is inspired by the urgent need to preserve linguistic diversity and ensure that speakers of underrepresented dialects are not excluded from technological advancements. By addressing the challenge of lack of established benchmarks for these Hindi dialects—this research highlights the potential of S2ST systems to foster inclusivity and accessibility. The resulting model not only paves the way for greater inclusivity and accessibility for speakers of low-resource dialects but also establishes a scalable and adaptable framework for extending speech technologies to other marginalized languages. This work contributes to a more equitable digital ecosystem, ensuring that linguistic heritage and diversity are upheld in the modern age. en_US
dc.language.iso en_US en_US
dc.publisher IIIT-Delhi en_US
dc.subject Speech-to-Speech Machine Translation en_US
dc.subject Low-Resource Languages en_US
dc.subject Hindi Belt Dialects en_US
dc.subject Acoustic Discrete Unit en_US
dc.title Speech-to-speech machine translation for Hindi dialects en_US
dc.type Other en_US


Files in this item

This item appears in the following Collection(s)

Show simple item record

Search Repository


Advanced Search

Browse

My Account