Abstract:
The Indo-Aryan language family, spoken by millions across the Hindi Belt, is a vital part of India’s culture and heritage. However, many of its dialects, like Awadhi, Bhojpuri, Braj, and Magahi, are largely ignored in modern language technologies. These dialects lack proper resources, making it difficult to create tools that support them. As a result, they risk being left behind in a world that increasingly depends on digital communication dominated by languages like Modern Standard Hindi (MSH). Our research addresses this gap by focusing on Speech-to-Speech Machine Translation (S2ST) be tween Modern Standard Hindi (MSH) and these dialects. We aim to design a bidirectional S2ST system that can handle low-resource scenarios with limited parallel data. To achieve this, we plan to use a cascaded S2ST approach, combining Automatic Speech Recognition (ASR), Machine Transla tion (MT), and Text-to-Speech (TTS) modules. Additionally, we will explore direct S2ST methods, leveraging recent advancements in unit-to-unit translation using acoustic discrete units to enhance translation for low-resource languages. For our research, we utilized a dataset curated by Dr. Bhimrao Ambedkar University as part of the SpeeD-IA project. This dataset provides parallel speech and transcription data for four dialects—Awadhi, Bhojpuri, Braj, and Magahi. Additionally, we plan to record corresponding Hindi audios. It comprises 369 Hindi sentences translated into the respective dialects and spoken by 10 speakers per dialect, resulting in approximately 8 hours of audio data. However, due to repeated and shuffled sentences, the dataset required extensive manual alignment and mapping to ensure accuracy. Our work is inspired by the urgent need to preserve linguistic diversity and ensure that speakers of underrepresented dialects are not excluded from technological advancements. By addressing the challenge of lack of established benchmarks for these Hindi dialects—this research highlights the potential of S2ST systems to foster inclusivity and accessibility. The resulting model not only paves the way for greater inclusivity and accessibility for speakers of low-resource dialects but also establishes a scalable and adaptable framework for extending speech technologies to other marginalized languages. This work contributes to a more equitable digital ecosystem, ensuring that linguistic heritage and diversity are upheld in the modern age.