Abstract:
Voice cloning (VC) has emerged as a transformative technology, enabling personalized Human Computer Interaction and enriching user experiences. However, the performance, efficiency, and environmental impact of these models require rigorous assessment and benchmarking. Current benchmarking techniques generally rely on the Mean Opinion Score (MOS), which is highly sub jective and impacted by individual biases. Therefore, this work proposes an automated framework, namely GreenVoice, for quantitative benchmarking of VC models using speaker verification. Green Voice is used to comprehensively benchmark four open-source state-of-the-art (SOTA) VC models based on the naturalness of their generated voice clones, costs, and the environmental impact while inferencing. GreenVoice has been used to investigate the performance of these models on out of-domain test sets to advance the development of inclusive models. Inclusive VC models show equitable performance for users from different demographics, including genders, accents, and low resource languages. Empirical results demonstrate a significant performance degradation of these models on out-of-domain datasets. Our findings reveal a higher naturalness in the generated clones for male speakers than female speakers. Additionally, we emphasize the importance of reporting the environmental impacts, such as carbon emissions, of using large AI models. With the view of having sustainable and inclusive VC technologies, we encourage researchers to follow the Green AI practices and work towards developing environment-friendly and robust speech processing models. Building on this, we further propose a GAN-like framework for fine-tuning any voice cloning model, and present some preliminary results for one of the models.