Please use this identifier to cite or link to this item: http://repository.iiitd.edu.in/xmlui/handle/123456789/1935
Title: Cost bias mitigation of audio deepfake detection
Authors: Kandpal, Sarthak
Gupta, Anubha (Advisor)
Keywords: Voice Conversion
Deepfake Detection
Equal Error Rate
Issue Date: 8-Dec-2024
Publisher: IIIT-Delhi
Abstract: With the rapid development and advance in the field of Speech Synthesis technology, differentiation of genuine and fake audios has become increasingly challenging. This semester the project focuses on benchmarking the voice conversion model as a preparatory step to create a dataset aimed at cost biased mitigation of audio deepfake detection but due to many constraints and inefficiency to finetune the model only one model was benchmarked correctly. To demonstrate progress and contribute meaningfully a dataset of 20.596 utterances was proposed named Kalpvani using the benchmark model. A user study was conducted where they were present with 6 fake and 6 real audios and evaluated cloned audio through subjective analysis. Participants were also asked to assess whether the given cloned audio is close to source audio or the target audio. Furthermore, speaker verification systems like Ecapa TDNN and Resnet TDNN were used to calculate Equal Error Rates (EER) for target-clone and source-clone pairs, providing an objective evaluation of voice similarity. This benchmarking lays the foundation for future work in cost bias mitigation of audio deep fake detection.
URI: http://repository.iiitd.edu.in/xmlui/handle/123456789/1935
Appears in Collections:Year-2024

Files in This Item:
File Description SizeFormat 
Btech Project Report - Sarthak Kandpal.pdf
  Restricted Access
673.22 kBAdobe PDFView/Open Request a copy


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.