Please use this identifier to cite or link to this item: http://repository.iiitd.edu.in/xmlui/handle/123456789/2048
Title: On recovering fair and accurate classifiers with imperfect distributions
Authors: Sharma, Mohit
Shah, Rajiv Ratn (Advisor)
Deshpande, Amit (Advisor)
Keywords: Machine learning
Bayes Optimal Classific
Fairness Constraints
Issue Date: Aug-2026
Publisher: IIIT-Delhi
Abstract: In the context of classification, the Bayes optimal classifier represents the best possible classification rule with respect to the 0-1 loss, for a given data distribution. In the setting of binary classification, it can often be expressed as a thresholding rule over the instance-dependent class posterior probability. Prior research works have studied the best classification performance and the corresponding Bayes optimal classification rule under fairness constraints, which can again be expressed as a group/instance- dependent thresholding rule. However, it is often observed that fairness and accuracy of a model are often at odds and exhibit a tradeoff. This tradeoff depends on the distribution and the fairness metric of interest, and understanding its cause and effects is crucial for real-world deployments. In this thesis, we study this tradeoff from several angles. One of the most important tools for all of our studies involves exploring the Bayes optimal fair classifiers to study the effects of data biases and derive theoretical guarantees. We first re-examine the critical stance on the fairness-accuracy tradeoff and instead study under what conditions fairness constraints can eventually help recover the uncon- strained Bayes optimal classifier, especially when we can characterize the factors of data bias in our given distribution. Next, we highlight how the Bayes optimal fair classification rule can be implemented at various stages of classification: Either as a pre-processing of the given distribution, a weighted risk minimization, or as a post-processing (thresholding) of the class posterior probability. We then show, via simulations on widely used fair classifiers, that with varying amounts of data bias, the theoretical equivalence does not translate into practice and, in fact, sometimes even fails to mitigate unfairness. Re-examining the tradeoff phenomena, we then ask whether it is possible to steer distributions most minimally and efficiently towards an ideal distribution, where the Bayes optimal classifiers are always fair by default. We mathematically characterize what an ideal distribution may look like when the distribution can be expressed with some parametric family (e.g., Gaussian distributions). Finally, we estimate the fairness-accuracy tradeoff using imperfect data access, i.e., in a model-free setting, only using class posterior probabilities and without access to features. In the spirit of prior works that use Bayes error estimation techniques to benchmark classification performance, we study how we can estimate and characterize the fairness-accuracy tradeoff for any distribution only using soft labels (class posterior probabilities). This allows us to compare various fairness-accuracy tradeoff estimation techniques in the literature against the theoretically best and serves as an effective bench- marking framework for future work in this space.
URI: http://repository.iiitd.edu.in/xmlui/handle/123456789/2048
Appears in Collections:Year-2026

Files in This Item:
File Description SizeFormat 
final_thesis_print_v2_sign_2.pdf13.68 MBAdobe PDFView/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.