IIIT-Delhi Institutional Repository

Embedding morality in AI systems

Show simple item record

dc.contributor.author Goel, Rahul
dc.contributor.author Mutharaju, Raghava (Advisor)
dc.date.accessioned 2026-08-31T10:39:06Z
dc.date.available 2026-08-31T10:39:06Z
dc.date.issued 2025-07-18
dc.identifier.uri http://repository.iiitd.edu.in/xmlui/handle/123456789/2056
dc.description.abstract The increasing integration of Artificial Intelligence into complex societal domains necessitates a deeper understanding of how to embed ethical reasoning into machines. This project addresses the challenge of quantifying and predicting alignment with established human ethical frameworks. Our methodology involved three key stages. First, we established a granular theoretical framework by defining 15 distinct ethical theories under the three major schools of thought: Consequentialism (α), Deontology (β), and Virtue Ethics (γ). Second, we developed a novel dataset of 450 ethical scenarios, specifically designed to elicit responses corresponding to each of the 15 theories. These scenarios were then annotated using Large Language Models (LLMs) to generate quantitative scores (α, β, γ) representing the relevance of each ethical school to a given case. Finally, we benchmarked a suite of classical machine learning models to predict these ethical alignments from textual features. Two primary experiments were conducted: a multi-output regression task to predict the (α, β, γ) scores and a multi-output classification task to predict the specific ethical school and theory. In the regression task, Linear Regression demonstrated the strongest explanatory power, achieving an R-squared (R2 ) value of 0.5804. For classification, Logistic Regression was the top- performing model, achieving an Exact Match Accuracy of 74.44% and 100% accuracy in identifying the high-level ethical school. This work provides a foundational dataset, a robust evaluation of classical machine learning models for ethical prediction, and confirms that quantitative representations of ethical schools are viable targets for computational modeling. The findings lay the groundwork for future research in building more transparent, auditable, and ethically-aligned AI systems. en_US
dc.language.iso en_US en_US
dc.publisher IIIT-Delhi en_US
dc.subject Computational ethics en_US
dc.subject Machine learning en_US
dc.subject Ethical frameworks en_US
dc.subject AI alignment en_US
dc.subject Consequentialism en_US
dc.title Embedding morality in AI systems en_US
dc.type Other en_US


Files in this item

This item appears in the following Collection(s)

Show simple item record

Search Repository


Advanced Search

Browse

My Account