Show simple item record

dc.contributor.author Singh, Amanpreet
dc.contributor.author Daipuria, Aditya
dc.contributor.author Sethi, Tavpritesh (Advisor)
dc.date.accessioned 2026-09-18T07:00:06Z
dc.date.available 2026-09-18T07:00:06Z
dc.date.issued 2024-07-25
dc.identifier.uri http://repository.iiitd.edu.in/xmlui/handle/123456789/2166
dc.description.abstract This project focuses on the SNOMED-CT mapping of over 9000 cancer-related datasets from Figshare and more than 2 million datasets from Zenodo FAIR Stations. The data was extracted from online sources using custom scripting, followed by the implementation of a multi-stage an alytical pipeline. This pipeline encompassed data cleaning, preprocessing, K-means clustering, word cloud generation, calculation of Jaccard and Kullback-Leibler divergence, and Bayesian Network modeling. Additionally, the BERT Sentence Model was utilized for calculating cosine scores to enhance the analysis. This comprehensive approach aimed to improve the accuracy and efficiency of SNOMED-CT mapping in cancer research, facilitating better data integration and interoperability within the medical and research communities. Our results demonstrate the effectiveness of these methods in handling large-scale datasets and providing valuable insights into cancer-related data. en_US
dc.language.iso en_US en_US
dc.publisher IIIT-Delhi en_US
dc.subject SNOMED-CT mapping en_US
dc.subject Cancer-related datasets en_US
dc.subject Figshare en_US
dc.subject Word cloud en_US
dc.subject Data Integration en_US
dc.title IFHP- tindering dataset en_US
dc.type Other en_US


Files in this item

This item appears in the following Collection(s)

Show simple item record

Search Repository


Advanced Search

Browse

My Account