Abstract:
This project focuses on the SNOMED-CT mapping of over 9000 cancer-related datasets from Figshare and more than 2 million datasets from Zenodo FAIR Stations. The data was extracted from online sources using custom scripting, followed by the implementation of a multi-stage an alytical pipeline. This pipeline encompassed data cleaning, preprocessing, K-means clustering, word cloud generation, calculation of Jaccard and Kullback-Leibler divergence, and Bayesian Network modeling. Additionally, the BERT Sentence Model was utilized for calculating cosine scores to enhance the analysis. This comprehensive approach aimed to improve the accuracy and efficiency of SNOMED-CT mapping in cancer research, facilitating better data integration and interoperability within the medical and research communities. Our results demonstrate the effectiveness of these methods in handling large-scale datasets and providing valuable insights into cancer-related data.