Abstract:
Predicting lipid-binding proteins is essential for understanding their roles in cellular functions and disease mechanisms. This research aims to develop a machine learning (ML) model to accurately identify lipid-binding proteins by leveraging structural biology data. Data was sourced from the RCSB Protein Data Bank (PDB), UniProt, and various research publications, resulting in an extensive database with features including UniProt IDs, structural angles, molecular dimensions, sequence lengths, and lipid-binding categories. The dataset was meticulously cleaned by removing duplicates and verified for binding sites using PDB files through keyword searches and structural analysis. Comprehensive mapping of proteins to their respective lipid types was achieved by reviewing research articles and utilizing specialized databases. This curated dataset lays the groundwork for building an ML model, which is expected to enhance the precision of lipid-binding protein predictions. The anticipated model will facilitate advancements in drug discovery and deepen our understanding of lipid-related diseases, demonstrating the significant potential of integrating structural biology with computational approaches in biomedical research.