Abstract:
Understanding protein function at the structural level requires accurate representation of their biologically active forms. However, structures deposited in the Protein Data Bank (PDB) often include only the asymmetric unit from crystallographic experiments, leading to discrepancies with the true oligomeric state. Such inconsistencies introduce significant challenges in down- stream analyses, particularly in cavity detection, where inter-subunit voids may be misidentified as functional binding sites. This work aims to address these challenges by constructing a curated dataset of protein struc- tures that accurately represent their functional oligomeric assemblies. PDB files were prepro- cessed to eliminate structural artifacts and to reconstruct biologically relevant assemblies. The cleaned structures were then analyzed using the CICLOP tool to identify cavity-lining residues, from which geometric and residue-level features were extracted. These features will serve as the foundation for machine learning models aimed at predicting protein function based on cavity characteristics. By integrating structural data correction with robust cavity analysis, this research establishes a reliable and interpretable pipeline for protein function prediction, with potential applications in structural bioinformatics and drug discovery.