| dc.description.abstract |
The way kinases interact with their substrates has been a longstanding problem in biology prediction and important to drug developing. This report is about a new computational system like Multimodal Siamese Graph Network that is helpful in predicting the ability of kinase substrate engagement effectively, using different data modalities. Our model is on both sequence and structural based features. The structural information in PDB files has been mapped into a graphical form and analyzed by structural features (binding pockets and secondary structures) along with sequence information as pre-processed structural information (including ESM-2 protein language model embeddings). Siamese network architecture is learning to differentiate among interacting pairs, and non-interacting pairs using a triplet Margin Loss syntax. On a held out validation set our model achieves a state of the art Area Under the Precision-Recall Curve (AUPRC) and Area Under the Receiver Operating Characteristic (AUROC) of 0.785 and 0.786, respectively, far outpacing the baseline models, including Random Forest and XGBoost. Impressively, the model is unsupervised in that the kinases in their existing biological families are clustered which implies that the network is learning biologically important representations. The application proves not only the effectiveness of the combination of deep sequence, structural information to model molecular recognition but also represents a solid foundation towards the future in silico discoveries of the phosphorylation phenomena by means of generative models |
en_US |