MolCAP: Molecular Chemical reActivity pretraining and prompted-finetuning enhanced molecular representation learning

Wang, Yu; Zhang, JingJie; Jin, Junru; Wei, Leyi

Quantitative Biology > Biomolecules

arXiv:2306.09187 (q-bio)

[Submitted on 13 Jun 2023]

Title:MolCAP: Molecular Chemical reActivity pretraining and prompted-finetuning enhanced molecular representation learning

Authors:Yu Wang, JingJie Zhang, Junru Jin, Leyi Wei

View PDF

Abstract:Molecular representation learning (MRL) is a fundamental task for drug discovery. However, previous deep-learning (DL) methods focus excessively on learning robust inner-molecular representations by mask-dominated pretraining framework, neglecting abundant chemical reactivity molecular relationships that have been demonstrated as the determining factor for various molecular property prediction tasks. Here, we present MolCAP to promote MRL, a graph pretraining Transformer based on chemical reactivity (IMR) knowledge with prompted finetuning. Results show that MolCAP outperforms comparative methods based on traditional molecular pretraining framework, in 13 publicly available molecular datasets across a diversity of biomedical tasks. Prompted by MolCAP, even basic graph neural networks are capable of achieving surprising performance that outperforms previous models, indicating the promising prospect of applying reactivity information for MRL. In addition, manual designed molecular templets are potential to uncover the dataset bias. All in all, we expect our MolCAP to gain more chemical meaningful insights for the entire process of drug discovery.

Subjects:	Biomolecules (q-bio.BM); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2306.09187 [q-bio.BM]
	(or arXiv:2306.09187v1 [q-bio.BM] for this version)
	https://doi.org/10.48550/arXiv.2306.09187

Submission history

From: Leyi Wei [view email]
[v1] Tue, 13 Jun 2023 13:48:06 UTC (808 KB)

Quantitative Biology > Biomolecules

Title:MolCAP: Molecular Chemical reActivity pretraining and prompted-finetuning enhanced molecular representation learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Quantitative Biology > Biomolecules

Title:MolCAP: Molecular Chemical reActivity pretraining and prompted-finetuning enhanced molecular representation learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators