SMT-DTA: Improving Drug-Target Affinity Prediction with Semi-supervised Multi-task Training

Pei, Qizhi; Wu, Lijun; Zhu, Jinhua; Xia, Yingce; Xia, Shufang; Qin, Tao; Liu, Haiguang; Liu, Tie-Yan

Quantitative Biology > Biomolecules

arXiv:2206.09818v1 (q-bio)

[Submitted on 20 Jun 2022 (this version), latest version 17 Oct 2023 (v3)]

Title:SMT-DTA: Improving Drug-Target Affinity Prediction with Semi-supervised Multi-task Training

Authors:Qizhi Pei, Lijun Wu, Jinhua Zhu, Yingce Xia, Shufang Xia, Tao Qin, Haiguang Liu, Tie-Yan Liu

View PDF

Abstract:Drug-Target Affinity (DTA) prediction is an essential task for drug discovery and pharmaceutical research. Accurate predictions of DTA can greatly benefit the design of new drug. As wet experiments are costly and time consuming, the supervised data for DTA prediction is extremely limited. This seriously hinders the application of deep learning based methods, which require a large scale of supervised data. To address this challenge and improve the DTA prediction accuracy, we propose a framework with several simple yet effective strategies in this work: (1) a multi-task training strategy, which takes the DTA prediction and the masked language modeling (MLM) task on the paired drug-target dataset; (2) a semi-supervised training method to empower the drug and target representation learning by leveraging large-scale unpaired molecules and proteins in training, which differs from previous pre-training and fine-tuning methods that only utilize molecules or proteins in pre-training; and (3) a cross-attention module to enhance the interaction between drug and target representation. Extensive experiments are conducted on three real-world benchmark datasets: BindingDB, DAVIS and KIBA. The results show that our framework significantly outperforms existing methods and achieves state-of-the-art performances, e.g., $0.712$ RMSE on BindingDB IC$_{50}$ measurement with more than $5\%$ improvement than previous best work. In addition, case studies on specific drug-target binding activities, drug feature visualizations, and real-world applications demonstrate the great potential of our work. The code and data are released at this https URL

Subjects:	Biomolecules (q-bio.BM); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2206.09818 [q-bio.BM]
	(or arXiv:2206.09818v1 [q-bio.BM] for this version)
	https://doi.org/10.48550/arXiv.2206.09818

Submission history

From: Lijun Wu [view email]
[v1] Mon, 20 Jun 2022 14:53:25 UTC (1,511 KB)
[v2] Wed, 22 Jun 2022 02:45:34 UTC (1,511 KB)
[v3] Tue, 17 Oct 2023 14:06:07 UTC (462 KB)

Quantitative Biology > Biomolecules

Title:SMT-DTA: Improving Drug-Target Affinity Prediction with Semi-supervised Multi-task Training

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Quantitative Biology > Biomolecules

Title:SMT-DTA: Improving Drug-Target Affinity Prediction with Semi-supervised Multi-task Training

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators