Enhanced Multimodal Representation Learning with Cross-modal KD

Chen, Mengxi; Xing, Linyu; Wang, Yu; Zhang, Ya

Computer Science > Computer Vision and Pattern Recognition

arXiv:2306.07646 (cs)

[Submitted on 13 Jun 2023]

Title:Enhanced Multimodal Representation Learning with Cross-modal KD

Authors:Mengxi Chen, Linyu Xing, Yu Wang, Ya Zhang

View PDF

Abstract:This paper explores the tasks of leveraging auxiliary modalities which are only available at training to enhance multimodal representation learning through cross-modal Knowledge Distillation (KD). The widely adopted mutual information maximization-based objective leads to a short-cut solution of the weak teacher, i.e., achieving the maximum mutual information by simply making the teacher model as weak as the student model. To prevent such a weak solution, we introduce an additional objective term, i.e., the mutual information between the teacher and the auxiliary modality model. Besides, to narrow down the information gap between the student and teacher, we further propose to minimize the conditional entropy of the teacher given the student. Novel training schemes based on contrastive learning and adversarial learning are designed to optimize the mutual information and the conditional entropy, respectively. Experimental results on three popular multimodal benchmark datasets have shown that the proposed method outperforms a range of state-of-the-art approaches for video recognition, video retrieval and emotion classification.

Comments:	Accepted by CVPR2023
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
Cite as:	arXiv:2306.07646 [cs.CV]
	(or arXiv:2306.07646v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2306.07646

Submission history

From: Mengxi Chen [view email]
[v1] Tue, 13 Jun 2023 09:35:37 UTC (2,623 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Enhanced Multimodal Representation Learning with Cross-modal KD

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Enhanced Multimodal Representation Learning with Cross-modal KD

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators