Clean Evaluations on Contaminated Visual Language Models

Lu, Hongyuan; Miao, Shujie; Lam, Wai

Computer Science > Computer Vision and Pattern Recognition

arXiv:2410.07030 (cs)

[Submitted on 9 Oct 2024]

Title:Clean Evaluations on Contaminated Visual Language Models

Authors:Hongyuan Lu, Shujie Miao, Wai Lam

View PDF HTML (experimental)

Abstract:How to evaluate large language models (LLMs) cleanly has been established as an important research era to genuinely report the performance of possibly contaminated LLMs. Yet, how to cleanly evaluate the visual language models (VLMs) is an under-studied problem. We propose a novel approach to achieve such goals through data augmentation methods on the visual input information. We then craft a new visual clean evaluation benchmark with thousands of data instances. Through extensive experiments, we found that the traditional visual data augmentation methods are useful, but they are at risk of being used as a part of the training data as a workaround. We further propose using BGR augmentation to switch the colour channel of the visual information. We found that it is a simple yet effective method for reducing the effect of data contamination and fortunately, it is also harmful to be used as a data augmentation method during training. It means that it is hard to integrate such data augmentation into training by malicious trainers and it could be a promising technique to cleanly evaluate visual LLMs. Our code, data, and model weights will be released upon publication.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
Cite as:	arXiv:2410.07030 [cs.CV]
	(or arXiv:2410.07030v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2410.07030

Submission history

From: Hongyuan Lu [view email]
[v1] Wed, 9 Oct 2024 16:13:19 UTC (29,074 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Clean Evaluations on Contaminated Visual Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Clean Evaluations on Contaminated Visual Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators