Can masking background and object reduce static bias for zero-shot action recognition?

Fukuzawa, Takumi; Hara, Kensho; Kataoka, Hirokatsu; Tamaki, Toru

doi:10.1007/978-981-96-2071-5_27

Computer Science > Computer Vision and Pattern Recognition

arXiv:2501.12681 (cs)

[Submitted on 22 Jan 2025]

Title:Can masking background and object reduce static bias for zero-shot action recognition?

Authors:Takumi Fukuzawa, Kensho Hara, Hirokatsu Kataoka, Toru Tamaki

View PDF HTML (experimental)

Abstract:In this paper, we address the issue of static bias in zero-shot action recognition. Action recognition models need to represent the action itself, not the appearance. However, some fully-supervised works show that models often rely on static appearances, such as the background and objects, rather than human actions. This issue, known as static bias, has not been investigated for zero-shot. Although CLIP-based zero-shot models are now common, it remains unclear if they sufficiently focus on human actions, as CLIP primarily captures appearance features related to languages. In this paper, we investigate the influence of static bias in zero-shot action recognition with CLIP-based models. Our approach involves masking backgrounds, objects, and people differently during training and validation. Experiments with masking background show that models depend on background bias as their performance decreases for Kinetics400. However, for Mimetics, which has a weak background bias, masking the background leads to improved performance even if the background is masked during validation. Furthermore, masking both the background and objects in different colors improves performance for SSv2, which has a strong object bias. These results suggest that masking the background or objects during training prevents models from overly depending on static bias and makes them focus more on human action.

Comments:	In proc. of MMM2025
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2501.12681 [cs.CV]
	(or arXiv:2501.12681v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2501.12681
Journal reference:	MMM2025
Related DOI:	https://doi.org/10.1007/978-981-96-2071-5_27

Submission history

From: Toru Tamaki [view email]
[v1] Wed, 22 Jan 2025 06:59:46 UTC (7,424 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Can masking background and object reduce static bias for zero-shot action recognition?

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Can masking background and object reduce static bias for zero-shot action recognition?

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators