Learning Object States from Actions via Large Language Models

Tateno, Masatoshi; Yagi, Takuma; Furuta, Ryosuke; Sato, Yoichi

Computer Science > Computer Vision and Pattern Recognition

arXiv:2405.01090 (cs)

[Submitted on 2 May 2024]

Title:Learning Object States from Actions via Large Language Models

Authors:Masatoshi Tateno, Takuma Yagi, Ryosuke Furuta, Yoichi Sato

View PDF HTML (experimental)

Abstract:Temporally localizing the presence of object states in videos is crucial in understanding human activities beyond actions and objects. This task has suffered from a lack of training data due to object states' inherent ambiguity and variety. To avoid exhaustive annotation, learning from transcribed narrations in instructional videos would be intriguing. However, object states are less described in narrations compared to actions, making them less effective. In this work, we propose to extract the object state information from action information included in narrations, using large language models (LLMs). Our observation is that LLMs include world knowledge on the relationship between actions and their resulting object states, and can infer the presence of object states from past action sequences. The proposed LLM-based framework offers flexibility to generate plausible pseudo-object state labels against arbitrary categories. We evaluate our method with our newly collected Multiple Object States Transition (MOST) dataset including dense temporal annotation of 60 object state categories. Our model trained by the generated pseudo-labels demonstrates significant improvement of over 29% in mAP against strong zero-shot vision-language models, showing the effectiveness of explicitly extracting object state information from actions through LLMs.

Comments:	19 pages of main content, 24 pages of supplementary material
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2405.01090 [cs.CV]
	(or arXiv:2405.01090v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2405.01090

Submission history

From: Masatoshi Tateno [view email]
[v1] Thu, 2 May 2024 08:43:16 UTC (7,662 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Learning Object States from Actions via Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Learning Object States from Actions via Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators