Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation

Xie, Junyu; Han, Tengda; Bain, Max; Nagrani, Arsha; Khandelwal, Eshika; Varol, Gül; Xie, Weidi; Zisserman, Andrew

Computer Science > Computer Vision and Pattern Recognition

arXiv:2504.01020 (cs)

[Submitted on 1 Apr 2025]

Title:Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation

Authors:Junyu Xie, Tengda Han, Max Bain, Arsha Nagrani, Eshika Khandelwal, Gül Varol, Weidi Xie, Andrew Zisserman

View PDF HTML (experimental)

Abstract:Our objective is the automatic generation of Audio Descriptions (ADs) for edited video material, such as movies and TV series. To achieve this, we propose a two-stage framework that leverages "shots" as the fundamental units of video understanding. This includes extending temporal context to neighbouring shots and incorporating film grammar devices, such as shot scales and thread structures, to guide AD generation. Our method is compatible with both open-source and proprietary Visual-Language Models (VLMs), integrating expert knowledge from add-on modules without requiring additional training of the VLMs. We achieve state-of-the-art performance among all prior training-free approaches and even surpass fine-tuned methods on several benchmarks. To evaluate the quality of predicted ADs, we introduce a new evaluation measure -- an action score -- specifically targeted to assessing this important aspect of AD. Additionally, we propose a novel evaluation protocol that treats automatic frameworks as AD generation assistants and asks them to generate multiple candidate ADs for selection.

Comments:	Project Page: this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2504.01020 [cs.CV]
	(or arXiv:2504.01020v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2504.01020

Submission history

From: Junyu Xie [view email]
[v1] Tue, 1 Apr 2025 17:59:57 UTC (25,653 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators