Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach

Khindkar, Vaishnavi; Balasubramanian, Vineeth; Arora, Chetan; Subramanian, Anbumani; Jawahar, C. V.

Computer Science > Computer Vision and Pattern Recognition

arXiv:2411.13302 (cs)

[Submitted on 20 Nov 2024]

Title:Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach

Authors:Vaishnavi Khindkar, Vineeth Balasubramanian, Chetan Arora, Anbumani Subramanian, C.V. Jawahar

View PDF HTML (experimental)

Abstract:With the increased importance of autonomous navigation systems has come an increasing need to protect the safety of Vulnerable Road Users (VRUs) such as pedestrians. Predicting pedestrian intent is one such challenging task, where prior work predicts the binary cross/no-cross intention with a fusion of visual and motion features. However, there has been no effort so far to hedge such predictions with human-understandable reasons. We address this issue by introducing a novel problem setting of exploring the intuitive reasoning behind a pedestrian's intent. In particular, we show that predicting the 'WHY' can be very useful in understanding the 'WHAT'. To this end, we propose a novel, reason-enriched PIE++ dataset consisting of multi-label textual explanations/reasons for pedestrian intent. We also introduce a novel multi-task learning framework called MINDREAD, which leverages a cross-modal representation learning framework for predicting pedestrian intent as well as the reason behind the intent. Our comprehensive experiments show significant improvement of 5.6% and 7% in accuracy and F1-score for the task of intent prediction on the PIE++ dataset using MINDREAD. We also achieved a 4.4% improvement in accuracy on a commonly used JAAD dataset. Extensive evaluation using quantitative/qualitative metrics and user studies shows the effectiveness of our approach.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2411.13302 [cs.CV]
	(or arXiv:2411.13302v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2411.13302

Submission history

From: Vaishnavi Khindkar [view email]
[v1] Wed, 20 Nov 2024 13:15:04 UTC (2,325 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Can Reasons Help Improve Pedestrian Intent Estimation? A Cross-Modal Approach

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators