A Novel Spike Transformer Network for Depth Estimation from Event Cameras via Cross-modality Knowledge Distillation

Zhang, Xin; Han, Liangxiu; Sobeih, Tam; Han, Lianghao; Dancey, Darren

Computer Science > Computer Vision and Pattern Recognition

arXiv:2404.17335 (cs)

[Submitted on 26 Apr 2024 (v1), last revised 24 Feb 2025 (this version, v3)]

Title:A Novel Spike Transformer Network for Depth Estimation from Event Cameras via Cross-modality Knowledge Distillation

Authors:Xin Zhang, Liangxiu Han, Tam Sobeih, Lianghao Han, Darren Dancey

View PDF HTML (experimental)

Abstract:Depth estimation is a critical task in computer vision, with applications in autonomous navigation, robotics, and augmented reality. Event cameras, which encode temporal changes in light intensity as asynchronous binary spikes, offer unique advantages such as low latency, high dynamic range, and energy efficiency. However, their unconventional spiking output and the scarcity of labelled datasets pose significant challenges to traditional image-based depth estimation methods. To address these challenges, we propose a novel energy-efficient Spike-Driven Transformer Network (SDT) for depth estimation, leveraging the unique properties of spiking data. The proposed SDT introduces three key innovations: (1) a purely spike-driven transformer architecture that incorporates spike-based attention and residual mechanisms, enabling precise depth estimation with minimal energy consumption; (2) a fusion depth estimation head that combines multi-stage features for fine-grained depth prediction while ensuring computational efficiency; and (3) a cross-modality knowledge distillation framework that utilises a pre-trained vision foundation model (DINOv2) to enhance the training of the spiking network despite limited data this http URL work represents the first exploration of transformer-based spiking neural networks for depth estimation, providing a significant step forward in energy-efficient neuromorphic computing for real-world vision applications.

Comments:	16 pages
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2404.17335 [cs.CV]
	(or arXiv:2404.17335v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2404.17335

Submission history

From: Xin Zhang [view email]
[v1] Fri, 26 Apr 2024 11:32:53 UTC (5,316 KB)
[v2] Wed, 1 May 2024 08:54:54 UTC (5,319 KB)
[v3] Mon, 24 Feb 2025 10:47:58 UTC (5,011 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:A Novel Spike Transformer Network for Depth Estimation from Event Cameras via Cross-modality Knowledge Distillation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:A Novel Spike Transformer Network for Depth Estimation from Event Cameras via Cross-modality Knowledge Distillation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators