Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing

Martinez-Ferrer, Pedro J.; Yzelman, Albert-Jan; Beltran, Vicenç

doi:10.1016/j.future.2024.107698

Computer Science > Distributed, Parallel, and Cluster Computing

arXiv:2501.03121 (cs)

[Submitted on 6 Jan 2025 (v1), last revised 25 Feb 2025 (this version, v2)]

Title:Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing

Authors:Pedro J. Martinez-Ferrer, Albert-Jan Yzelman, Vicenç Beltran

View PDF HTML (experimental)

Abstract:The tensor-vector contraction (TVC) is the most memory-bound operation of its class and a core component of the higher-order power method (HOPM). This paper brings distributed-memory parallelization to a native TVC algorithm for dense tensors that overall remains oblivious to contraction mode, tensor splitting and tensor order. Similarly, we propose a novel distributed HOPM, namely dHOPM3, that can save up to one order of magnitude of streamed memory and is about twice as costly in terms of data movement as a distributed TVC operation (dTVC) when using task-based parallelization. The numerical experiments carried out in this work on three different architectures featuring multi-core and accelerators confirm that the performances of dTVC and dHOPM3 remain relatively close to the peak system memory bandwidth (50%-80%, depending on the architecture) and on par with STREAM benchmark figures. On strong scalability scenarios, our native multi-core implementations of these two algorithms can achieve similar and sometimes even greater performance figures than those based upon state-of-the-art CUDA batched kernels. Finally, we demonstrate that both computation and communication can benefit from mixed precision arithmetic also in cases where the hardware does not support low precision data types natively.

Comments:	16 pages, 7 figures, 3 tables, accepted manuscript at Journal of Future Generation Computer Systems
Subjects:	Distributed, Parallel, and Cluster Computing (cs.DC)
MSC classes:	15-04
ACM classes:	I.6.3; G.1.3
Cite as:	arXiv:2501.03121 [cs.DC]
	(or arXiv:2501.03121v2 [cs.DC] for this version)
	https://doi.org/10.48550/arXiv.2501.03121
Journal reference:	Future Generation Computer Systems (166) 107698, May 2025
Related DOI:	https://doi.org/10.1016/j.future.2024.107698

Submission history

From: Pedro J. Martinez-Ferrer [view email]
[v1] Mon, 6 Jan 2025 16:29:19 UTC (489 KB)
[v2] Tue, 25 Feb 2025 14:56:49 UTC (482 KB)

Computer Science > Distributed, Parallel, and Cluster Computing

Title:Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Distributed, Parallel, and Cluster Computing

Title:Distributed and heterogeneous tensor-vector contraction algorithms for high performance computing

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators