In search of the most efficient and memory-saving visualization of high dimensional data

Minch, Bartosz

Computer Science > Machine Learning

arXiv:2303.05455 (cs)

[Submitted on 27 Feb 2023]

Title:In search of the most efficient and memory-saving visualization of high dimensional data

Authors:Bartosz Minch

View PDF

Abstract:Interactive exploration of large, multidimensional datasets plays a very important role in various scientific fields. It makes it possible not only to identify important structural features and forms, such as clusters of vertices and their connection patterns, but also to evaluate their interrelationships in terms of position, distance, shape and connection density. We argue that the visualization of multidimensional data is well approximated by the problem of two-dimensional embedding of undirected nearest-neighbor graphs. The size of complex networks is a major challenge for today's computer systems and still requires more efficient data embedding algorithms. Existing reduction methods are too slow and do not allow interactive manipulation. We show that high-quality embeddings are produced with minimal time and memory complexity. We present very efficient IVHD algorithms (CPU and GPU) and compare them with the latest and most popular dimensionality reduction methods. We show that the memory and time requirements are dramatically lower than for base codes. At the cost of a slight degradation in embedding quality, IVHD preserves the main structural properties of the data well with a much lower time budget. We also present a meta-algorithm that allows the use of any unsupervised data embedding method in a supervised manner.

Comments:	PhD thesis on searching the most efficient and memory-saving visualization of high dimensional data. arXiv admin note: substantial text overlap with arXiv:1902.01108, arXiv:1602.00370 by other authors; text overlap with arXiv:2109.02508 by other authors
Subjects:	Machine Learning (cs.LG); Human-Computer Interaction (cs.HC)
Cite as:	arXiv:2303.05455 [cs.LG]
	(or arXiv:2303.05455v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2303.05455

Submission history

From: Bartosz Minch [view email]
[v1] Mon, 27 Feb 2023 20:56:13 UTC (116,114 KB)

Computer Science > Machine Learning

Title:In search of the most efficient and memory-saving visualization of high dimensional data

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:In search of the most efficient and memory-saving visualization of high dimensional data

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators