On Distribution Dependent Sub-Logarithmic Query Time of Learned Indexing

Zeighami, Sepanta; Shahabi, Cyrus

Computer Science > Databases

arXiv:2306.10651 (cs)

[Submitted on 19 Jun 2023]

Title:On Distribution Dependent Sub-Logarithmic Query Time of Learned Indexing

Authors:Sepanta Zeighami, Cyrus Shahabi

View PDF

Abstract:A fundamental problem in data management is to find the elements in an array that match a query. Recently, learned indexes are being extensively used to solve this problem, where they learn a model to predict the location of the items in the array. They are empirically shown to outperform non-learned methods (e.g., B-trees or binary search that answer queries in $O(\log n)$ time) by orders of magnitude. However, success of learned indexes has not been theoretically justified. Only existing attempt shows the same query time of $O(\log n)$, but with a constant factor improvement in space complexity over non-learned methods, under some assumptions on data distribution. In this paper, we significantly strengthen this result, showing that under mild assumptions on data distribution, and the same space complexity as non-learned methods, learned indexes can answer queries in $O(\log\log n)$ expected query time. We also show that allowing for slightly larger but still near-linear space overhead, a learned index can achieve $O(1)$ expected query time. Our results theoretically prove learned indexes are orders of magnitude faster than non-learned methods, theoretically grounding their empirical success.

Comments:	Proceedings of 40th International Conference on Machine Learning, ICML 2023
Subjects:	Databases (cs.DB); Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG)
Cite as:	arXiv:2306.10651 [cs.DB]
	(or arXiv:2306.10651v1 [cs.DB] for this version)
	https://doi.org/10.48550/arXiv.2306.10651

Submission history

From: Sepanta Zeighami [view email]
[v1] Mon, 19 Jun 2023 00:00:37 UTC (626 KB)

Computer Science > Databases

Title:On Distribution Dependent Sub-Logarithmic Query Time of Learned Indexing

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Databases

Title:On Distribution Dependent Sub-Logarithmic Query Time of Learned Indexing

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators