Vanilla Gradient Descent for Oblique Decision Trees

Panda, Subrat Prasad; Genest, Blaise; Easwaran, Arvind; Suganthan, Ponnuthurai Nagaratnam

Computer Science > Machine Learning

arXiv:2408.09135v2 (cs)

[Submitted on 17 Aug 2024 (v1), revised 22 Aug 2024 (this version, v2), latest version 15 Oct 2024 (v3)]

Title:Vanilla Gradient Descent for Oblique Decision Trees

Authors:Subrat Prasad Panda, Blaise Genest, Arvind Easwaran, Ponnuthurai Nagaratnam Suganthan

View PDF HTML (experimental)

Abstract:Decision Trees (DTs) constitute one of the major highly non-linear AI models, valued, e.g., for their efficiency on tabular data. Learning accurate DTs is, however, complicated, especially for oblique DTs, and does take a significant training time. Further, DTs suffer from overfitting, e.g., they proverbially "do not generalize" in regression tasks. Recently, some works proposed ways to make (oblique) DTs differentiable. This enables highly efficient gradient-descent algorithms to be used to learn DTs. It also enables generalizing capabilities by learning regressors at the leaves simultaneously with the decisions in the tree. Prior approaches to making DTs differentiable rely either on probabilistic approximations at the tree's internal nodes (soft DTs) or on approximations in gradient computation at the internal node (quantized gradient descent). In this work, we propose DTSemNet, a novel semantically equivalent and invertible encoding for (hard, oblique) DTs as Neural Networks (NNs), that uses standard vanilla gradient descent. Experiments across various classification and regression benchmarks show that oblique DTs learned using DTSemNet are more accurate than oblique DTs of similar size learned using state-of-the-art techniques. Further, DT training time is significantly reduced. We also experimentally demonstrate that DTSemNet can learn DT policies as efficiently as NN policies in the Reinforcement Learning (RL) setup with physical inputs (dimensions $\leq32$). The code is available at {\color{blue}\textit{\url{this https URL}}}.

Comments:	Published in ECAI-2024. Full version (includes supplementary material)
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2408.09135 [cs.LG]
	(or arXiv:2408.09135v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2408.09135

Submission history

From: Subrat Panda [view email]
[v1] Sat, 17 Aug 2024 08:18:40 UTC (518 KB)
[v2] Thu, 22 Aug 2024 03:28:39 UTC (518 KB)
[v3] Tue, 15 Oct 2024 12:58:35 UTC (519 KB)

Computer Science > Machine Learning

Title:Vanilla Gradient Descent for Oblique Decision Trees

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Vanilla Gradient Descent for Oblique Decision Trees

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators