CHG Shapley: Efficient Data Valuation and Selection towards Trustworthy Machine Learning

Cai, Huaiguang

Computer Science > Computer Science and Game Theory

arXiv:2406.11730 (cs)

[Submitted on 17 Jun 2024 (v1), last revised 18 Jun 2024 (this version, v2)]

Title:CHG Shapley: Efficient Data Valuation and Selection towards Trustworthy Machine Learning

Authors:Huaiguang Cai

View PDF HTML (experimental)

Abstract:Understanding the decision-making process of machine learning models is crucial for ensuring trustworthy machine learning. Data Shapley, a landmark study on data valuation, advances this understanding by assessing the contribution of each datum to model accuracy. However, the resource-intensive and time-consuming nature of multiple model retraining poses challenges for applying Data Shapley to large datasets. To address this, we propose the CHG (Conduct of Hardness and Gradient) score, which approximates the utility of each data subset on model accuracy during a single model training. By deriving the closed-form expression of the Shapley value for each data point under the CHG score utility function, we reduce the computational complexity to the equivalent of a single model retraining, an exponential improvement over existing methods. Additionally, we employ CHG Shapley for real-time data selection, demonstrating its effectiveness in identifying high-value and noisy data. CHG Shapley facilitates trustworthy model training through efficient data valuation, introducing a novel data-centric perspective on trustworthy machine learning.

Subjects:	Computer Science and Game Theory (cs.GT); Machine Learning (cs.LG)
Cite as:	arXiv:2406.11730 [cs.GT]
	(or arXiv:2406.11730v2 [cs.GT] for this version)
	https://doi.org/10.48550/arXiv.2406.11730

Submission history

From: Huaiguang Cai [view email]
[v1] Mon, 17 Jun 2024 16:48:31 UTC (978 KB)
[v2] Tue, 18 Jun 2024 07:38:31 UTC (978 KB)

Computer Science > Computer Science and Game Theory

Title:CHG Shapley: Efficient Data Valuation and Selection towards Trustworthy Machine Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Science and Game Theory

Title:CHG Shapley: Efficient Data Valuation and Selection towards Trustworthy Machine Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators