Learning to Abstain From Uninformative Data

Zhang, Yikai; Zheng, Songzhu; Dalirrooyfard, Mina; Wu, Pengxiang; Schneider, Anderson; Raj, Anant; Nevmyvaka, Yuriy; Chen, Chao

Computer Science > Machine Learning

arXiv:2309.14240 (cs)

[Submitted on 25 Sep 2023]

Title:Learning to Abstain From Uninformative Data

Authors:Yikai Zhang, Songzhu Zheng, Mina Dalirrooyfard, Pengxiang Wu, Anderson Schneider, Anant Raj, Yuriy Nevmyvaka, Chao Chen

View PDF

Abstract:Learning and decision-making in domains with naturally high noise-to-signal ratio, such as Finance or Healthcare, is often challenging, while the stakes are very high. In this paper, we study the problem of learning and acting under a general noisy generative process. In this problem, the data distribution has a significant proportion of uninformative samples with high noise in the label, while part of the data contains useful information represented by low label noise. This dichotomy is present during both training and inference, which requires the proper handling of uninformative data during both training and testing. We propose a novel approach to learning under these conditions via a loss inspired by the selective learning theory. By minimizing this loss, the model is guaranteed to make a near-optimal decision by distinguishing informative data from uninformative data and making predictions. We build upon the strength of our theoretical guarantees by describing an iterative algorithm, which jointly optimizes both a predictor and a selector, and evaluates its empirical performance in a variety of settings.

Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2309.14240 [cs.LG]
	(or arXiv:2309.14240v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2309.14240

Submission history

From: Songzhu Zheng [view email]
[v1] Mon, 25 Sep 2023 15:55:55 UTC (3,027 KB)

Computer Science > Machine Learning

Title:Learning to Abstain From Uninformative Data

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Learning to Abstain From Uninformative Data

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators