Impact of Biases in Big Data

Glauner, Patrick; Valtchev, Petko; State, Radu

Computer Science > Machine Learning

arXiv:1803.00897 (cs)

[Submitted on 2 Mar 2018]

Title:Impact of Biases in Big Data

Authors:Patrick Glauner, Petko Valtchev, Radu State

View PDF

Abstract:The underlying paradigm of big data-driven machine learning reflects the desire of deriving better conclusions from simply analyzing more data, without the necessity of looking at theory and models. Is having simply more data always helpful? In 1936, The Literary Digest collected 2.3M filled in questionnaires to predict the outcome of that year's US presidential election. The outcome of this big data prediction proved to be entirely wrong, whereas George Gallup only needed 3K handpicked people to make an accurate prediction. Generally, biases occur in machine learning whenever the distributions of training set and test set are different. In this work, we provide a review of different sorts of biases in (big) data sets in machine learning. We provide definitions and discussions of the most commonly appearing biases in machine learning: class imbalance and covariate shift. We also show how these biases can be quantified and corrected. This work is an introductory text for both researchers and practitioners to become more aware of this topic and thus to derive more reliable models for their learning problems.

Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:1803.00897 [cs.LG]
	(or arXiv:1803.00897v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1803.00897
Journal reference:	Proceedings of the 26th European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN 2018)

Submission history

From: Patrick Glauner [view email]
[v1] Fri, 2 Mar 2018 15:35:18 UTC (304 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2018-03

Change to browse by:

References & Citations

DBLP - CS Bibliography

listing | bibtex

Patrick O. Glauner
Petko Valtchev
Radu State

export BibTeX citation

Computer Science > Machine Learning

Title:Impact of Biases in Big Data

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Impact of Biases in Big Data

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators