Cross-Modal Data Programming Enables Rapid Medical Machine Learning

Dunnmon, Jared; Ratner, Alexander; Khandwala, Nishith; Saab, Khaled; Markert, Matthew; Sagreiya, Hersh; Goldman, Roger; Lee-Messer, Christopher; Lungren, Matthew; Rubin, Daniel; Ré, Christopher

Computer Science > Machine Learning

arXiv:1903.11101 (cs)

[Submitted on 26 Mar 2019]

Title:Cross-Modal Data Programming Enables Rapid Medical Machine Learning

Authors:Jared Dunnmon, Alexander Ratner, Nishith Khandwala, Khaled Saab, Matthew Markert, Hersh Sagreiya, Roger Goldman, Christopher Lee-Messer, Matthew Lungren, Daniel Rubin, Christopher Ré

View PDF

Abstract:Labeling training datasets has become a key barrier to building medical machine learning models. One strategy is to generate training labels programmatically, for example by applying natural language processing pipelines to text reports associated with imaging studies. We propose cross-modal data programming, which generalizes this intuitive strategy in a theoretically-grounded way that enables simpler, clinician-driven input, reduces required labeling time, and improves with additional unlabeled data. In this approach, clinicians generate training labels for models defined over a target modality (e.g. images or time series) by writing rules over an auxiliary modality (e.g. text reports). The resulting technical challenge consists of estimating the accuracies and correlations of these rules; we extend a recent unsupervised generative modeling technique to handle this cross-modal setting in a provably consistent way. Across four applications in radiography, computed tomography, and electroencephalography, and using only several hours of clinician time, our approach matches or exceeds the efficacy of physician-months of hand-labeling with statistical significance, demonstrating a fundamentally faster and more flexible way of building machine learning models in medicine.

Subjects:	Machine Learning (cs.LG); Image and Video Processing (eess.IV); Machine Learning (stat.ML)
Cite as:	arXiv:1903.11101 [cs.LG]
	(or arXiv:1903.11101v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1903.11101

Submission history

From: Jared Dunnmon [view email]
[v1] Tue, 26 Mar 2019 18:12:34 UTC (3,659 KB)

Computer Science > Machine Learning

Title:Cross-Modal Data Programming Enables Rapid Medical Machine Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Cross-Modal Data Programming Enables Rapid Medical Machine Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators