On the Impact of Hard Adversarial Instances on Overfitting in Adversarial Training

Liu, Chen; Huang, Zhichao; Salzmann, Mathieu; Zhang, Tong; Süsstrunk, Sabine

Computer Science > Machine Learning

arXiv:2112.07324 (cs)

[Submitted on 14 Dec 2021 (v1), last revised 17 Dec 2024 (this version, v2)]

Title:On the Impact of Hard Adversarial Instances on Overfitting in Adversarial Training

Authors:Chen Liu, Zhichao Huang, Mathieu Salzmann, Tong Zhang, Sabine Süsstrunk

View PDF

Abstract:Adversarial training is a popular method to robustify models against adversarial attacks. However, it exhibits much more severe overfitting than training on clean inputs. In this work, we investigate this phenomenon from the perspective of training instances, i.e., training input-target pairs. Based on a quantitative metric measuring the relative difficulty of an instance in the training set, we analyze the model's behavior on training instances of different difficulty levels. This lets us demonstrate that the decay in generalization performance of adversarial training is a result of fitting hard adversarial instances. We theoretically verify our observations for both linear and general nonlinear models, proving that models trained on hard instances have worse generalization performance than ones trained on easy instances, and that this generalization gap increases with the size of the adversarial budget. Finally, we investigate solutions to mitigate adversarial overfitting in several scenarios, including fast adversarial training and fine-tuning a pretrained model with additional data. Our results demonstrate that using training data adaptively improves the model's robustness.

Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2112.07324 [cs.LG]
	(or arXiv:2112.07324v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2112.07324
Journal reference:	Journal of Machine Learning Research, Vol 25, 2024

Submission history

From: Chen Liu [view email]
[v1] Tue, 14 Dec 2021 12:19:24 UTC (12,128 KB)
[v2] Tue, 17 Dec 2024 08:17:26 UTC (15,860 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2021-12

Change to browse by:

References & Citations

DBLP - CS Bibliography

listing | bibtex

Chen Liu
Zhichao Huang
Mathieu Salzmann
Tong Zhang
Sabine Süsstrunk

export BibTeX citation

Computer Science > Machine Learning

Title:On the Impact of Hard Adversarial Instances on Overfitting in Adversarial Training

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:On the Impact of Hard Adversarial Instances on Overfitting in Adversarial Training

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators