Untangling tradeoffs between recurrence and self-attention in neural networks

Kerg, Giancarlo; Kanuparthi, Bhargav; Goyal, Anirudh; Goyette, Kyle; Bengio, Yoshua; Lajoie, Guillaume

Computer Science > Machine Learning

arXiv:2006.09471 (cs)

[Submitted on 16 Jun 2020 (v1), last revised 10 Dec 2020 (this version, v2)]

Title:Untangling tradeoffs between recurrence and self-attention in neural networks

Authors:Giancarlo Kerg, Bhargav Kanuparthi, Anirudh Goyal, Kyle Goyette, Yoshua Bengio, Guillaume Lajoie

View PDF

Abstract:Attention and self-attention mechanisms, are now central to state-of-the-art deep learning on sequential tasks. However, most recent progress hinges on heuristic approaches with limited understanding of attention's role in model optimization and computation, and rely on considerable memory and computational resources that scale poorly. In this work, we present a formal analysis of how self-attention affects gradient propagation in recurrent networks, and prove that it mitigates the problem of vanishing gradients when trying to capture long-term dependencies by establishing concrete bounds for gradient norms. Building on these results, we propose a relevancy screening mechanism, inspired by the cognitive process of memory consolidation, that allows for a scalable use of sparse self-attention with recurrence. While providing guarantees to avoid vanishing gradients, we use simple numerical experiments to demonstrate the tradeoffs in performance and computational resources by efficiently balancing attention and recurrence. Based on our results, we propose a concrete direction of research to improve scalability of attentive networks.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:2006.09471 [cs.LG]
	(or arXiv:2006.09471v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2006.09471

Submission history

From: Giancarlo Kerg [view email]
[v1] Tue, 16 Jun 2020 19:24:25 UTC (1,893 KB)
[v2] Thu, 10 Dec 2020 09:58:29 UTC (1,892 KB)

Computer Science > Machine Learning

Title:Untangling tradeoffs between recurrence and self-attention in neural networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Untangling tradeoffs between recurrence and self-attention in neural networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators