Computer Science > Machine Learning

arXiv:2410.01686 (cs)

[Submitted on 2 Oct 2024]

Title:Positional Attention: Out-of-Distribution Generalization and Expressivity for Neural Algorithmic Reasoning

Authors:Artur Back de Luca, George Giapitzakis, Shenghao Yang, Petar Veličković, Kimon Fountoulakis

Abstract:There has been a growing interest in the ability of neural networks to solve algorithmic tasks, such as arithmetic, summary statistics, and sorting. While state-of-the-art models like Transformers have demonstrated good generalization performance on in-distribution tasks, their out-of-distribution (OOD) performance is poor when trained end-to-end. In this paper, we focus on value generalization, a common instance of OOD generalization where the test distribution has the same input sequence length as the training distribution, but the value ranges in the training and test distributions do not necessarily overlap. To address this issue, we propose that using fixed positional encodings to determine attention weights-referred to as positional attention-enhances empirical OOD performance while maintaining expressivity. We support our claim about expressivity by proving that Transformers with positional attention can effectively simulate parallel algorithms.

Comments:	37 pages, 22 figures
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Data Structures and Algorithms (cs.DS)
Cite as:	arXiv:2410.01686 [cs.LG]
	(or arXiv:2410.01686v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2410.01686

Submission history

From: Artur Back de Luca [view email]
[v1] Wed, 2 Oct 2024 15:55:08 UTC (2,576 KB)

Computer Science > Machine Learning

Title:Positional Attention: Out-of-Distribution Generalization and Expressivity for Neural Algorithmic Reasoning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Positional Attention: Out-of-Distribution Generalization and Expressivity for Neural Algorithmic Reasoning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators