Publications

Sparse Masked Attention Policies for Reliable Generalization

Caroline Horsch, Laurens Engwegen, Max Weltevrede, Matthijs T. J. Spaan, and Wendelin Böhmer. Sparse Masked Attention Policies for Reliable Generalization. arXiv:2602.19956, 2026.

Download

pdf 

Abstract

In reinforcement learning, abstraction methods that remove unnecessary information from the observation are commonly used to learn policies which generalize better to unseen tasks. However, these methods often overlook a crucial weakness: the function which extracts the reduced-information representation has unknown generalization ability in unseen observations. In this paper, we address this problem by presenting an information removal method which more reliably generalizes to new states. We accomplish this by using a learned masking function which operates on, and is integrated with, the attention weights within an attention-based policy network. We demonstrate that our method significantly improves policy generalization to unseen tasks in the Procgen benchmark compared to standard PPO and masking approaches.

BibTeX Entry

@Misc{Horsch26arxiv,
  author =       {Caroline Horsch and Laurens Engwegen and Max
                  Weltevrede and Matthijs T. J. Spaan and Wendelin
                  B{\"o}hmer},
  title =        {Sparse Masked Attention Policies for Reliable
                  Generalization},
  howpublished = {arXiv:2602.19956},
  year =         2026
}

Note: This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright. In most cases, these works may not be reposted without the explicit permission of the copyright holder.

Generated by bib2html.pl (written by Patrick Riley) on Fri Aug 28, 2026 12:56:06 UTC