PublicationsVariBASed: Variational Bayes-Adaptive Sequential Monte-Carlo Planning for Deep Reinforcement LearningJoery A. De Vries, Jinke He, Yaniv Oren, Pascal R. Van der Vaart, Mathijs M. De Weerdt, and Matthijs T. J. Spaan. VariBASed: Variational Bayes-Adaptive Sequential Monte-Carlo Planning for Deep Reinforcement Learning. arXiv:2602.18857, 2026. DownloadAbstractOptimally trading-off exploration and exploitation is the holy grail of reinforcement learning as it promises maximal data-efficiency for solving any task. Bayes-optimal agents achieve this, but obtaining the belief-state and performing planning are both typically intractable. Although deep learning methods can greatly help in scaling this computation, existing methods are still costly to train. To accelerate this, this paper proposes a variational framework for learning and planning in Bayes-adaptive Markov decision processes that coalesces variational belief learning, sequential Monte-Carlo planning, and meta-reinforcement learning. In a single-GPU setup, our new method VariBASeD exhibits favorable scaling to larger planning budgets, improving sample- and runtime-efficiency over prior methods. BibTeX Entry@Misc{DeVries26arxiv,
author = {De Vries, Joery A. and He, Jinke and Oren, Yaniv and
Van der Vaart, Pascal R. and De Weerdt, Mathijs
M. and Spaan, Matthijs T. J.},
title = {{VariBASed}: Variational {B}ayes-Adaptive Sequential
{M}onte-{C}arlo Planning for Deep Reinforcement
Learning},
howpublished = {arXiv:2602.18857},
year = 2026
}
Note: This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright. In most cases, these works may not be reposted without the explicit permission of the copyright holder. Generated by bib2html.pl (written by Patrick Riley) on Fri Aug 28, 2026 12:56:06 UTC |