Next Article in Journal
Entropy Method for Decision Making with Uncertainty
Next Article in Special Issue
RIB-Guard: A Risk-Aware Information Bottleneck Defense for Black-Box Large Language Models
Previous Article in Journal
Entropy Bathtub for Living Systems: A Markovian Perspective
Previous Article in Special Issue
MVIB-Lip: Multi-View Information Bottleneck for Visual Speech Recognition via Time Series Modeling
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Information-Theoretic Intrinsic Motivation for Reinforcement Learning in Combinatorial Routing

1
Krieger School of Arts and Sciences, Johns Hopkins University, Washington, DC 20001, USA
2
School of Integrated Circuit Engineering, Guangdong University of Technology, Guangzhou 510006, China
3
School of Computer Science, University of Liverpool, Liverpool L69 3DR, UK
*
Authors to whom correspondence should be addressed.
Entropy 2026, 28(2), 140; https://doi.org/10.3390/e28020140
Submission received: 28 December 2025 / Revised: 18 January 2026 / Accepted: 21 January 2026 / Published: 27 January 2026
(This article belongs to the Special Issue The Information Bottleneck Method: Theory and Applications)

Abstract

Intrinsic motivation provides a principled mechanism for driving exploration in reinforcement learning when external rewards are sparse or delayed. A central challenge, however, lies in defining meaningful novelty signals in high-dimensional and combinatorial state spaces, where observation-level density estimation and prediction-error heuristics often become unreliable. In this work, we propose an information-theoretic framework for intrinsically motivated reinforcement learning grounded in the Information Bottleneck principle. Our approach learns compact latent state representations by explicitly balancing the compression of observations and the preservation of predictive information about future state transitions. Within this bottlenecked latent space, intrinsic rewards are defined through information-theoretic quantities that characterize the novelty of state–action transitions in terms of mutual information, rather than raw observation dissimilarity. To enable scalable estimation in continuous and high-dimensional settings, we employ neural mutual information estimators that avoid explicit density modeling and contrastive objectives based on the construction of positive–negative pairs. We evaluate the proposed method on two representative combinatorial routing problems, the Travelling Salesman Problem and the Split Delivery Vehicle Routing Problem, formulated as Markov decision processes with sparse terminal rewards. These problems serve as controlled testbeds for studying exploration and representation learning under long-horizon decision making. Experimental results demonstrate that the proposed information bottleneck-driven intrinsic motivation improves exploration efficiency, training stability, and solution quality compared to standard reinforcement learning baselines.
Keywords: intrinsically-motivated reinforcement learning; information bottleneck; curiosity-driven exploration; combinatorial routing problems intrinsically-motivated reinforcement learning; information bottleneck; curiosity-driven exploration; combinatorial routing problems

Share and Cite

MDPI and ACS Style

Xi, R.; Ni, Y.; Wu, W. Information-Theoretic Intrinsic Motivation for Reinforcement Learning in Combinatorial Routing. Entropy 2026, 28, 140. https://doi.org/10.3390/e28020140

AMA Style

Xi R, Ni Y, Wu W. Information-Theoretic Intrinsic Motivation for Reinforcement Learning in Combinatorial Routing. Entropy. 2026; 28(2):140. https://doi.org/10.3390/e28020140

Chicago/Turabian Style

Xi, Ruozhang, Yao Ni, and Wangyu Wu. 2026. "Information-Theoretic Intrinsic Motivation for Reinforcement Learning in Combinatorial Routing" Entropy 28, no. 2: 140. https://doi.org/10.3390/e28020140

APA Style

Xi, R., Ni, Y., & Wu, W. (2026). Information-Theoretic Intrinsic Motivation for Reinforcement Learning in Combinatorial Routing. Entropy, 28(2), 140. https://doi.org/10.3390/e28020140

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop