1. Introduction
Belief Propagation (BP) [
1,
2] plays a pivotal role as a message-passing algorithm in graphical models. Its applications range from computing the partition function of a Markov random field [
3] to estimating marginal distributions [
4] and decoding low-density parity-check (LDPC) codes [
5]. Constraint optimization problems (COPs) [
6,
7] provide a versatile mathematical framework for modeling real-world challenges in transportation, supply chain management, energy, finance, and scheduling [
7,
8,
9,
10]. In the realm of COPs, BP, also recognized as Min-sum message passing [
11], seeks cost-optimal solutions by propagating cost information throughout the factor graph. Beyond standard BP, variational, bounding, and elimination-based perspectives provide additional context for inference in loopy or high-treewidth settings, including Tree-Reweighted BP (TRW) and Max-Product Linear Programming (MPLP) for bounds [
12,
13], as well as bucket elimination and mini-bucket approximations for structure-exploiting tradeoffs [
14,
15].
However, vanilla BP lacks convergence guarantees on factor graphs with loops and can converge to suboptimal fixed points in COPs with cyclic structure. Recognizing this challenge, significant research has focused on stabilizing loopy BP [
16,
17,
18,
19,
20,
21]. Notably, Damped Belief Propagation (DBP) [
21] has gained attention for improving convergence behavior on loopy graphs. These stabilization techniques are complementary to variational and LP-based reparameterizations that control oscillation and provide certificates when available [
12,
13].
Despite the advantages of damping, fine-tuning damping factors per instance remains laborious. Recent work integrates deep learning to automate these choices. In particular, Deep Adaptive Belief Propagation (DABP) leverages self-supervised learning with gradient-based optimization to determine damping factors, and further introduces dynamic damping based on real-time optimization status to improve per-iteration decisions. This trend is part of a broader line on neuralized message passing that parameterizes or augments BP updates while preserving their semantics, including the Belief Propagation Neural Network (BPNN) [
22], Neural-Enhanced Belief Propagation (NEBP) [
23], and the Factor Graph Neural Network (FGNN) for higher-order factors [
24], surveyed in learning for combinatorial reasoning and solver guidance [
25].
However, DABP faces scalability and efficiency challenges, primarily due to high GPU memory consumption and its reliance on autoregressive damping-factor predictions, which limit applicability to larger instances. To address these limitations, we introduce a paradigm that decouples damping-factor learning from the BP process. By eliminating the need for a recurrent neural network (RNN) over sequential BP messages, our approach enables joint prediction of damping factors, akin to conventional deep models. This reduces the number of deep neural network (DNN) calls, yields substantial memory savings, and improves efficiency over DABP. Additionally, our method leverages mixed precision to further reduce GPU memory. We refine learning by introducing priors, constraining the damping space, and learning adaptive variations rather than directly optimizing raw values, which enhances stability and convergence. Overall, our approach substantially improves both the efficiency and scalability of DABP. Our design aligns with recent efficiency-first trends in machine learning for discrete optimization that aim to limit sequential neural computation, for example, diffusion-based solvers such as DIFUSCO and Fast tokens-to-tokens (Fast T2T) and industrial-scale pipelines like DISCO [
26,
27,
28], as well as restart policies and learning-guided control for robust anytime behavior [
29,
30]. From a theoretical perspective, message-passing graph neural networks (GNNs) can capture optimal approximation algorithms for broad classes of Max-CSPs, with OptGNN providing a practical instantiation [
31].
The efficacy of our algorithm is substantiated through extensive experiments across diverse benchmarks, with an emphasis on larger problem instances. Our approach not only handles significantly larger problem sizes but also achieves an average speedup of 2.87× over DABP at equivalent restart counts. In terms of solution quality, our model remains competitive for smaller instances (variable size less than 150) and surpasses baselines on larger instances, where DABP falters due to memory limitations.
Our contributions are threefold: (1) we propose FDBP, a scalable BP framework that rethinks how damping factors are learned; by leveraging online self-supervised learning and parallelizable message processing, FDBP accelerates COP solving; (2) FDBP eliminates RNN hidden-state storage and per-iteration neural passes, which reduces GPU memory usage and enables scalability to large COP instances where prior DABP fails due to memory constraints; and (3) we conduct extensive experiments across diverse benchmarks, demonstrating that FDBP achieves a favorable accuracy–efficiency tradeoff and consistent gains over existing methods.
4. Methodologies
In each step, the DBP algorithm uses a fixed, manually chosen damping factor to balance the contribution of new and old messages when updating the message from a variable node to a function node in a factor graph. However, the new message is composed of messages received from the neighboring function nodes of the current variable node, and the importance of these messages can vary. To address this limitation, a more flexible approach should allow for the assignment of variable-specific damping factors and neighbor-specific weights for each variable node during the BP process. Specifically, dynamic damping factors and neighbor-specific weights can be integrated into the BP process. In step
t, a variable node
computes the message to function node
using the following expression:
Here, and represent learnable damping factors and neighbor weights, respectively, for the message from to at iteration t.
Suppose a problem instance has N variables, and its corresponding factor graph has an average node degree of M. If the BP process runs for T steps until termination, the complexity of updating the factors in adaptive BP is . This represents a computationally demanding task, making manual handcrafting impractical.
In response, DABP [
33] was proposed, leveraging DNNs parameterized by
to automatically infer these factors in an autoregressive manner:
Here, G represents the factor graph, and and denote the BP messages from variable nodes to function nodes and from function nodes to variable nodes in the previous steps, respectively.
However, the entire DABP procedure operates autoregressively, requiring the invocation of DNNs at each step to infer damping factors and neighbor weights. This recurrent dependency on DNNs can significantly increase computational overhead.
4.1. Fast Deep Belief Propagation
We propose an efficient learning-based algorithm for solving constraint optimization problems, referred to as Fast Deep Belief Propagation (FDBP). An illustration of the proposed approach is provided in
Figure 1. The algorithm will execute
R iterations, effectively restarting
R times. The restart policy has demonstrated significant potential in addressing COPs [
34]. Researchers have observed that combinatorial search algorithms often exhibit highly unpredictable runtime and solution quality across problem instances. By incorporating a restart policy, the algorithm minimizes the risk of getting trapped in local minima, thereby improving its overall efficiency and effectiveness.
In each iteration r, the algorithm runs for T steps. At each step t, it updates messages by using the messages from the previous step and inferred parameters . These parameters are computed by a Graph Attention Neural Network (GAT), which takes as input the factor graph and the messages from the previous iteration, specifically .
This design introduces a key improvement: instead of inferring parameters at each step t during the current iteration (as in the DABP algorithm), our approach infers all parameters for the entire iteration r in a single step, using a single invocation of the GAT. By doing so, the algorithm avoids the need to repeatedly invoke DNNs times per iteration, as required in DABP. This significantly reduces computational overhead and leverages GPU parallelism, making our approach far more efficient while retaining the effectiveness of parameter inference.
The time complexity of our algorithm per iteration matches that of DBP in the worst-case scenario. However, in practice, our approach often achieves faster performance due to its enhanced convergence behavior (as demonstrated in
Table 1 and
Table 2 of
Section 5), which reduces the required number of steps to reach convergence. Compared to DABP, our algorithm’s efficiency advantage becomes particularly evident when employing the same restart strategy, as all parameters for an entire iteration are computed simultaneously, significantly streamlining the process.
It is worth noting that our algorithm adopts a different approach to parameter inference compared to DABP. In DABP, the parameters for step t are inferred autoregressively by utilizing messages from all preceding steps (0 to ) within the same iteration, where a RNN is used to encode these messages. By contrast, our algorithm infers parameters for step using only the corresponding messages from the previous iteration, . This design assumes that the messages from step t of the previous iteration already encapsulate the relevant information from earlier steps. While not directly intervening in the BP process, this approach simplifies the inference process and complements our algorithm’s overall focus on efficiency and scalability.
In summary, our FDBP introduces a significantly more efficient and scalable framework for parameter inference in learning-based belief propagation. By addressing the computational bottlenecks inherent in DABP, our approach enhances both the speed and scalability of solving COPs. Specifically, FDBP overcomes the limitations of DABP by reducing the need for repeated, computationally expensive deep learning network invocations at each iteration, thereby lowering memory overhead and avoiding sequential computation bottlenecks. This leads to improved overall performance, making it well-suited for real-world, large-scale applications where both memory efficiency and computational speed are critical.
4.2. The FDBP Algorithm
The GAT model in our FDBP algorithm is trained using an online self-supervised learning approach, removing the dependency on labor-intensive, human-labeled datasets. This design allows our method to be directly applied to solving COPs without requiring prior model pretraining. Specifically, at each iteration
t, given the BP messages
m, we optimize the following self-supervised objective
:
where
denotes the belief-induced probability of the event
, defined by the following:
Intuitively, an assignment
is assigned a higher probability
when it has a lower belief cost
. It preserves the minimization semantics of the
argmin in Equation (
4) and is consistent with the global objective in Equation (
1). By leveraging this self-supervised loss, our algorithm adaptively refines its predictions in a principled manner, facilitating efficient and effective training.
Our FDBP algorithm is presented in Algorithm 1. The algorithm runs for R iterations (lines 3–19), with each iteration comprising the following steps:
Initialization: initialize messages and parameters for the factor graph (line 4).
Message Updates: perform sequential message updates using Equations (
6) and (
3) with the inferred parameters
(lines 6–7).
Solution Evaluation: periodically evaluate the current solution and compare its cost to the best solution found so far (lines 8–10).
Buffer Storage: store selected message pairs in a buffer for loss computation and gradient updates (lines 11–12).
Adaptive Optimization: optimize parameters adaptively through backpropagation using the stored messages and the defined loss function (lines 15–19).
Note that we use mixed precision to further economize GPU memory and refine the learning process by constraining damping factors within the range of 0.8 to 1.0. In a departure from conventional methodologies, we learn varying factors instead of directly optimizing damping factors, aligning with empirical observations suggesting optimal values around 0.9, as noted by [
21].
| Algorithm 1: The fast deep belief propagation algorithm |
![Mathematics 13 03349 i001 Mathematics 13 03349 i001]() |
4.3. Time Complexity
We denote by b the maximum scope size of a factor, d the maximum variable domain size, and e the average degree of a variable node in the factor graph. Let be the cost of one call to the GAT-based controller (a single forward pass). We follow the implementation setting where factors and variables are processed in parallel on a single GPU; the expressions below therefore reflect wall-time under this parallelism (work complexity would multiply by or accordingly).
Factor-to-variable update Equation (
3). Evaluating a factor message over a scope of size
b requires iterating over assignments of the other
variables, giving
.
Loss/objective evaluation Equation (
8). Dominated by factor evaluations; with factors processed in parallel on GPU, the wall-time is
.
Variable-to-factor (weighted BP) update Equation (
6). Aggregation over the
e incoming messages yields
(variables processed in parallel).
Decoding Equation (
4). Selecting the best label for one variable is
(variables processed in parallel).
Per BP iteration, the dominant cost is, therefore,
Our controller (GAT) is invoked once every K iterations, so each restart performs controller calls, contributing . We evaluate the objective once per restart, adding .
Total time complexity. With
R restarts and at most
T BP iterations per restart, the total wall-time is
For methods that infer weights/damping every iteration (e.g., DABP), replace with , which explains the higher per-iteration compute/memory footprint relative to our interval/restart-level controller.
5. Empirical Evaluations
In this section, we present a comprehensive empirical study on the effectiveness of our proposed FDBP. We commence by providing insights into the experimental setup and implementation details. Subsequently, we showcase the superiority of FDBP over existing state-of-the-art methods.
Benchmarks and Baselines. We evaluate on four canonical benchmarks: random COPs, scale-free networks, small-world networks, and Weighted Graph Coloring Problems (WGCPs) [
7]. For random COPs and WGCPs, constraint edges are sampled i.i.d. with graph density
. Scale-free instances are generated using the Barabási–Albert (BA) model with parameters
. Small-world instances follow the Newman–Watts–Strogatz construction with
and
. See Cohen et al. [
21] for further details on problem instance generation.
We benchmark FDBP against state-of-the-art solvers for COPs: (1) Toulbar2 with a timeout of 1200 s (7200s for large-scale problems); (2) Mini-bucket Elimination (MBE) with an i-bound of 8 (i-bound of 7 for specific random COPs); (3) GAT-PCM-LNS with a destroy probability of 0.2; (4) DBP with a damping factor of 0.9 and a splitting ratio of 0.95; and (5) DABP with different restarts.
All experiments are conducted on a server equipped with an Intel(R) Xeon(R) Gold 6148 CPU (Intel Corporation, Santa Clara, CA, USA), NVIDIA GeForce RTX 3090 GPUs (NVIDIA Corporation, Santa Clara, CA, USA), and 125 GB memory. The reported results represent the best solution cost for each run, and for each experiment, the results are averaged over 100 random problem instances. For larger-scale problems, a 2-hour runtime limit is imposed on all methods.
Implementation. Our FDBP model first applies a learned linear projection to a batch of BP messages to obtain 8-dimensional embeddings. These embeddings are then processed by a stack of four Graph Attention Network (GAT) layers. Each GAT layer produces eight feature channels using four attention heads. In line with DBP and DABP, we operate on a
Splitting Constraint Factor Graph (SCFG) with a splitting ratio of 0.95. The implementation is built in
PyTorch Geometric and trained with the Adam optimizer, using a learning rate of
and weight decay
. For each problem instance, we allocate
independent restarts and cap message passing at
iterations per restart (the restart budget and iteration cap were selected via a small grid search on random COPs with
, sweeping
and
. The best trade-off between accuracy and efficiency was achieved at
and
, with only minor gains beyond
). Dynamic weights and damping factors are re-estimated from the current BP messages every
iterations for most settings; for random COPs with
, the update schedule is tightened to every
iterations. We provide a list of parameters used in
Appendix B.
Performance Comparison. Table 1 and
Table 2 compare the performance of various methods across four standard benchmarks.
Table 1 focuses on smaller problem instances, while
Table 2 addresses larger problem instances.
Table 1.
Comparison of methods using normalised cost and relative gap. Cost is per constraint . Gap is with lower values preferred. Best results are highlighted in both bold and blue. OOM means out of memory.
Table 1.
Comparison of methods using normalised cost and relative gap. Cost is per constraint . Gap is with lower values preferred. Best results are highlighted in both bold and blue. OOM means out of memory.
| Random COPs |
|---|
| | | | |
| Methods | Cost | Gap | Time | Cost | Gap | Time | Cost | Gap | Time |
| Toulbar2 | 29.17 | 7.53% | 20m | 32.32 | 7.75% | 20m | 34.31 | 7.19% | 20m |
| MBE | 32.42 | 19.51% | 21s | 34.94 | 16.50% | 41s | 36.86 | 15.14% | 59s |
| GAT-PCM-LNS | 28.03 | 3.31% | 4m35s | 30.81 | 2.72% | 10m9s | 32.78 | 2.41% | 19m23s |
| DBP | 27.62 | 1.80% | 38s | 30.41 | 1.39% | 1m30s | 32.45 | 1.37% | 2m45s |
| DABP () | 27.21 | 0.28% | 57s | 30.10 | 0.35% | 1m8s | 32.09 | 0.25% | 1m25s |
| DABP () | 27.17 | 0.14% | 1m53s | 30.05 | 0.19% | 2m15s | 32.04 | 0.09% | 2m56s |
| DABP () | 27.13 | 0.00% | 3m36s | 29.99 | 0.00% | 4m23s | 32.01 | 0.00% | 5m44s |
| FDBP () | 27.24 | 0.42% | 17s | 30.13 | 0.45% | 29s | 32.08 | 0.23% | 53s |
| FDBP () | 27.22 | 0.32% | 32s | 30.08 | 0.29% | 55s | 32.05 | 0.14% | 1m48s |
| FDBP () | 27.17 | 0.14% | 1m1s | 30.06 | 0.21% | 1m43s | 32.02 | 0.03% | 3m42s |
| WGCPs |
| Toulbar2 | 0.19 | 0.00% | 20m | 1.19 | 36.70% | 20m | 2.09 | 41.14% | 20m |
| MBE | 2.04 | 1001.03% | 0s | 2.86 | 227.35% | 0s | 3.50 | 136.26% | 0s |
| GAT-PCM-LNS | 0.52 | 182.02% | 44s | 1.17 | 34.10% | 2m5s | 1.78 | 19.84% | 3m28s |
| DBP | 0.42 | 126.52% | 0m17s | 1.06 | 21.08% | 1m12s | 1.68 | 13.72% | 2m55s |
| DABP () | 0.32 | 70.94% | 1m43s | 0.90 | 2.49% | 2m53s | 1.59 | 7.55% | 5m38s |
| DABP () | 0.30 | 63.29% | 3m25s | 0.88 | 0.95% | 5m35s | 1.49 | 0.80% | 11m10s |
| DABP () | 0.29 | 55.07% | 6m35s | 0.87 | 0.00% | 11m8s | 1.48 | 0.00% | 22m17s |
| FDBP () | 0.33 | 79.84% | 31s | 0.90 | 3.23% | 1m | 1.51 | 1.65% | 2m56s |
| FDBP () | 0.32 | 73.69% | 1m3s | 0.89 | 1.47% | 1m59s | 1.50 | 0.96% | 5m35s |
| FDBP () | 0.31 | 69.37% | 2m1s | 0.88 | 0.47% | 3m57s | 1.49 | 0.32% | 11m9s |
| Scale-free networks |
| Toulbar2 | 30.90 | 7.21% | 20m | 31.64 | 8.01% | 20m | 32.34 | 9.24% | 20m |
| MBE | 33.97 | 17.87% | 22s | 34.55 | 17.95% | 31s | 34.84 | 17.69% | 40s |
| GAT-PCM-LNS | 29.72 | 3.11% | 5m28s | 30.41 | 3.82% | 8m48s | 31.10 | 5.06% | 13m25s |
| DBP | 29.28 | 1.60% | 41s | 29.79 | 1.70% | 1m | 30.03 | 1.43% | 1m29s |
| DABP () | 28.90 | 0.29% | 1m1s | 29.39 | 0.33% | 1m3s | 29.69 | 0.30% | 1m5s |
| DABP () | 28.86 | 0.12% | 1m58s | 29.34 | 0.16% | 2m4s | 29.63 | 0.11% | 2m6s |
| DABP () | 28.82 | 0.00% | 3m50s | 29.29 | 0.00% | 4m17s | 29.60 | 0.00% | 4m12s |
| FDBP () | 28.97 | 0.52% | 14s | 29.45 | 0.53% | 19s | 29.78 | 0.61% | 23s |
| FDBP () | 28.92 | 0.36% | 28s | 29.42 | 0.46% | 36s | 29.74 | 0.47% | 48s |
| FDBP () | 28.90 | 0.29% | 54s | 29.39 | 0.34% | 1m14s | 29.70 | 0.35% | 1m34s |
| Small-world networks |
| Toulbar2 | 28.51 | 11.01% | 20m | 28.80 | 12.37% | 20m | 28.37 | 10.67% | 20m |
| MBE | 29.46 | 14.69% | 16s | 29.40 | 14.73% | 22s | 29.31 | 14.32% | 28s |
| GAT-PCM-LNS | 26.49 | 3.13% | 3m24s | 26.43 | 3.14% | 5m6s | 26.26 | 2.45% | 7m19s |
| DBP | 25.87 | 0.70% | 35s | 25.88 | 0.99% | 56s | 25.68 | 0.16% | 1m17s |
| DABP () | 25.77 | 0.33% | 1m52s | 25.71 | 0.34% | 1m39s | 25.73 | 0.35% | 2m |
| DABP () | 25.72 | 0.13% | 3m42s | 25.67 | 0.16% | 3m18s | 25.67 | 0.12% | 3m57s |
| DABP () | 25.69 | 0.00% | 7m18s | 25.63 | 0.00% | 6m33s | 25.64 | 0.00% | 7m57s |
| FDBP () | 25.81 | 0.49% | 20s | 25.75 | 0.47% | 31s | 25.75 | 0.46% | 41s |
| FDBP () | 25.76 | 0.27% | 40s | 25.68 | 0.22% | 1m4s | 25.71 | 0.28% | 1m19s |
| FDBP () | 25.72 | 0.11% | 1m18s | 25.63 | 0.02% | 2m5s | 25.67 | 0.14% | 2m34s |
Table 2.
Comparison of methods using normalised cost and relative gap. Cost is per constraint . Gap is with lower values preferred. Best results are highlighted in both bold and blue. OOM means out of memory.
Table 2.
Comparison of methods using normalised cost and relative gap. Cost is per constraint . Gap is with lower values preferred. Best results are highlighted in both bold and blue. OOM means out of memory.
| Random COPs |
|---|
| | | | | |
| Methods | Cost | Gap | Time | Cost | Gap | Time | Cost | Gap | Time | Cost | Gap | Time |
| Toulbar2 | 37.16 | 5.49% | 2h | 38.92 | 4.69% | 2h | 40.14 | 4.27% | 2h | 41.05 | 3.90% | 2h |
| MBE | 39.48 | 12.057% | 2m32s | 41.00 | 10.29% | 4m40s | 41.98 | 9.03% | 27s | 42.71 | 8.09% | 40s |
| GAT-PCM-LNS | 35.86 | 1.81% | 1h5m | 37.86 | 1.84% | 2h | 39.32 | 2.13% | 2h | 40.39 | 2.22% | 2h |
| DBP | 35.56 | 0.96% | 9m50s | 37.50 | 0.87% | 22m38s | 38.76 | 0.67% | 45m43s | 39.73 | 0.57% | 1h90m |
| DABP () | | | | |
| DABP () | OOM | OOM | OOM | OOM |
| DABP () | | | | |
| FDBP () | 35.30 | 0.21% | 2m7s | 37.23 | 0.14% | 3m53s | 38.54 | 0.12% | 4m25s | 39.55 | 0.10% | 5m10s |
| FDBP () | 35.26 | 0.10% | 4m19s | 37.20 | 0.07% | 7m42s | 38.52 | 0.06% | 8m44s | 39.53 | 0.06% | 10m24s |
| FDBP () | 35.23 | 0.00% | 8m47s | 37.18 | 0.00% | 15m11s | 38.50 | 0.00% | 17m17s | 39.51 | 0.00% | 20m54s |
| WGCPs |
| Toulbar2 | 3.43 | 25.83% | 20m | 4.25 | 17.89% | 2h | 4.88 | 14.73% | 2h | 5.36 | 11.61% | 2h |
| MBE | 4.55 | 66.84% | 1s | 5.20 | 44.15% | 1s | 5.70 | 34.07% | 2s | 6.06 | 26.22% | 3s |
| GAT-PCM-LNS | 2.97 | 9.04% | 10m39s | 3.78 | 4.70% | 22m1s | 4.37 | 2.84% | 41m49s | 4.81 | 0.31% | 1h8m |
| DBP | 3.00 | 10.08% | 11m42s | 4.01 | 11.05% | 27m54s | 4.71 | 10.79% | 53m48s | 5.24 | 9.14% | 1h42m |
| DABP() | | | | |
| DABP () | OOM | OOM | OOM | OOM |
| DABP () | | | | |
| FDBP () | 2.79 | 2.53% | 5m48s | 3.70 | 2.42% | 11m43s | 4.38 | 3.00% | 18m54s | 4.99 | 3.89% | 24m57s |
| FDBP () | 2.75 | 0.78% | 11m32s | 3.64 | 1.01% | 23m17s | 4.31 | 1.22% | 38m9s | 4.89 | 1.94% | 50m |
| FDBP () | 2.73 | 0.00% | 23m18s | 3.61 | 0.00% | 46m15s | 4.25 | 0.00% | 1h16m | 4.80 | 0.00% | 1h40m |
| Scale-free networks |
| Toulbar2 | 33.09 | 10.27% | 20m | 33.07 | 9.12% | 2h | 33.18 | 8.81% | 2h | 33.45 | 9.50% | 2h |
| MBE | 35.26 | 17.48% | 1m3s | 35.40 | 16.79% | 1m25s | 35.51 | 16.46% | 1m51s | 35.65 | 16.71% | 2m18s |
| GAT-PCM-LNS | 32.32 | 7.68% | 30m46s | 33.27 | 9.76% | 53m15s | 33.79 | 10.82% | 1h20m | 34.26 | 12.14% | 1h54m |
| DBP | 30.36 | 1.17% | 2m34s | 30.59 | 0.93% | 4m9s | 30.69 | 0.65% | 5m15s | 30.78 | 0.74% | 6m29s |
| DABP() | 30.08 | 0.23% | 1m14s | | | |
| DABP () | 30.04 | 0.11% | 2m23s | OOM | OOM | OOM |
| DABP () | 30.01 | 0.00% | 4m41s | | | |
| FDBP () | 30.14 | 0.43% | 36s | 30.38 | 0.23% | 36s | 30.53 | 0.13% | 47s | 30.59 | 0.12% | 47s |
| FDBP () | 30.11 | 0.32% | 1m8s | 30.33 | 0.08% | 1m10s | 30.52 | 0.08% | 1m33s | 30.57 | 0.06% | 1m34s |
| FDBP () | 30.08 | 0.22% | 2m10s | 30.31 | 0.00% | 2m21s | 30.49 | 0.00% | 3m4s | 30.55 | 0.00% | 3m10s |
| Small-world networks |
| Toulbar2 | 28.16 | 12.66% | 20m | 28.77 | 12.13% | 20m | 27.95 | 8.80% | 2h | 27.74 | 8.06% | 2h |
| MBE | 28.47 | 13.90% | 42s | 29.29 | 14.17% | 55s | 29.87 | 16.26% | 1m9s | 29.63 | 15.43% | 1m23s |
| GAT-PCM-LNS | 25.66 | 2.65% | 14m41s | 26.32 | 2.60% | 22m50s | 26.91 | 4.74% | 32m56s | 26.65 | 3.84% | 45m3s |
| DBP | 25.00 | 0.00% | 2m49s | 25.66 | 0.00% | 4m24s | 26.18 | 1.89% | 6m31 | 25.95 | 1.10% | 8m20s |
| DABP () | 25.73 | 2.92% | 3m5s | 25.74 | 0.31% | 3m50s | 25.76 | 0.28% | 4m52s | 25.73 | 0.24% | 5m40s |
| DABP () | 25.69 | 2.77% | 6m6s | 25.69 | 0.14% | 7m36s | 25.73 | 0.15% | 9m59s | 25.69 | 0.09% | 11m27s |
| DABP () | 25.65 | 2.60% | 12m28s | 25.66 | 0.01% | 14m54s | 25.69 | 0.00% | 19m57s | 25.67 | 0.00% | 22m56s |
| FDBP () | 25.75 | 3.01% | 1m4s | 25.77 | 0.42% | 1m44s | 25.79 | 0.39% | 3m5s | 25.77 | 0.41% | 2m55s |
| FDBP () | 25.70 | 2.83% | 2m7s | 25.72 | 0.22% | 3m35s | 25.75 | 0.22% | 6m1s | 25.74 | 0.29% | 5m49s |
| FDBP () | 25.67 | 2.69% | 4m18s | 25.68 | 0.08% | 7m6s | 25.72 | 0.12% | 11m49s | 25.71 | 0.15% | 11m22s |
Even with generous time budgets (20 min and 2 h), Toulbar2 lags behind on most benchmarks. This outcome is consistent with its largely systematic search strategy: as instance size grows, the effective branching factor renders global exploration impractical, even with pruning and heuristics. By contrast, MBE often achieves shorter runtimes; however, memory limits force small bucket sizes and weaker relaxations, so its solution quality is substantially lower than the stronger competitors.
The results further highlight the difficulty of obtaining certifiably optimal labels for DNN-based BP variants using an exact solver. GAT-PCM-LNS, which combines iterative large neighborhood search with a machine-learned repair policy, often delivers higher-quality solutions than Toulbar2 and MBE. This advantage comes at a cost: wall-clock time is markedly longer because each LNS iteration destroys a subset of variables and requires multiple rounds of model inference to repair and reassign them.
Due to constraint function splitting, DBP frequently outperforms Toulbar2, MBE, and GAT-PCM-LNS. Additionally, DBP exhibits shorter runtime compared to most other methods. By seamlessly integrating BP and deep learning models to infer dynamic weights and damping factors, DABP achieves better solutions than DBP within an acceptable increase in time. However, it faces limitations in solving larger-scale problems due to GPU memory constraints.
Our proposed FDBP demonstrates significant superiority across various benchmarks:
Computational Efficiency: FDBP achieves an average speedup of 2.87x over DABP at equivalent restart counts (R). For 100-variable random COPs, FDBP () completes in 3m42s vs. DABP’s 5m44s (35% faster) with only a 0.03% cost gap.
Scalability Advantage: DABP encounters out-of-memory failures beyond 150 variables, while FDBP successfully handles 300-variable instances (
Table 2). On 300-variable WGCPs, FDBP (
) achieves optimal normalized cost (4.80) in 1h40m versus DBP’s suboptimal 5.24 (9.17% gap) in 1h42m.
Performance-Efficiency Tradeoff: While DABP () maintains 0.00% gap on smaller instances, FDBP () shows marginally higher gaps (0.03–0.35%) but with 35–82% runtime reductions. This tradeoff becomes favorable for FDBP at scale—it solves 300-variable scale-free networks in 3m10s (0.00% gap) where DABP fails.
DABP integrates GRUs and multi-head GAT to infer dynamic weights and damping at each BP iteration, adding per-iteration compute and memory; FDBP avoids such per-iteration neural passes (one lightweight controller call per update interval/restart), trading a small, small-instance gap for substantial speed and lower peak memory. This design choice explains why FDBP is 35–82% faster and stays within 15.2 GB at
, while DABP hits OOM at
(see
Table A1). Note that we report peak GPU memory usage in
Appendix A.
Anytime Performance Analysis. We conduct further comparison of solution quality among different methods over time for problem instances with
in
Figure 2. We have the following key insights:
FDBP consistently outperforms all other methods across all problem types, demonstrating the fastest convergence to the optimal or near-optimal solution with the lowest normalized cost. This suggests that FDBP is the most efficient algorithm overall.
DABP and DBP generally show slower but steady improvement, with DABP typically outperforming DBP.
GAT-PCM-LNS and Toulbar show slower progress and are less effective in terms of both speed and quality of the solution compared to FDBP, DABP, and DBP.
MBE consistently reaches a plateau early but at a much higher cost than the other methods, indicating that while it converges faster, it is less effective at finding low-cost solutions.
Overall, the figures highlight the superiority of FDBP in terms of both speed and the quality of the solution across all problem instances. The other algorithms, while effective to some extent, generally converge slower and achieve higher costs, with MBE being the least competitive in terms of performance.
Albation Study. In loopy BP, messages at step
t summarize the influence of steps
. Conditioning the inference network for restart
r on
, therefore, provides near-sufficient context for predicting
and
without constructing an intra-iteration autoregressive encoder. This preserves standard BP updates while avoiding sequential neural calls, reducing latency and memory. To confirm this, we also implemented an intra-iteration variant (FDBP-AR). On
, FDBP-AR attains similar cost/gap but requires
runtime and 15–
higher peak GPU memory (see
Table 3).
Sensitivity Analysis. We add a sensitivity study (
Table 4). For the update interval
K on
with
, we observe a clear accuracy–efficiency trade-off—
yields cost
, gap
, and 160 s;
yields
,
, and 121 s;
yields
,
, and 108 s;
yields
,
, and 83 s—suggesting
as a good balance. For damping, the range
is stable and fast, whereas
occasionally oscillates on WGCPs. For restarts, returns diminish beyond
; for
, increasing from
to
increases time from 108 s to 222 s for only a
gap improvement. Finally, for GAT capacity, a smaller configuration
is
faster but slightly less accurate, a larger configuration
is
slower but slightly more accurate, and the default
provides the best accuracy–efficiency balance.
6. Conclusions
This work addresses critical limitations in modern BP methods for COPs. While DBP improved convergence through manual damping heuristics, and DABP automated these via deep learning, both suffered from fundamental bottlenecks: DABP’s reliance on sequential RNNs for damping prediction introduced prohibitive memory overheads and limited scalability, while DBP’s manual tuning proved impractical for large-scale applications. We propose FDBP, a scalable BP framework that fundamentally rethinks how damping factors are learned. By decoupling damping inference from BP iterations and replacing RNNs with periodic message aggregation (every K steps), FDBP achieves the following:
Memory Efficiency: Eliminates RNN hidden state storage, significantly reducing GPU memory usage compared to DABP, allowing scalability to 300-variable COPs, where DABP fails due to OOM issues.
Computational Efficiency: Parallelizable message processing reduces runtime by an average of 65% at equivalent restart counts , solving 100-variable COPs in 3m42s compared to DABP’s 5m44s, with only a 0.03% loss in solution quality.
Practical Optimality: Maintains near-DABP performance (0.00–0.35% gaps) on small instances while outperforming all baselines on large-scale problems (e.g., 300-variable WGCPs: 4.80 normalized cost vs DBP’s 5.24).
Empirical results across synthetic and topology-driven benchmarks confirm that FDBP’s design choices successfully balance the accuracy–efficiency tradeoff.
Limitations and future work. This work focuses on static, single-agent COPs; time-varying structures and decentralized coordination are out of scope. In future work, we will extend the framework to dynamic COPs and multi-agent systems by adding online message updates, warm starts, and communication-efficient, decentralized controllers, enabling more adaptive and distributed problem solving.
Broader implications. This work advances applied mathematics, computing, and machine learning by coupling model-based belief propagation with a lightweight learned controller, which preserves interpretability while improving efficiency. The framework is GPU-friendly and memory-efficient, aligning with high-performance computing for AI. It supports uncertainty-aware decisions via belief estimates and applies broadly to discrete-optimization tasks in science and engineering, including scheduling, routing, network design, and error-correcting codes. This strengthens applied machine learning and AI for scientific discovery under fixed hardware budgets.