3.1. Rationale and Classification Methodology
An algorithm-family taxonomy is straightforward to construct, but it records which model was used rather than what the model changed about the routing decision.
Table 2 shows why: Q-learning selects the maximum-value neighbor within a classical scaffold in one study [
64] and, over city-grid states, selects the next grid toward the destination in another [
28]; the same recurrent architecture forecasts traffic in one design [
7] and classifies malicious nodes without forecasting in a second [
20].
Therefore, we investigate the corpus as a set of recurring answers to the five questions: the role and decision authority of the learner, the network-state representation, the information horizon, the temporal horizon, and policy organization and coordination. Nearly every corpus design that acts on the future does so through an explicit forecast, so we refer to the temporal-horizon dimension simply as prediction.
Figure 3 maps the declared extensions, citation links, and inferred continuations among the principal works.
Three fundamental properties of this analytical framework warrant explicit mention at the beginning. Firstly, the five dimensions are cross-cutting analytical descriptors rather than mutually exclusive classes: the same model can serve different dimensions in different papers, different models can answer the same dimension, and each design can in principle be characterized along all five. Secondly, for the countable summaries of
Section 3.8 and
Figure 4 only, we additionally assign each paper one primary-contribution label (its primary dimension hereafter), the dimension on which the paper’s stated novelty most directly changes the routing mechanism. The label was determined sequentially from (1) the contribution explicitly claimed by the original authors, (2) the routing component directly changed by the proposed mechanism, judged by the state–action interface the design actually changes, and (3) the principal variable that its own evaluation manipulates; citation by later studies and historical importance were not used. Cases that remain ambiguous after these tests are ruled on individually in
Section 3.7. Within a dimension a paper receives one position, not several labels; across dimensions it receives exactly one primary label, and any further relevance is recorded as secondary, retained in the qualitative analysis of
Section 3.2,
Section 3.3,
Section 3.4,
Section 3.5,
Section 3.6 and
Section 3.7 but excluded from the frequency counts. Hybrid designs that combine, for instance, route selection, link-quality prediction, and adaptive forwarding are therefore placed by the component their authors present as the contribution, with the other components noted as secondary.
Table 3 gives the operational test for each dimension and its boundary with its neighbors, and a worked example follows the classification rules below. Thirdly, we classify by a paper’s defining contribution, not by its most conspicuous algorithm: a paper in which the learner makes the next-hop choice is not automatically a learner-role paper because next-hop selection occurs across multiple dimensions; it belongs to this dimension only when a defining contribution is the relocation of decision authority from the classical protocol to the learner. Applying this rule to all 92 papers yields the distribution of
Figure 4: the role and decision authority of the learner accounts for 45 papers, the temporal horizon (prediction) 18, policy organization and coordination 11, network-state representation 11, and the information horizon seven.
Section 3.2,
Section 3.3,
Section 3.4,
Section 3.5 and
Section 3.6 trace the principal works of each dimension and the documented development relations among them. For the detailed mechanism comparison and evaluation audit,
Section 4 and
Section 5 use a focused analytical core of 37 papers; how the corpus was assembled and how the core was drawn from it are set out at the end of this section.
Figure 3 places the principal works by publication year, with three arrow types recording the relationships between them. Isolation is informative in its own right: PQR [
60], for instance, has no edge to any work shown, arriving at similar mechanisms with no direct citation relationship. Such cases suggest that similar solutions may emerge independently and substantial scope for consolidation remains.
A worked example. DeepCQ+ [
18] shows why the rule is needed. Its learner decides only broadcast versus unicast inside the CQ+ scaffold (role) from a fixed-size encoding of best-neighbor statistics (state) using the one-hop tables that CQ+ already keeps (information horizon) on the measured present (temporal horizon) with one policy shared by every node, trained centrally and transferred unchanged across network sizes (coordination). Organized by action authority alone it would fall under the learner’s role; its primary dimension is coordination because the shared size-agnostic policy is the contribution the paper presents as its own.
Classification procedure and its limits. Because the studies in which learning touches the routing decision are numerous and heterogeneous, a classification was needed to organize and count them, and the five dimensions were adopted so that the framework concerns the routing decision rather than the algorithm. All 92 papers were classified by the first author alone through multiple full-text review passes with boundary cases re-read against the rules before the assignment was fixed. No second rater took part and no inter-rater statistic was computed, and secondary relevance was recorded only for the boundary cases of
Section 3.7; both are limitations of the taxonomy, and no claim of formal inter-rater reliability is made. Two measures improve transparency and auditability: every boundary ruling is reported in
Section 3.7, and
Section 3.8 provides the complete primary assignment so that any alternative reading can be applied to the same corpus.
Corpus construction. Candidate studies were retrieved with an AI-assisted semantic literature search engine (Undermind) through eight iterative search paths, supplemented by forward and backward citation chasing (
Appendix A); no publication-date restriction was imposed, and the retained studies span 2004 to 2026. A study was retained if a learned or predictive component directly informs a routing decision in a MANET, VANET, or FANET or monitors or governs the operation of a learned router; learning that serves only an adjacent function, such as medium-access control or resource allocation, was excluded. In a small number of studies the predictive or scoring component is a statistical or fuzzy model rather than a trained learner (for example, the fuzzy road-suitability scoring of [
62]); these were retained because they occupy the same decision point as their learned counterparts, and the term learning-based is used in this survey to cover both. Preprint, conference, and journal versions of one study were merged and the published version is cited. The result is the 92-paper corpus classified in
Section 3.8;
Table 4 summarizes the steps.
The analytical core. Section 4 and
Section 5 concentrate on 37 papers selected by the first author for mechanism-level comparison on three considerations: extractability, the decisive one, meaning that the paper reports enough to identify the deployment scenario, the learned or predictive component, its interface with the routing decision, and at least one quantitative evaluation, the remaining audit fields of
Section 5 being coded as not reported where necessary; influence and venue, with well-cited studies in well-regarded journals and conferences preferred as a preference rather than a threshold since the designs of 2024 and 2025 could not yet have accumulated citations; and coverage, the five lineages of
Section 3.2,
Section 3.3,
Section 3.4,
Section 3.5 and
Section 3.6 having been reconstructed first from the full corpus and the core and then chosen to retain their principal works together with at least one design for every dimension and each network family. The core coincides with the designs examined in
Section 4;
Table A2 lists every corpus paper by its standing. The remaining 55 papers are classified in
Section 3.8 and discussed in
Section 3, where they extend a lineage, but their evaluations were not coded.
3.2. Role and Decision Authority of the Learner
The first dimension examines the learner’s role and decision-making authority within the routing stack. Rather than focusing on model depth, it addresses where the line of responsibility is drawn between the classical protocol and the learning agent. At one extreme is the learning-assisted protocol in which the learner tunes a single parameter, scores links, ranks routes, or makes the hop-by-hop choice, while route discovery, route repair, and control signaling remain the responsibility of a classical protocol. At the opposite extreme is the learned forwarding policy, which may extend to a clean-slate design in which routing is posed directly as a learning problem and the policy itself is the forwarding logic, with the classical core reduced to a baseline or a feature source or removed entirely. With 45 of the 92 papers, this dimension is the largest class in the corpus. The most common defining contribution is still the extent of the routing decision assigned to the learner.
Both extremes were visible by the mid-2000s. At the assisted extreme, for example, Usaha and Barria [
15] used a Monte Carlo reinforcement learning method to tune a single parameter (the initial probing-ticket allocation, chosen within a pre-specified maximum) of the bounded QoS search scheme of
Section 2.2, leaving the discovery framework otherwise intact. At the learned-forwarding extreme, SAMPLE [
13], a collaborative-RL routing scheme designed without a classical routing scaffold, chooses every packet’s next hop probabilistically from distributed Q-values, leaving no classical route-computation core for learning to “assist.” These two independent studies illustrate the two endpoints of this dimension.
The corpus’s most clearly documented lineage grew at the protocol-scaffold extreme in VANETs. Wu et al. embedded Q-learning into the AODV scaffold (QLAODV) to score path quality from vehicle mobility and available bandwidth and to switch to better routes pre-emptively [
12]. Their subsequent extension replaced the hand-crafted combination of link metrics with fuzzy-logic fusion, moved Q-value selection into route discovery itself, and removed the hard dependence on position and physical-layer information, assessing links from hello message reception when positioning is unavailable [
50]. The same group’s third step broadened the scope of learning beyond route selection to medium-access control, representing a shift toward broader system-level integration [
65], a continuation that is visible in the shared authorship and the field-tested system rather than an explicit citation of its predecessor. Deep reinforcement learning subsequently appeared at both ends of the spectrum rather than in sequence. DRQR [
16] retains AODV’s control machinery yet assigns the hop-by-hop forwarding decision to a deep Q-network in cognitive-radio MANETs (
Section 4.1). By contrast, [
44] and its continuation [
45] extended the learned-forwarding approach with relational features and deep RL over temporally extended actions, reported to generalize across traffic patterns, connectivity regimes, and, in [
45], node mobility. These two works occupy the same end of the spectrum as [
13], now using deep networks. Their defining generalization result is taken up under the fifth dimension (
Section 3.6).
Around this core lineage, a substantial number of studies fall within the assisted category. Many vehicular schemes occupy it with different scaffolds: ARPRL treats each packet as an agent choosing the maximum-Q neighbor [
64], and HAEQR augments QLAODV’s selection with an explicit link-maintenance-time model [
69]. RL-LAR integrates a per-vehicle Q-learner into the request zone of location-aided routing (LAR) [
70], and a routing-table-free variant maintains no end-to-end route, relying only on an adaptively updated Q-table [
71]. Others adapt the contention window of the routing protocol for low-power and lossy networks (RPLs) with Q-learning in vehicular deployments [
72], drive named-data VANET forwarding with deep prioritized SARSA [
40], and add a learned UAV-relay load-balancing layer above classical multipath routing [
73]. Collectively, these studies broaden the variables and operational objectives considered by the learner, including mobility, link persistence, contention, and load, without changing the learner’s broad functional role within the surrounding routing architecture. Deep networks enabled more complex decision functions without changing the learner’s functional role within the routing stack: a deep Q-network (DQN) determines the grid-to-grid routing trajectory, while low-level packet forwarding remains greedy [
74]; a convolutional network classifies the optimal next hop from a stacked neighbor-feature matrix on the GPSR scaffold [
75]; and a security-oriented stack uses an ensemble-LSTM node classifier to blacklist malicious vehicles before a bio-inspired optimizer routes [
20].
The non-vehicular literature covers similar levels of learner decision authority. In general MANET settings the learner variously replaces the hop-count metric with a cross-layer time metric [
76] or groups neighbors to estimate delay and delivery [
77]. In other schemes it adds mobility-, energy-, and link-quality-aware multipath choice [
78], decouples exploration from data traffic in a hybrid on-demand/proactive Q-router [
79], or learns a probabilistic multipath policy via SARSA [
39]. In a further scheme it solves routing as a partially observable Markov decision process (POMDP) under incomplete information about neighbors’ private parameters [
80]. Across these studies the learner is given a different quantity to optimize, while the surrounding protocol continues to supply the candidate set. Among the deep variants, several give the learner direct control over next-hop selection while retaining classical signaling, neighbor discovery, or other routing support. Examples include prioritized-replay double-DQN forwarding [
81], PPO-based single-hop selection replacing GPSR’s greedy/perimeter rule [
57], and deep-Q geographic routing that accounts for channel fading and stale beacons [
54]. A GRU-DDPG controller emits continuous per-link weights under an SDN [
82], and a DQN geographic policy serves high-speed robotic networks [
83]. One robotic-network scheme learns QGeo’s reward function by inverse RL instead of hand-crafting it [
47]. DeepMPR replaces hand-designed multi-point-relay selection, familiar from OLSR, with a learned multicast forwarding policy [
33]. Two opportunistic/delay-tolerant schemes learn the next forwarder through fuzzy-plus-Q-learning [
27] or a supervised classifier over routing features [
26]. In FANETs the dimension is just as densely populated: Q-FANET weights its Q-value updates by channel quality and episode recency [
51]; a traffic-balancing Q-network penalizes queue backlog on GPSR [
84]; QNGPSR reduces the use of perimeter mode, which can lengthen paths, by scoring neighbors with a Q-network [
85]; AR-GAIL learns a forwarding policy using AODV as the expert [
46]; a full-echo Q-router updates every neighbor’s value and anneals its exploration [
86]; and a further cluster adds adaptive Q-learning/AODV switching [
87], rate control [
88], distance-priority multi-Q updates that accelerate geographic route learning [
89], dueling-DQN link-quality routing [
90], resilient prioritized double-dueling forwarding under node failures [
91], attention-based DQN over IP-encoded positions [
92], centralized weight-tuning for cluster-head merit [
93], and a public-safety scheme in which UAVs relay for static victim clusters, a mixed aerial–ground setting that
Section 3.8 accordingly files under MANET/general [
6].
Only a few works occupy the learned-forwarding extreme: SAMPLE [
13] and the relational forwarding-policy line [
44,
45]. Most studies remain within the assisted category. One boundary case, a scheme that learns to reposition a relay node while classical AODV continues to route [
94], is retained in the assisted category: the learner does not select next hops, but its observations are router-observable signals, its reward is end-to-end delivery, and its actions determine whether the multi-hop path that AODV can use exists, so it shapes the routing decision through the connectivity it creates rather than the forwarding choice itself.
3.3. Network-State Representation Used by the Learner
The second dimension, the network-state representation used by the learner, focuses on what network entity, and which of its attributes, the learner represents as the state for its routing decisions: an individual vehicle or UAV, a local neighbor set, a grid cell, an intersection, a road segment, or a group of roadside units, described by attributes such as position, velocity, queue length, link lifetime, or residual energy. Here, “network state” refers to the information represented at the learner’s decision point, not to the complete global state space of the entire network. The choice has proved especially consequential in urban VANETs, where using individual vehicles as states scales poorly with network size: the state space grows with network size, slowing convergence [
28], and vehicle movement can render vehicle-level Q-tables obsolete before they converge [
53]. The papers on this dimension redefine the state representation to address this scalability problem, and what matters is not model depth (every scheme here is tabular Q-learning except for the single DQN variant [
55]) but that the chosen unit governs the size of the state space, how faithfully road topology is expressed, how infrastructure enters the decision, and whether the table can be adapted online. There are 11 papers in this category, all of which target vehicular applications with the exception of the most recent one.
QGrid [
28,
30] made the foundational change by changing the unit of abstraction: the city is divided into uniform geographic grids, and the grids—not the vehicles—serve as the states of Q-learning. A table of learned Q-values, trained offline on historical taxi GPS traces by exploiting the day-to-day stability of traffic within each grid, selects the next grid toward the destination, while a non-learned heuristic picks the specific relay vehicle inside it. Two design choices were central to QGrid: the grid geometry and the offline static table. Subsequent studies modified both. IV2XQ [
52] identifies four limitations of QGrid: grids ignore intersections and building occlusion; the offline table cannot observe network load; only a single fixed destination is supported; and always choosing the maximum-Q grid overloads popular regions. Benchmarking directly against QGrid, it re-grounds the abstraction in the road network itself: intersections become states, road segments become actions, and RSUs at intersections serve as relays, although the Q-table remains offline and congestion is handled by a fixed rule. HQGR [
53] then applies the same treatment to IV2XQ: taking it as the main baseline, criticizing its slow convergence and a Q-table overhead that scales with the product of RSU count and neighbor count, and redesigning the hierarchy around RSU groups whose Q-tables are learned online. The three teams—QGrid’s [
28,
30], IV2XQ’s [
52], and HQGR’s [
53]—share no authors. This indicates that the progression was not confined to a single research group.
Two additional lines of research developed independently. QTAR, proposed in [
29], should not be confused with the FANET protocol of the same name in
Section 3.4, so we refer to it as QTAR-VANET where context does not disambiguate. QTAR-VANET also argues against QGrid’s offline table—although the two are never compared empirically—running Q-learning fully online at two levels, between vehicles within a road segment and between RSUs across intersections. IQRRL [
66] arrives at the same two-level division without a direct citation relationship to these studies. However, it allocates learning and explicit scoring in the opposite way to IV2XQ, suggesting that the division can arise independently. Its next intersection is chosen by explicit QoS scoring rather than learning, while RL is confined to selecting the relay vehicle within a segment. IQRRL instead cites ARPRL [
64], its only learned baseline and the acknowledged source of its learning-rate setting. A cluster of related hierarchical designs elaborates the same idea: hello packets carry group Q-vectors and link rewards from which each RSU builds its group and local Q-tables [
95]; a software-defined directional grid platform folds positions and historical trajectories into grid directionality and two-hop relay selection [
96]; and a global-plus-local intersection scheme runs a central Q-learning server [
97]. A DQN variant centers the state representation on the current intersection, augmenting it with packet details and a timestamp. It then selects the next relay intersection as its action, extending this abstraction into a deep learning framework [
55].
The single non-vehicular entry provides particularly clear evidence that the principle is not confined to roads: a software-defined FANET protocol partitions the mission space into regular hexagonal cells and learns at the cell level, so the Q-table scales with the number of cells rather than the number of UAVs [
67]. This is the first example in the corpus to apply the same state-abstraction principle to an aerial network, suggesting that the dimension is not intrinsically vehicular but appears to have emerged first in vehicular settings. Across the corpus, the represented unit ranges from individual vehicles to grids, intersections, road segments, RSU groups, and aerial cells, while the learner remained largely fixed.
Section 4.2 examines what this re-engineering gains in convergence and costs in fidelity and adaptivity.
3.4. The Information Horizon
The third dimension, the information horizon, focuses on how much of the network a learner can observe and the control effort required to access it rather than what the state signifies. The distinction is the observation boundary, that is, how much present information the decision acquires and, at its inner edge, how the set of neighbors it evaluates is delimited: a one-hop neighborhood, a two-hop neighborhood, sensed local topology dynamics, or a candidate set narrowed by geometric or external filtering. This dimension is most clearly delineated in FANETs, where information obtained through beacon exchanges can become outdated within seconds and expanding the observation range incurs additional control overhead [
4,
32]; five of its seven members form the most closely related group of studies in the corpus, a sixth [
62] cites into it, and only the deep-learning member [
42] stands apart.
The earliest work in this lineage is QGeo [
34], framed for “unmanned robotic networks” spanning aerial and ground vehicles rather than FANETs proper. It recast one-hop geographic forwarding as distributed Q-learning: each node selects the next hop among its one-hop neighbors, guided by a reward based on packet travel speed that also accounts for link and location error, with periodic hello messages carrying position, current Q-value, link condition, and location error. Its closest successor in this corpus is QMR [
68], which keeps the one-hop view but performs multi-objective optimization—jointly minimizing delay and energy—replacing QGeo’s fixed learning rate and two-level discount factor with a learning rate adapted to observed delay variation and a discount factor adapted to neighbor-set change; QGeo is the only protocol baseline in QMR’s reported simulations.
Within this lineage, the boundary was subsequently widened only once. QTAR [
32] argues that a one-hop view can lead to blind paths and routing holes and extends the exchanged information to two-hop neighborhoods augmented with estimated link durations and delay, velocity, link-quality, and residual-energy information. It takes QGeo as a baseline rather than an acknowledged predecessor, and its adaptive learning-rate and discount-factor rules share their mathematical form with QMR’s. Although this relationship is not explicitly discussed, the two studies use structurally similar update equations. TARRAQ [
4] calls QTAR “a promising solution” and takes it as a baseline. Yet TARRAQ rejects QTAR’s defining two-hop exchange as excessive overhead and quantifies the complexity gap. It reverts to one-hop information, compensating with adaptive sensing intervals and Kalman-filter prediction of each link’s residual lifetime. The two protocols therefore address the same problem, topology awareness under rapid change, but not by the same method: TARRAQ substitutes temporal prediction for spatial breadth. Later work continued the substitution rather than resuming the expansion: LPMD-GPSR benchmarks against QGeo, QTAR, and QRF [
35] yet abandons the two-hop widening, and PSIF cites QTAR yet does not position itself as an extension of this lineage; both predictive designs are treated under the prediction dimension (
Section 3.5).
Meanwhile, a parallel line narrows the candidate space within a one-hop view: QRF [
35] filters the candidate set geometrically into a spherical sector before Q-learning to counter the state-space growth it associates with QTAR’s two-hop exchange. The one deep-learning member independently revisits the use of two-hop information: a DQN whose state stacks explicit one-hop and two-hop feature matrices while requiring only partial two-hop information, an accommodation of the incompleteness resulting from unstable links [
42]. The sole vehicle-based member of this dimension extends the decision scope vertically instead of laterally. Specifically, a UAV-assisted VANET framework relies on an aerial tier to calculate a global route using fuzzy logic scoring and depth-first search, which then enables ground vehicles to filter out off-route or congested neighbor nodes when choosing their next hop [
62].
The overall trajectory of this dimension is thus not monotone growth in what the router knows: within the main lineage the observation boundary expanded from one hop to two and then contracted again under overhead pressure, while the information supplied to the decision grew richer along other dimensions from measured neighbor tables to predicted link lifetimes and physically sensed surroundings.
3.5. Temporal Horizon: Measured Present Versus Predicted Future
The fourth dimension, the temporal horizon, asks whether the decision rests on the network as measured now or as it is predicted to become. This categorization depends on the temporal horizon of the routing input rather than the specific prediction model used. Papers belong in this category if they feed an explicit future-state forecast, or a prospective representation of future network states, into the routing decision regardless of whether that forecast is generated by an ANN, an LSTM, a statistical mobility model, a gradient-boosting regressor, or a clustering-based approach. (One boundary case instead projects into the future by planning across recorded encounter opportunities without generating an explicit forecast, as detailed in
Section 3.7). Counted by defining contribution rather than algorithm, this dimension is the second-largest, with 18 papers.
The clearest early instantiations of this prediction-before-routing architecture are vehicular. CRS-MP [
1] uses an SDN controller to train an ANN on real traffic-detector data to forecast the vehicle arrival rate on each road segment; RSUs and the base station convert that forecast through a stochastic traffic model into each request’s success probability and expected delay “ahead of time,” choosing between infrastructure-assisted vehicle-to-infrastructure (V2I) and multi-hop vehicle-to-vehicle (V2V) service modes so as to minimize overall service delay. PQR [
60] independently implements a related prediction-before-routing architecture with no direct citation link between the two papers: an acceleration-based statistical trajectory model estimates each link’s remaining lifetime while a gradient-boosting model fed by online measurements predicts route quality, and the protocol switches routes before the current one breaks or degrades. Two further branches complete the vehicular core: an ANN-based forwarder [
59] replaces the multi-metric forwarding criterion of the authors’ own protocol 3MRP with per-candidate delivery-probability prediction, citing [
1] only as background, and a delay-tolerant branch [
19] has each neighboring vehicle predict its own trajectory with an LSTM, so a sender can transfer its data to the vehicle predicted to reach a base station earliest.
A similar prediction-before-routing architecture subsequently emerged in FANETs, although as a result of engineering reuse rather than explicit conceptual lineage. PAP [
7], the earliest FANET entry in this prediction-before-routing line, uses an LSTM to forecast each UAV’s packet arrival rate and folds the forecast into a routing decision factor together with queue backlog, distance, and hop count; it cites [
1] for its SDN control architecture and channel model but not for the underlying prediction concept. JPE [
58] is treated here as an inferred continuation of this line on the basis of shared authorship, notation, and forecasting framework. Five of JPE’s seven authors overlap with those of PAP [
7], whose notation and LSTM framework carry over. JPE criticizes PAP by name (“only traffic-related prediction is not rewarding to multi-metric sensitivity routing decisions”) and retains it as a baseline. At the same time, prediction widens to joint mobility-and-traffic forecasting. Beyond that two-generation core the FANET line branches rather than continues: LPMD-GPSR [
17] cites JPE only once to call its link-duration computation “relatively complex” and instead trains an LSTM to predict link stability directly from timestamped neighbor tables on the classical greedy scaffold, while PSIF [
31], which in fact predates [
17] and arrives at the same substitution outside the PAP–JPE authorial line, migrates prediction-driven routing to GPS-denied low-altitude airspace, fusing obstacle intensity from the onboard SLAM map with historical signal strength into an LSTM link-quality forecast consumed by a DQN forwarder.
This dimension’s wider population organizes by what is predicted. Around the vehicular core sit an RSU-based system that predicts each vehicle’s next turn and the next RSU on its path by clustering [
49], a compatibility-aware scheme that predicts the connectivity duration between vehicles and quantizes it into programmable reliability classes [
98], and an SDN scheme whose random forest unit forecasts each vehicle’s next intersections and arrival times so that its deep relay-selector can favor relay vehicles whose predicted encounters are feasible [
61]. In the aerial literature, PARRoT folds trajectory prediction into the reinforcement-learning update through a time-varying mobility-aware discount factor [
99]; a prediction-driven 3D protocol switches its forwarding mode on a forecast of each UAV’s future distance to the ground station [
100]; and FANET schemes evaluate the future evolution of link states before relay selection [
101], predict UAV motion to drive adaptive beaconing and DDQN forwarding [
102], and anticipate future signal quality in a Q-learning router [
103]. Two general works complete the dimension: a link-failure-resilient Q-router predicts each neighbor’s remaining reachability time and folds it into a time-varying discount factor [
104], and a delay-tolerant scheme builds a temporal-graph state whose survival action lets the agent defer a present forwarding opportunity in favor of a future encounter represented in the graph, planning over the future without explicitly forecasting it [
105].
A broad tendency is observable across the vehicular and aerial prediction-based studies: prediction targets shift from aggregate traffic flows toward per-link lifetimes and future link quality, the input space widens to physically sensed surroundings, and the forecast moves ever closer to the forwarding decision itself.
Section 4.4 examines the benefit of computing a forecast before the routing decision and the impact of prediction errors.
3.6. Policy Organization and Coordination
The fifth dimension, policy organization and coordination, applies when a network contains many learners at once and asks how the resulting policies are organized, shared, coordinated, generalized, and deployed alongside classical rules. It concerns the lifecycle and governance of learned routing policies across nodes and operating conditions, including how policies are trained, shared, coordinated, transferred, selected, monitored, and replaced. It is broader than MARL: it covers a single size-agnostic policy shared under CTDE, explicit multi-agent cooperation, generalizable forwarding policies, and hybrid policy pools with rule-based fallback [
2,
14,
18,
45]. A paper belongs here when the organization of policies, not merely the fact that many nodes run the same algorithm, is its signature. This dimension is important because a network of learners raises problems that improved single-node reward design alone does not address: how many policies to maintain, how to keep them cooperating under a changing topology, how to generalize them to unseen network sizes, and which fallback mechanism should be used when a learned policy should no longer be trusted. It accounts for 11 papers.
Its starting point in this corpus, DeepCQ+ [
18], already addressed several aspects of the question at once. A single policy trained with PPO is shared by every node (“one policy for all”) and generalizes across network sizes through a size-agnostic neighbor encoding (the mechanism is detailed in
Section 4.5.1); yet the learner has a deliberately limited decision role, deciding only whether to broadcast or unicast each packet while next-hop choice, acknowledgement bookkeeping, and loop suppression remain the responsibility of the rule-based CQ+ protocol it is embedded in. Scalable generalization and learned–rule hybridity were thus present at this dimension’s origin in the form of one learned decision deliberately embedded inside a classical protocol.
Multi-agent studies then developed coordination further in different environments. For UAV swarms, [
2] uses DeepCQ+ in three roles: as background, as the target of a named criticism (jointly with another learned scheme for ignoring channel quality and node load), and as an experimental baseline. It responds with a cooperative stochastic game solved by MADDPG with recurrent actor and critic networks. For tactical sensor networks under jamming, [
5] proceeds in parallel—with no citation contact with DeepCQ+ or [
2]—and addresses the corruption that jamming introduces into the learning signal itself; its place on this dimension rests on organizing distributed policies under such corrupted feedback. Alongside these, a geographic FANET routing scheme employs cooperative Q-learning with a Nash-Q update mechanism across single-hop agents [
43]. A cooperative packet-routing algorithm makes each router a double-DQN agent over local and neighbor-queue state [
56], and a vehicular scheme lets independent Q-learning agents share and update a single Q-network [
41]. The relational forwarding-policy studies discussed under the first dimension are also relevant to this dimension: [
44] and its same-team continuation [
45], whose frozen-policy generalization spans device counts and, in [
45], mobility models. The key contribution of these studies lies in policy organization and coordination rather than merely establishing learned-forwarding ownership. The relational and DeepCQ+ lines are linked at one point, where [
45] cites the DeepCQ+ work among approaches confined to “relatively limited network scenarios with a few fixed flows and up to 50 devices.”
Chimera [
14] integrates these approaches and reorganizes the hybrid. It engages [
44] directly as one of its three learned baselines and the target of a named criticism (distributed training that can be trapped in local optima; an unrealistic assumption that nodes know all others’ real-time positions). Chimera addresses that criticism through its architecture rather than through increased model complexity: a policy pool in which nine learned policies and a classical protocol coexist. A device-agnostic vector of network conditions selects among the pooled policies. When conditions are too far from anything seen in training, the scheme falls back to an AODV-style traditional policy. A further member governs rather than routes: DeepADMR is not a router but a real-time anomaly detector that monitors the temporal-difference errors of a deployed DeepCQ+ policy and flags when the learned policy is operating outside its trusted operating regime [
63]. This makes the underlying issue of trust along this dimension explicit. The most recent work on this dimension at the corpus cut-off, HCPMR [
3], is a largely independent FANET branch: a GNN topology encoder with hierarchical multi-agent PPO within clusters and hand-crafted utilities between cluster leaders. Its experimental baselines (IQMR and CQMR) lie outside the corpus, and its only direct citation contact with the 10 works discussed above is a related-work citation of the Nash-Q geographic router [
43]. Overall, the evolution across this dimension centers on the structural architecture of coexistence between learning and rules rather than whether they coexist as they have since the beginning.
3.7. Cross-Dimension Relationships and Boundary Rulings
Cross-dimension relevance. Because the five dimensions are cross-cutting descriptors rather than disjoint bins, many papers have a primary contribution on one dimension and secondary relevance on others; the primary assignment follows the defining-contribution rule of
Section 3.1 and does not imply that a paper’s contributions are confined to that dimension. DRQR is primary on the learner-role dimension since its defining contribution is relocating hop-by-hop selection into a DQN inside the AODV scaffold. It also widens the information the decision consumes to a cross-layer physical/MAC/network state [
16]. JPE is primary on the prediction dimension for its joint forecast but also strengthens multi-metric path scoring [
58]. PSIF is primary on the prediction dimension for its future-link-quality forecast yet simultaneously enriches the information that is available to the routing decision with physically sensed obstacle context, which is why it recurs in both the information and prediction discussions above [
31]. LPMD-GPSR predicts link stability (temporal) but does so on a classical greedy scaffold, remaining partly within the assisted band of the learner-role dimension [
17]. The two-level QTAR-VANET and IQRRL are primary on state representation yet clearly also concern where learning sits since each also lets the learner control part of the forwarding [
29,
66]; the hexagonal-cell FANET scheme is primary on state representation but achieves scalability by decoupling Q-table size from the number of UAVs through its cell-level state [
67].
Boundary rulings. The relational forwarding-policy line illustrates most sharply why the primary label must follow the defining contribution. Its learned forwarding policy places it near the far end of the learner-role dimension, but the contribution its authors present is a forwarding policy that generalizes across networks without retraining, which
Table 3 lists under policy organization as size-agnostic generalization; we therefore assign that dimension as primary and record the learner-role placement as secondary [
44,
45]. SAMPLE receives the reverse assignment: it is assigned primarily to role and decision authority because its defining contribution is to make the learned policy itself the forwarding mechanism; its collaborative policy-sharing machinery is recorded as secondary [
13].
The remaining boundary decisions mark where this classification departs from a by-algorithm reading. Four works are assigned primarily to the temporal-horizon dimension, although a by-model view would classify them as generic learned routing: the next-RSU predictor [
49], the connectivity-duration predictor [
98], the future-link-state evaluator [
101], and the proactive link-failure predictor [
104]. Each places an explicit future-state forecast in front of the routing decision. The temporal-graph opportunistic learner is retained on the same dimension as a boundary case: it deploys no explicit predictor, but its temporal-graph state and survival action assign value to future encounter opportunities learned from temporal patterns [
105]. One work is excluded from the prediction dimension for the opposite reason: the security stack whose LSTM is a present-tense malicious-node classifier, not a forecaster, is read under the learner-role dimension [
20]. A Q-learner that selects the next hop is not primarily a learner-role paper when its distinctive move is a two-hop observation window [
42]; and DeepMPR, whose signature is re-learning MPR-style relay selection for multicast forwarding, is read under the learner-role dimension with coordination secondary [
33]. Where an isolated feature-based reading suggests placement on another dimension (e.g., QGeo and QMR as information horizon [
34,
68]; QTAR-VANET and IQRRL as state representation [
29,
66]), we retain the primary assignment determined by the defining-contribution rule of
Section 3.1 and record the alternative relevance as secondary. The primary assignment is used only for countable classification; secondary relevance is retained for the cross-dimension trade-offs analyzed in
Section 4.