1. Introduction
Urban air mobility (UAM) has the potential to reshape urban transportation by introducing electric vertical takeoff and landing (eVTOL) aircraft as a new mode of travel. Cities seeking to adopt this technology must prepare through forward-looking planning that anticipates infrastructure needs and ensures that aerial mobility integrates with existing urban systems [
1,
2]. Central to this task is vertiport site selection, which determines the accessibility, efficiency, and long-term viability of UAM networks. Traditional methods for site selection, such as multi-criteria decision analysis [
3], optimization models [
4], and GIS-based suitability mapping [
5], provide structured evaluation but are constrained in flexibility.
These methods rely on static criteria and fixed workflows, which makes it difficult to accommodate diverse planning strategies or respond quickly to new objectives. When faced with multi-objective and dynamic planning contexts, these approaches lack the generalization capacity needed for scalable and adaptive decision-making. Many existing studies focus on isolated vertiport site selection [
6], while integrating siting with ground transportation—such as traffic flows, metro stations, and bus stops—remains underexplored, despite its importance for first/last-mile access and passenger uptake [
7]. Emerging work shows how co-locating vertiports with high-accessibility nodes or along strategic corridors can materially improve reach and network performance, arguing for joint siting with transit and roadway systems rather than treating UAM as an isolated layer [
8].
Large language models (LLMs) offer a new pathway to address these limitations. Their capabilities in natural-language understanding and reasoning allow planning criteria expressed in policy text or expert guidelines to be translated directly into structured constraints. By integrating heterogeneous spatial and transportation datasets and orchestrating tool-based geospatial operations, LLMs enable workflows that are more flexible and adaptable across scenarios. Recent work positions LLMs as interfaces that connect regulations, domain knowledge, and quantitative evaluation pipelines—linking rules to computable indicators and models [
9]. A geospatial LLM, trained to generate multi-step tool chains, turns natural-language specifications into reliable GIS workflows, demonstrating accuracy on complex spatial tasks and highlighting the feasibility of LLM-driven urban analytics [
10].
To address these challenges, we propose a strategy-aware LLM-based framework for vertiport site selection that integrates ground transportation and rapidly adapts to diverse planning strategies. Our agent (i) integrates heterogeneous spatial data and aligns it with an input planning strategy; (ii) introduces LLM reasoning and planning into the siting pipeline; (iii) establishes a reflective loop connecting planner, executor, and validator for iterative refinement; and (iv) orchestrates domain-specific tools for data processing, optimization, and visualization. Empirically, we evaluated the framework with two site selection methods—MCDM and a genetic algorithm—each instantiated under three strategies (coverage-, equity-, and efficiency-oriented). The outcome evaluation and process evaluation results show strategy–site alignment, reproducible tool chains, and gains on the outcome metrics, supporting the feasibility and promise of LLM agents for vertiport siting. The framework delivers reproducible, auditable pipelines that formalize policy text into measurable indicators and constraints, select relevant siting metrics, and orchestrate geospatial tools. This supports rapid what-if analyses, serving as a practical assistant for planners. By supporting flexible and adaptive siting decisions, the framework contributes to intelligent urban planning practices and provides a decision-support tool for smart city development.
2. Review
2.1. Vertiport Site Selection Problem
Urban Air Mobility (UAM) has emerged as a prospective transport layer that integrates electric Vertical Take-off and Landing (eVTOL) aircraft into metropolitan systems. EVTOL vehicles offer point-to-point connectivity with low noise, reduced emissions, and minimal ground footprint, making them suitable for dense urban settings [
10]. Several pilot trials and demonstration flights have already been conducted in cities such as Los Angeles, Dubai and Shenzhen, indicating that deployment is shifting from concept to practice. As the backbone of UAM, vertiports serve as the primary ground infrastructure for eVTOL operations [
11]. Their siting determines accessibility, integration with existing transport systems, and operational safety [
6].
Vertiport siting has therefore become a central challenge. Existing studies often frame the siting task using multi-criteria decision analysis (MCDA), optimization models, or GIS-based suitability mapping. Feasibility analyses apply GIS-based multi-criteria evaluation to identify candidate sites; for example, in the Seoul metropolitan region, 148 potential sites were screened to 56 suitable locations after accounting for airspace, demand, and accessibility constraints [
3]. Capacity-oriented studies emphasize throughput requirements, showing that high-demand nodes may need 600–1200 takeoffs and landings per hour—and careful pad/turnaround sizing—to sustain reliable operations and avoid bottlenecks [
4,
5]. System-level evaluations link vertiport siting to urban accessibility, finding that optimized layouts can yield up to 50% reductions in travel time for passengers [
12] and more conservative 2–5% improvements in overall network efficiency, depending on multimodal integration [
13].
These findings illustrate that vertiport siting is simultaneously a land use, operations, and transport-integration problem. Existing approaches, however, often rely on static demand assumptions and fixed weighting of criteria, limiting their adaptability to heterogeneous data and evolving policy requirements. This gap motivates more flexible frameworks capable of incorporating diverse constraints and dynamically refining candidate sites in response to strategic planning needs.
2.2. Large Language Models in Urban Planning
Large language models are entering urban planning as assistants that generate planning text, retrieve evidence, and evaluate documents [
14,
15]. Recent perspectives identify LLMs as interfaces integrating regulatory text, domain knowledge, and quantitative evaluation workflows. PlanGPT instantiates this role by combining a planning-specific language model with efficient retrieval and a planner-style document workflow; it answers policy queries, drafts regulation-conforming prose, and runs reviews against codified checklists with an agent that calls external tools when needed [
16]. GeoGPT formalizes geospatial task prompts and routes requests to spatial tools and code synthesis, enabling coordinate handling, projection conversion, and dataset querying from plain language [
17]. LLMs act as mediators between natural-language inputs and structured planning workflows, linking policy text, spatial data, and automated verification into integrated urban planning support systems.
LLMs also have applications in site selection, using natural-language understanding and reasoning to translate textual siting criteria into geospatial operations [
18]. For evaluation, a civilian analog evaluates parking facility locations at the city scale using an LLM-based agent to score candidate sites under operational criteria, indicating transferability of agentic siting to public infrastructure [
19]. UrbanMUDA adapts an LLM agent with domain instructions, explicit geospatial tool calls, and a self-refinement loop to place urban military units. Tool-augmented reflection raises geospatial reasoning accuracy [
20]. LocationReasoner provides a deterministic benchmark for real-world siting with heterogeneous constraints and a fixed tool sandbox. Results show current models improve when grounded in external geospatial tools and automatic verification, while pure prompting underperforms on multi-constraint queries [
21]. However, site selection remains challenging because LLMs frequently mismanage multi-constraint logic, applying filters in sequence rather than evaluating conditions holistically. This leads to valid sites being discarded and weakens trust in the results. Addressing this gap requires workflows that explicitly extract constraints, coordinate tool-based operations, and refine outputs.
3. Problem Formulation
To demonstrate strategy-aware LLM orchestration across both interpretable scoring and combinatorial optimization, we instantiate two site selection methodologies—multi-criteria decision analysis (MCDA) and a genetic algorithm (GA). MCDA provides a transparent, policy-traceable baseline that maps planning guidelines to weighted indicators and rapidly ranks candidates. GA tackles the inherently nonlinear, constrained nature of set selection under spatial and operational rules, exploring system-level trade-offs that simple ranking cannot capture.
Let denote the universe of parcels/cells. A candidate set is obtained by land-use screening; we denote the final candidate index set by . Decision variables indicate whether site is selected. be the site budget, and encode feasibility constraints. A strategy is converted by the LLM into (i) a set of criteria, (ii) their weights, and (iii) constraint parameters; these feed either MCDA or GA below.
3.1. Multi-Criteria Decision Analysis
MCDA provides a transparent mechanism to translate planning guidance into quantitative scores for feasible candidates [
22]. Each indicator is classified as either a benefit criterion (larger is preferable) or a cost criterion (smaller is preferable) and normalized to the range [0, 1] using min–max scaling (
avoids division by zero). For candidate
and indicator
with raw value
, the normalized value is calculated in Equations (1) and (2).
Each criterion is assigned a non-negative weight
with
. A composite score for site
is then obtained as Equation (3).
Selection seeks to maximize the overall score subject to the site budget and feasibility constraints as Equation (4).
where
is the site budget, i.e., the number of sites planned for construction. The feasibility set
may include: (i) a land-use mask
with
, where
indicates parcel
is eligible after land-use screening; (ii) pairwise spacing constraints
whenever
here
is the inter-site distance and
e minimum separation threshold; and (iii) additional exclusion buffers.
In practice, we employ a rank-then-filter heuristic: sort candidates by and greedily admit sites that satisfy until are selected. This approach yields an auditable and policy-traceable mapping from indicators and weights to the final siting outcome.
3.2. Genetic Algorithm
The genetic algorithm (GA) addresses the nonlinear and combinatorial nature of vertiport site selection by searching directly over feasible subsets of sites [
23]. Each individual encodes a candidate solution as a binary vector, and the fitness function evaluates its performance against strategy-defined objectives. The search proceeds iteratively through selection, crossover, mutation, and repair, gradually improving population quality. The fitness function is expressed as Equation (5).
where
is the strategy-instantiated utility of solution
,
measures the violation of constraint
, and
is the penalty weight. Elitism ensures that the best-performing individuals are retained across generations, while swap-based mutation and repair operators maintain feasibility and the site budget.
In the GA process, the purpose is to capture nonlinear trade-offs and explore globally optimal site sets. The planner configures the objective, feasibility constraints, and algorithmic parameters to translate strategy into an optimization problem. Strategy information adjusts the relative importance of coverage, equity, and economic terms in the GA fitness function by modifying their scalar weights, and the underlying functional form remains stable across runs unless the strategy requires additional terms or altered penalty strength. The executor adapts and runs code modules for initialization, crossover, mutation, and fitness evaluation to search the combinatorial space of candidate subsets, and parameters may be refined when the planner identifies inconsistencies between strategy requirements and intermediate outputs. The validator applies the same performance criteria to ensure comparability across methods and robustness of the results. This workflow highlights the LLM’s ability to adapt algorithmic components while preserving a consistent optimization structure, enabling strategy-aware exploration of siting configurations beyond the scope of simple ranking.
4. Framework
To support flexible and strategy-aware vertiport site selection, we propose an LLM-based agent workflow that integrates heterogeneous spatial data with planning objectives. The workflow consists of two modules, as shown in
Figure 1: candidate generation through land-use screening and site optimization using either multi-criteria decision analysis or a genetic algorithm. A knowledge base provides references, parameters, and planning strategies to ground the process in domain knowledge and regulatory context. Heterogeneous spatial inputs—including income, land use, traffic flows, demand, and transit stops—are incorporated to capture socio-spatial conditions that affect feasibility and accessibility. OD origins and destinations are treated as point features and are assigned to parcels through point-to-polygon spatial joins, which transfer their associated demand and travel time attributes to the parcel layer. Ground transportation variables, including first- and last-mile access and transit availability, are also derived from point data and are linked to parcels using the same join procedure. This produces a unified candidate table in which each parcel carries a consistent set of attributes derived from the original .shp, .csv, and .txt sources. The harmonized representation ensures that feasibility and accessibility indicators originating from different formats can be evaluated within a single spatial framework.
At the core of the workflow is an LLM agent that orchestrates task decomposition, execution, and validation in an iterative manner. A reflective loop ensures that outputs can be refined and aligned with strategy requirements, while a standardized tool set supports data integration, constraint handling, and visualization. This design allows the framework to adapt across planning strategies and datasets, providing a generalizable pipeline for site selection under different urban mobility contexts.
4.1. Site Selection Process
After heterogeneous datasets are integrated, candidate parcels are first screened by land use, which provides the regulatory feasibility filter that defines which parcels are legally and physically eligible for vertiport development, before the framework advances to site selection through MCDA or GA. Both are coordinated by the agent but emphasize different roles of LLM orchestration: MCDA relies more on structured function calling to normalize indicators and compute scores, while GA requires code generation and adaptation to implement evolutionary operators and explore feasible site combinations.
In the MCDA process, the purpose is to provide a transparent, auditable mapping from planning strategy to selected sites. As shown in
Figure 2, the planner interprets the strategy to identify relevant indicators, assign weights, and determine the number of vertiports. The executor normalizes indicator values and aggregates them into composite scores through function calling to operationalize planning criteria into quantitative evaluation. A distance filter is then applied to enforce spatial feasibility, and visualization tools present the selected sites to support interpretation and communication of results. The validator assesses the outcome against demand coverage, equity, economic return, and accessibility to ensure that siting decisions align with policy objectives. This workflow demonstrates the LLM’s ability to transform textual strategies into structured indicators and to coordinate tool-based selection under explicit constraints.
In the GA process, the purpose is to capture nonlinear trade-offs and explore globally optimal site sets. The planner configures the objective, feasibility constraints, and algorithmic parameters to translate strategy into an optimization problem. Strategy information modifies the relative importance of coverage, equity, and economic terms in the GA fitness function by adjusting their scalar weights. The structure of the fitness function remains fixed; only the strategy-dependent weights vary. The executor adapts and runs code modules for initialization, crossover, mutation, and fitness evaluation to search the combinatorial space of candidate subsets. The validator applies the same performance criteria to ensure comparability across methods and robustness of the results. This workflow highlights the LLM’s ability to adapt algorithmic components dynamically, enabling strategy-aware exploration of siting configurations beyond the scope of simple ranking.
4.2. Prompt Engineer and Reflective Loop
Prompt engineering in the framework draws on structured prompting and chain-of-thought reasoning [
24] and is organized around three roles: task planner, task executor, and validator. This design also resonates with the principle of combining reasoning and acting [
25], where planning and tool use are interleaved to improve task reliability. Each stage is guided by prompts that specify the goal, critical input information, expected workflow, and structured output, allowing the LLM to act systematically rather than through free-form generation.
As shown in
Figure 3, the task planner uses the reasoning ability of the LLM to decompose high-level planning goals into executable subtasks. This step aligns the planning strategy with available inputs and specifies indicators, constraints, and algorithmic settings in a structured manner. The task executor implements these subtasks by refining or generating code and by invoking domain-specific tools. In the MCDA pipeline, this involves function calls for normalization, scoring, and site selection, while in the GA pipeline, it requires adapting algorithmic operators and parameters. The validator checks whether the selected sites satisfy spacing constraints, land-use feasibility, and the strategy’s quantitative objectives before accepting the solution. It inspects the outputs against requirements, evaluates performance using metrics such as demand coverage, equity, ROI, and accessibility, and provides explicit pass/fail feedback with reasons.
To enhance robustness, we embed a reflective loop connecting planner, executor, and validator. Feedback from the validator is returned to the planner, which revises subtasks and prompts as needed, enabling iterative refinement. The loop resolves cases where intermediate outputs lack required parameters for execution, and the planner adjusts the task specification while the underlying datasets remain fixed. This loop prevents error propagation, adapts to missing or inconsistent data, and ensures that the siting outputs remain aligned with the planning strategy. Once a solution is validated as satisfactory, the agent proceeds to visualization and quantitative assessment, providing both interpretable outcomes and measurable performance of the site selection plan.
4.3. Function Calling
Function calling provides a structured interface for the agent to translate planning strategies into computational workflows. Rather than producing ad hoc scripts, the LLM invokes predefined functions that encapsulate key tasks such as candidate screening, indicator integration, spatial feasibility checks, visualization, and performance evaluation. This modular approach grounds siting decisions in reproducible procedures while maintaining alignment with policy objectives.
In the MCDA pipeline, function calling is central to operationalization. Candidate parcels are filtered by eligibility, and multiple indicators are normalized and aggregated through the data integration function. This module incorporates outlier-robust scaling and benefit–cost orientation, ensuring that composite scores reflect both statistical robustness and explicit planning priorities. Downstream, visualization functions enforce spatial spacing rules and generate interpretable maps, producing an auditable link from weights and constraints to siting outcomes. By contrast, the GA workflow emphasizes code execution, with optimization routines adapted through the planner but still relying on common evaluation functions to ensure comparability of results.
Analysis functions provide a unified evaluation layer across both pipelines, computing economic, accessibility, and equity metrics from selected sites. Encapsulating these operations within callable functions ensures that assessments remain consistent and transparent. In the smart city context, this design enhances the portability of site selection methods across urban scenarios and creates a verifiable bridge between textual planning strategies and algorithmic decision support.
5. Experiment
5.1. Data Source
The empirical analysis focuses on the Los Angeles metropolitan area, selected for its polycentric structure, high travel demand, and strong policy relevance for early-stage UAM deployment. The data used is shown in
Figure 4. OD travel demand and network data are drawn from ETH Zurich’s Los Angeles transport dataset [
26], which provides detailed OD flows and a calibrated road network suitable for microsimulation. Ground travel times are benchmarked with the Google Maps Distance Matrix API under peak-hour conditions, while eVTOL travel times are estimated from Euclidean distances combined with assumed cruise speeds and terminal operations. OD pairs where eVTOL offers at least a 40% reduction in travel time [
27] are retained as demand-generating pairs, ensuring that demand estimation emphasizes cases of substantial travel-time advantage.
Tract-level household income data are obtained from the U.S. Census Bureau and incorporated into demand modeling to capture socioeconomic heterogeneity in adoption potential. Land-use screening of candidate parcels relies on the Southern California Association of Governments (SCAG) parcel-level land-use inventory, which defines eligibility according to zoning and development categories. Transit accessibility is represented by bus and rail stop data released in GTFS format by LA Metro, enabling evaluation of first- and last-mile connectivity around candidate sites. Together, these datasets provide the socioeconomic, spatial, and transport-system context required to evaluate both the feasibility and the equity implications of vertiport siting strategies.
5.2. Outcome Evaluation
To evaluate whether the proposed agent can reason effectively under diverse planning strategies, we design outcome experiments that align siting decisions with policy-relevant objectives. The purpose is to test whether the agent can translate high-level strategies into quantitative criteria and deliver site configurations that exhibit consistent trade-offs across coverage, equity, efficiency, and accessibility. The comparison across strategies provides a basis for examining how different planning priorities redistribute performance across coverage, equity, economic return, and accessibility.
To achieve this, we define three representative planning strategies with distinct priorities. The coverage-oriented strategy aims to extend service reach, combining central hubs with peripheral sites for broad inclusion. The equity-oriented strategy seeks to reduce disparities by prioritizing underserved neighborhoods while retaining core connectivity. The efficiency-oriented strategy emphasizes economic return, concentrating facilities in high-demand, well-connected areas to keep the network cost-effective. Each strategy is translated by the agent into weighted indicators and constraints, with outcomes compared across MCDA and GA pipelines. The baseline represents a conventional MCDA workflow in which criteria, thresholds and weights are fixed manually according to established planning guidelines, and candidate sites are produced through standard geospatial screening without any language model involvement. Evaluation uses four metrics—demand coverage, equity (Gini coefficient), economy (ROI), and accessibility (transit stop coverage)—allowing assessment of both siting quality and the agent’s ability to align textual planning prompts with quantitative outcomes.
As shown in
Table 1, the MCDA pipeline exhibits clear alignment between strategies and outcomes. The coverage-oriented strategy achieves the highest demand coverage (24.04%) and accessibility (0.100), reflecting its explicit focus on extending service reach. The equity-oriented strategy produces the lowest Gini coefficient (0.367), indicating improved fairness in distribution across income groups. The efficiency-oriented strategy returns the highest ROI (0.873), consistent with its emphasis on economic performance. These spatial patterns are also illustrated in
Figure 5, where the distribution of facilities under different strategies visually reflects the corresponding performance outcomes. These shifts across strategies suggest that MCDA effectively operationalizes textual planning goals into quantifiable siting outcomes. By contrast, the genetic algorithm shows weaker differentiation across strategies. Coverage-oriented runs again yield the highest demand coverage (24.04%) and accessibility (0.128), but equity- and efficiency-oriented strategies do not demonstrate similarly distinct responses. For example, the Gini coefficients (0.309–0.303) and ROIs (0.910–0.353) are relatively close across strategies, suggesting that GA emphasizes global optimization but is less sensitive to the explicit strategic framing provided in prompts.
To further assess robustness, we evaluate three LLM backbones—Gemini 2.5 Flash, Qwen2.5-32B-Instruct, and DeepSeek V3.1—under two prompt settings that differ only in whether ground transportation is explicitly incorporated. The selection reflects a cross-section of contemporary LLM architectures that are widely used in planning-related agent workflows. Gemini 2.5 Flash is a lightweight, high-speed variant designed for efficient reasoning with moderate context capacity. Qwen2.5-32B-Instruct represents a larger-scale instruction-tuned model with stronger performance on structured tasks. DeepSeek V3.1 is a recent open-source backbone optimized for reasoning and planning. Together, these models cover different scales, training paradigms, and accessibility settings, forming a representative testbed for evaluating whether the strategy-aware orchestration layer remains stable across heterogeneous backbones while keeping computational requirements manageable for repeated planner–executor–validator cycles.
In the with-ground setting, the planner prompt requires integration with transit hubs and traffic flows, accounting for first/last-mile penalties and accessibility constraints. In the without-ground setting, these cues are removed, and the agent optimizes aerial siting without explicit surface-network considerations. By holding data, tools, and evaluation constant, this design isolates the effect of transportation integration on outcomes and enables cross-model comparison. The purpose is to determine whether differences in reasoning or grounding capabilities affect the mapping from strategies to siting outcomes and whether the reflective loop mitigates variability across LLMs.
The purpose is to test prompt sensitivity and verify that integration signals propagate into indicator choice, constraint formation, and tool use. We evaluate the same four metrics (demand coverage, equity/Gini, ROI, accessibility). The results in
Table 2 indicate that ground integration prompts do affect siting outcomes, though the impact is not consistently positive across all metrics. For Gemini 2.5 Flash, incorporating ground transportation considerations increases coverage and accessibility but comes with reduced equity and ROI. Qwen2.5-32B-Instruct remains relatively stable across both settings, suggesting limited responsiveness to integration but consistently high ROI performance. DeepSeek V3.1 shows a trade-off where ground integration improves equity but reduces coverage and economic return. These findings suggest that the influence of ground transportation is multifaceted: it can shift the balance among coverage, equity, efficiency, and accessibility, but not in a uniformly beneficial direction.
5.3. Process Evaluation
To evaluate the internal reliability of the agent, we design a process evaluation that focuses on its ability to orchestrate tool use and iterative reasoning. The aim is not only to measure whether final siting outcomes are reasonable but also to quantify how effectively the agent manages the intermediate workflow of planning, execution, and validation. Four indicators are defined: (1) Function Calling Accuracy measures whether the agent correctly selects the intended function and provides valid parameters; this reflects the precision of mapping planning instructions into tool operations; (2) Function Calling Success Rate records whether the function call returns a valid computational result, capturing robustness against runtime errors; (3) Delivery Rate evaluates whether the agent is able to produce a usable output for the user regardless of validation, representing end-to-end operability; (4) Perfect Pass measures the fraction of cases where the validator accepts the result in a single iteration, reflecting the capacity of the reflective loop to converge efficiently without further adjustment. Function Calling Accuracy is evaluated by comparing each generated call with the full schema required by the planner prompt. A call is counted as accurate when its function name and all mandatory parameters are complete, well-typed, and aligned with the expected structure. This process evaluation provides an internal perspective on orchestration quality.
To evaluate process performance, we compare three usage paradigms: zero-shot prompting, an ablated agent without the reflective loop, and the strategy-aware agent workflow. Each paradigm is assessed through indicators that capture whether the agent can reliably transform planning strategies into executable outcomes. Because MCDA depends primarily on structured function calling, its evaluation emphasizes Function Calling Accuracy, which measures whether the correct functions and parameters are selected, and Perfect Pass, which records the share of cases validated in a single iteration. By contrast, GA requires code generation and modification, so Delivery Rate, reflecting whether the generated code can be executed without errors, and Perfect Pass, indicating validator-approved results on the first attempt, are more representative.
Table 3 highlights that the strategy-aware agent achieves the strongest performance, with higher accuracy and a greater share of validated results across both MCDA and GA workflows. The ablated agent without the reflective loop also shows improvement over zero-shot prompting, as it preserves the structured planner–executor–validator workflow even without iterative feedback. These results suggest that while workflow structure already brings notable gains compared with direct prompting, the addition of reflective feedback further strengthens reliability and alignment with planning objectives.
To evaluate whether the proposed framework functions consistently across different backbone models, we test Gemini 2.5 Flash, Qwen2.5-32B-Instruct, and DeepSeek V3.1 using the same process indicators. The purpose is to examine whether the reflective agent workflow depends strongly on the underlying LLM or whether it generalizes across models. As shown in
Table 4, all three models achieve comparable levels of function calling accuracy, success rate, and delivery, with Qwen2.5-32B-Instruct and DeepSeek V3.1 slightly higher on Perfect Pass. These results suggest that the framework provides a stable orchestration layer that can be applied to multiple LLMs with only modest variation in performance.
6. Discussion
6.1. Strategy-Aware
The experiments suggest that LLMs can act as a bridge between high-level planning strategies and quantitative siting outcomes. The observed alignment in the MCDA runs illustrates that when the agent is asked to prioritize coverage, equity, or efficiency, the LLM is able to operationalize this textual framing into corresponding indicator weights and constraints. This shows that the model is not only executing predefined functions but also interpreting the intent of the strategy and mapping it into computationally meaningful parameters. This demonstrates the ability of LLM agents to transform qualitative specifications into structured decision logic, a key feature in planning tasks [
28]. This aligns with recent typologies of LLM-based multi-agent systems that emphasize structured workflows, role decomposition and iterative coordination [
29].
The weaker differentiation in GA outcomes highlights a limitation in how strategy signals propagate through more complex workflows. Here, the LLM is tasked with modifying or generating code to implement constraints and objectives, and the stochasticity of evolutionary search can dilute the explicit link between prompt and outcome. Rather than a drawback of GA per se, this reflects the current challenge of aligning free-text strategies with algorithmic operators when the mapping is less direct. These results underscore that the value of LLMs in site selection lies not only in producing feasible solutions but also in maintaining traceability between policy prompts and technical execution—something that requires careful design of interfaces between natural-language input and optimization procedures. The empirical setup is built around a single metropolitan region with constrained data variation, and the experimental design emphasizes how planning strategies are expressed through optimization outcomes. Under this structure, the tables present how variations in strategic intent shape the resulting siting patterns and how these differences are reflected in the quantitative indicators.
The results also make the structure of the trade-offs across strategies more explicit. In the MCDA runs, prioritizing coverage leads to broader spatial reach and higher accessibility, but this comes with reduced economic return compared with the efficiency-oriented strategy. The equity-oriented strategy improves distributional outcomes but reduces both coverage and accessibility relative to the coverage-oriented runs. These shifts reflect how the same set of parcels is redistributed when emphasis moves among spatial inclusion, socioeconomic fairness, and economic performance. Across the GA runs, the direction of these trade-offs becomes weaker, since the search process balances objectives in a way that reduces the separation among strategies. This pattern indicates that the agent can express strategy signals clearly in MCDA, while in GA, the interaction between stochastic search and code-level adaptation makes the propagation of strategic intent less direct.
6.2. Ground Transportation Integration
The integration of ground transportation emerges as an important factor in shaping siting outcomes. By explicitly embedding transit hubs, traffic flows, and first/last-mile penalties into the prompts, the agent is encouraged to consider multimodal accessibility rather than focusing solely on aerial efficiency. The experimental results suggest that these integration signals do influence outcomes, but their effects differ across models and metrics. For instance, some models show improved coverage or accessibility when ground transportation is included, while others demonstrate shifts in equity or economic return. Recent work on vertiport siting confirms that multimodal access and ground-infrastructure integration materially influence feasible UAM catchment areas [
30]. These patterns indicate that transportation integration reshapes the relative strengths of the strategies rather than reinforcing a single dominant direction and that its influence depends on the backbone model and the metric considered.
These variations highlight that the role of ground integration is not to guarantee uniformly superior outcomes but to redirect the balance among competing objectives. They also illustrate how the incorporation of external transport-related inputs can broaden the reasoning space of LLM agents, enriching the search process without prescribing a uniform direction of improvement [
31]. In this sense, the framework makes the effect of integration transparent, enabling planners to observe how multimodal considerations translate into siting trade-offs. Taken together, the results suggest that prompting for transportation integration does matter, though the exact direction of impact depends on both the evaluation metric and the model’s responsiveness. This reinforces the value of LLM agents as exploratory tools, capable of surfacing how different strategic emphases, including multimodal alignment, manifest in quantifiable outcomes.
6.3. Function of the Reflective Loop
The reflective loop emerges as a central mechanism that enables LLM agents to move from fragile, one-shot reasoning toward adaptive and accountable workflows. Without the loop, the agent may produce outputs that are syntactically valid—for example, executable code in GA or a computed score table in MCDA—but fail to satisfy the strategic intent or evaluation criteria. With the loop, validator feedback is routed back to the planner, prompting revisions of indicators, parameters, or constraints until outcomes realign with planning goals. This iterative process allows the LLM to act as an orchestrator of reasoning cycles, not merely a one-time generator. It also reduces reliance on model scale by grounding reasoning in tool-based operations and validation. As a result, incomplete or inconsistent intermediate outputs are corrected before site selection advances. Comparable architectures that combine chain-of-thought prompting with multi-agent tool orchestration have demonstrated improvements in planning reliability and interpretability [
32].
As shown in
Table 3, the benefits of the reflective loop are not fully captured by improvements in delivery rate. Rather, its more meaningful impact lies in Perfect Pass. In the GA pipeline especially, omitting the loop may increase the proportion of attempts that “execute” (i.e., delivery rate), but many of those runs fail validation or misalign with strategic criteria. The reflective loop increases the proportion of runs that pass validation on the first attempt (Perfect Pass), improving the consistency and strategic accuracy of the final outputs. This mechanism mirrors broader evidence that iterative self-reflection enhances the quality and stability of problem-solving in LLM agents [
33]. In location site selection, where multiple data sources, objectives, and constraints must be integrated, this property is vital. The reflective loop helps to catch and correct early misinterpretations, embedding domain oversight into the agent’s reasoning process. More broadly, it showcases how LLM-based systems can support deliberative decision-making—where feedback, correction, and adaptability are built in—suiting planning domains where accountability, interpretability, and robustness are as critical as computational performance. The two pipelines illustrate how the framework instantiates strategy-aware workflows through different methodological forms. MCDA expresses strategies through indicator weighting and ranking, while GA expresses strategies inside a constrained combinatorial search. The evaluation examines how each method reflects the planning strategy within its own workflow.
6.4. Future Direction and Limitations
This study demonstrates the potential of LLM-based agents to operationalize planning strategies into siting decisions, but several limitations provide avenues for future work. The first concerns data resolution and scope. The experiments relied on tract-level socioeconomic data, aggregate OD flows, and simplified eVTOL performance assumptions; finer-grained datasets on household travel behavior, multimodal accessibility, and infrastructure costs would allow more nuanced evaluation of siting outcomes. A second limitation lies in the representation of demand: while time savings were used as a proxy for adoption potential, incorporating behavioral models of traveler choice could provide more realistic forecasts of uptake. A third limitation relates to the type of travel-demand data used. The OD flows in this study are derived from taxi trips, which capture a specific segment of urban travel behavior. In contexts where commuting or other systematic travel patterns dominate potential eVTOL demand, the use of home-to-work OD datasets or household travel surveys could yield different spatial distributions and support a more comprehensive assessment of siting performance. These constraints do not undermine the methodological contribution but highlight the importance of richer empirical inputs for practical deployment. The generalizability of the workflow also depends on the structure of the local urban system. Applying the framework to cities with different land-use patterns, multimodal networks, and regulatory contexts may require additional adaptation, and cross-city experiments would clarify how transferable the agent’s reasoning and tool orchestration are. Furthermore, real-world deployment involves challenges that go beyond the computational pipeline, including integration with existing planning procedures and the need to maintain oversight when tool-derived constraints are translated into practice. These limitations do not diminish the methodological contribution but underscore the need for richer data in practice.
Future research can extend the framework in several directions. To provide a stronger empirical benchmark, future work could incorporate machine learning models, such as Random Forest or Neural Networks, to capture complex nonlinear patterns and compare their data-driven predictions with the strategy-aware outcomes of LLM-orchestrated workflows. To enhance generalizability, the workflow could be tested in different metropolitan contexts, where variations in land use, transit integration, and socioeconomic conditions would stress-test the adaptability of the agent. To deepen strategy alignment, more sophisticated mechanisms could be developed for mapping policy objectives into optimization formulations, particularly for algorithms like GA, where the link between text and outcome is less direct. Finally, advancing interfaces between LLM agents and existing planning software would improve transparency and adoption in professional practice, enabling planners to experiment with alternative strategies while retaining oversight of the decision logic. Future extensions that embed the workflow into broader planning toolchains—such as transportation modeling platforms or visualization environments—could further clarify how the agent supports end-to-end decision processes in real applications.
7. Conclusions
This study proposed a strategy-aware LLM-based framework for vertiport site selection in urban air mobility, embedding ground transportation integration and iterative validation into the planning pipeline. The framework combines heterogeneous datasets—covering land use, demand, income, and transit—with planning strategies expressed in natural language and translates them into structured constraints and tool-based operations. A reflective loop linking planner, executor, and validator ensures that outputs are iteratively refined until they align with strategic objectives, providing a systematic mechanism for transparent and adaptive siting.
The experimental evaluation in Los Angeles demonstrates that the framework can reproduce strategy-aware outcomes, particularly in the MCDA pipeline where coverage-, equity-, and efficiency-oriented strategies map into distinct siting patterns and quantitative trade-offs. The GA pipeline highlights the potential of evolutionary optimization but also underscores the challenge of maintaining strong alignment between textual strategies and algorithmic operators. Across models, results show that including ground transportation considerations shifts siting outcomes in meaningful but non-uniform ways, reflecting the inherent complexity of balancing aerial efficiency with multimodal accessibility. Process evaluations further indicate that the reflective loop substantially improves orchestration quality, not simply by increasing delivery rates but by raising the proportion of results validated on the first iteration, which strengthens both robustness and accountability.
By demonstrating that LLM agents can formalize qualitative strategies into quantitative siting outcomes, this work highlights their potential as interpretable and adaptive components of urban planning workflows. For real-world implementation, this framework can function as an interactive decision-support tool, allowing planners to conduct rapid “what-if” analyses by testing different natural-language strategies and instantly observing the resulting siting trade-offs. Its natural-language interface also makes it suitable for public participation, helping communicate complex planning scenarios to stakeholders. Furthermore, the agent’s orchestration layer could be integrated as a module within existing Geographic Information Systems (GIS) or transportation planning software, translating high-level policy into executable workflows and thus bridging the gap between strategic intent and technical execution.
While demonstrated here for vertiport planning, the approach points more broadly to how LLMs may support transparent, adaptive, and accountable decision-making in urban systems, contributing to the development of intelligent planning support for future smart cities.
Author Contributions
Conceptualization, Y.J. and J.M.; methodology, Y.J. and J.M.; software, Y.J. and J.M.; validation, Y.J. and J.M.; formal analysis, Y.J. and J.M.; investigation, Y.J. and J.M.; resources, Y.J. and J.M.; data curation, Y.J. and J.M.; writing—original draft preparation, Y.J.; writing—review and editing, Y.J. and J.M.; visualization, Y.J. and J.M.; supervision, J.M.; project administration, Y.J. and J.M. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Data Availability Statement
Data will be made available on request.
Acknowledgments
During the preparation of this work, the authors used Claude Sonnet 4.5 to improve readability and language. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.
Conflicts of Interest
The authors declare no conflict of interest.
References
- Sengupta, R.; Bulusu, V.; Mballo, C.E.; Onat, E.B.; Cao, S. Urban air mobility research challenges and opportunities. Annu. Rev. Control. Robot. Auton. Syst. 2025, 8, 407–431. [Google Scholar] [CrossRef] [Scilit]
- Yan, Y.; Wang, K.; Qu, X. Urban air mobility (UAM) and ground transportation integration: A survey. Front. Eng. Manag. 2024, 11, 734–758. [Google Scholar] [CrossRef] [Scilit]
- Yoon, D.; Jeong, M.; Lee, J.; Kim, S.; Yoon, Y. Integrating Urban Air Mobility with Highway Infrastructure: A Strategic Approach for Vertiport Location Selection in the Seoul Metropolitan Area. arXiv 2025, arXiv:2502.00399. [Google Scholar] [CrossRef] [Scilit]
- Yu, Y.; Wang, M.; Mesbahi, M.; Topcu, U. Vertiport selection in hybrid air–ground transportation networks via mathematical programs with equilibrium constraints. IEEE Trans. Control. Netw. Syst. 2023, 10, 2108–2119. [Google Scholar] [CrossRef] [Scilit]
- Jin, Z.; Ng, K.K.; Zhang, C. Robust optimisation for vertiport location problem considering travel mode choice behaviour in urban air mobility systems. J. Air Transp. Res. Soc. 2024, 2, 100006. [Google Scholar] [CrossRef] [Scilit]
- Jiang, X.; Tang, Y.; Cao, J.; Bulusu, V.; Yang, H.; Peng, X.; Sengupta, R. Simulating integration of urban air mobility into existing transportation systems: Survey. J. Air Transp. 2024, 32, 97–107. [Google Scholar] [CrossRef] [Scilit]
- Garrow, L.A.; German, B.J.; Leonard, C.E. Urban air mobility: A comprehensive review and comparative analysis with autonomous and electric ground transportation for informing future research. Transp. Res. Part C Emerg. Technol. 2021, 132, 103377. [Google Scholar] [CrossRef] [Scilit]
- Kotwicz Herniczek, M.T.; German, B.J. Scalable Combinatorial Vertiport Placement Method for an Urban Air Mobility Commuting Service. Transp. Res. Rec. 2024, 2678, 1624–1641. [Google Scholar] [CrossRef] [Scilit]
- Zheng, Y.; Xu, F.; Lin, Y.; Santi, P.; Ratti, C.; Wang, Q.R.; Li, Y. Urban planning in the era of large language models. Nat. Comput. Sci. 2025, 5, 727–736. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Li, J.; Wang, Z.; He, Z.; Guan, Q.; Lin, J.; Yu, W. Geospatial large language model trained with a simulated environment for generating tool-use chains autonomously. Int. J. Appl. Earth Obs. Geoinf. 2025, 136, 104312. [Google Scholar] [CrossRef] [Scilit]
- Straubinger, A.; Rothfeld, R.; Shamiyeh, M.; Büchter, K.D.; Kaiser, J.; Plötner, K.O. An overview of current research and developments in urban air mobility–Setting the scene for UAM introduction. J. Air Transp. Manag. 2020, 87, 101852. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Yang, C.; Liu, J.; Jones, S. Assessing the Feasibility and Viability of Multimodal Urban Air Mobility in Small and Medium-Sized Urban Areas: Integration of Shared Autonomous Vehicles and Vertical Take-Off and Landing Vehicles. Available online: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4995752 (accessed on 26 November 2025).
- Coppola, P.; De Fabiis, F.; Silvestri, F. Urban Air Mobility demand forecasting: Modeling evidence from the case study of Milan (Italy). Eur. Transp. Res. Rev. 2025, 17, 2. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Z.; Lin, Y.; Li, Y. Large language model empowered participatory urban planning. arXiv 2024, arXiv:2402.01698. [Google Scholar] [CrossRef] [Scilit]
- Han, J.; Ning, Y.; Yuan, Z.; Ni, H.; Liu, F.; Lyu, T.; Liu, H. Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications. arXiv 2025, arXiv:2507.00914. [Google Scholar] [CrossRef] [Scilit]
- Zhu, H.; Zhang, W.; Huang, N.; Li, B.; Niu, L.; Fan, Z.; Liu, X. PlanGPT: Enhancing urban planning with tailored language model and efficient retrieval. arXiv 2024, arXiv:2402.19273. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Wei, C.; He, Z.; Yu, W. GeoGPT: An assistant for understanding and processing geospatial tasks. Int. J. Appl. Earth Obs. Geoinf. 2024, 131, 103976. [Google Scholar] [CrossRef] [Scilit]
- Li, Z.; Ning, H. Autonomous GIS: The next-generation AI-powered GIS. Int. J. Digit. Earth 2023, 16, 4668–4686. [Google Scholar] [CrossRef] [Scilit]
- Jin, Y.; Ma, J. Large language model as parking planning agent in the context of mixed period of autonomous vehicles and Human-Driven vehicles. Sustain. Cities Soc. 2024, 117, 105940. [Google Scholar] [CrossRef] [Scilit]
- Peng, B.; Wang, Y.; Feng, C.; Xia, X.; Li, P. UrbanMUDA: An LLM Agent-based Site Selection Approach for Urban Military Unit Deployment. In Proceedings of the 2025 IEEE 26th China Conference on System Simulation Technology and Its Applications (CCSSTA), Shenzhen, China, 11–13 July 2025; pp. 235–240. [Google Scholar]
- Koda, M.; Zheng, Y.; Ma, R.; Sun, M.; Pansare, D.; Duarte, F.; Santi, P. LocationReasoner: Evaluating LLMs on Real-World Site Selection Reasoning. arXiv 2025, arXiv:2506.13841. [Google Scholar]
- Ishizaka, A.; Nemery, P. Multi-Criteria Decision Analysis: Methods and Software; John Wiley & Sons: Hoboken, NJ, USA, 2013. [Google Scholar]
- Holland, J.H. Genetic algorithms. Sci. Am. 1992, 267, 66–73. [Google Scholar] [CrossRef] [Scilit]
- Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q.; Zhou, D. Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inf. Process. Syst. 2022, 35, 24824–24837. [Google Scholar]
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. React: Synergizing reasoning and acting in language models. In Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- LAmb¨uhl, M.; Menendez, M.C. Gonz’alez, Understanding congestion propagation by combining percolation theory with the macroscopic fundamental diagram. Commun. Phys. 2023, 6, 26. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Y.; Feng, T. Strategic integration of vertiport planning in multimodal transportation for urban air mobility: A case study in Beijing, China. J. Clean. Prod. 2024, 467, 142988. [Google Scholar] [CrossRef] [Scilit]
- Li, X. A review of prominent paradigms for llm-based agents: Tool use, planning (including rag), and feedback learning. In Proceedings of the 31st International Conference on Computational Linguistics, Abu Dhabi, United Arab Emirates, 19–24 January 2025; pp. 9760–9779. [Google Scholar]
- Li, X.; Wang, S.; Zeng, S.; Wu, Y.; Yang, Y. A survey on LLM-based multi-agent systems: Workflow, infrastructure, and challenges. Vicinagearth 2024, 1, 9. [Google Scholar] [CrossRef] [Scilit]
- Lu, Y.; Zeng, W.; Wei, W.; Wu, W.; Jiang, H. Vertiport Location Selection and Optimization for Urban Air Mobility in Complex Urban Scenes. Aerospace 2025, 12, 709. [Google Scholar] [CrossRef] [Scilit]
- Huang, X.; Liu, W.; Chen, X.; Wang, X.; Wang, H.; Lian, D.; Wang, Y.; Tang, R.; Chen, E. Understanding the planning of LLM agents: A survey. arXiv 2024, arXiv:2402.02716. [Google Scholar] [CrossRef] [Scilit]
- Cao, H.; Ma, R.; Zhai, Y.; Shen, J. Llm-collab: A framework for enhancing task planning via chain-of-thought and multi-agent collaboration. Appl. Comput. Intell. 2024, 4, 328–348. [Google Scholar] [CrossRef] [Scilit]
- Renze, M.; Guven, E. Self-reflection in llm agents: Effects on problem-solving performance. arXiv 2024, arXiv:2405.06682. [Google Scholar]
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).