GeoPPO—A Location-Allocation Method of Superstores Based on Deep Reinforcement Learning—A Case Study of Xi’an
Abstract
1. Introduction
- Contribution One: Proposed a new DW-MCLP spatial optimization model by integrating multi-source data, including population, point-of-interest distribution [15], and urban transportation networks, to characterize spatial demand and constraints.
- Contribution Two: Provided a feasible technical path for handling highly dynamic urban facility spatial configuration with geographic mask constraints [16]. The application in Xi’an demonstrates that the framework can identify location plans aligned with the city’s evolving multi-center spatial structure under the given constraints.
- Contribution Three: Developed a novel deep reinforcement learning-based solving framework named GeoPPO [17]. This framework, integrating an encoder-decoder architecture [18] with a dynamic coverage information gating mechanism, demonstrates superior solving speed and solution quality compared to Vanilla PPO, GA, and PSO, offering efficient and adaptive decision support for complex, large-scale urban spatial optimization problems.
2. Related Work
3. Materials and Methods
3.1. Data Sources
- Population density data: obtained from the LandScan Global 2023 30-arcsecond population distribution dataset released by Oak Ridge National Laboratory, Oak Ridge, TN, USA [30].
- Administrative division data: obtained from Tianditu Open Platform (National Platform for Common Geospatial Information Services, Beijing, China).
- Road traffic network data: obtained from the OpenStreetMap (OSM) national road network dataset [31], containing the spatial location and topological information of all levels of roads in Xi’an.
- Commercial POI data: retrieved and collected through the Gaode Map API (Alibaba Group, Hangzhou, China) in June 2025, covering 10 types of commercial facility points such as shopping centers, communities, and transportation hubs in the study area. All POI data were spatially located based on the coordinate information provided by Gaode Map.
3.2. Study Area
- 1.
- Stable and large-scale customer traffic base
- 2.
- Young Consumers
- 3.
- Mature transportation system
3.3. Environmental Modeling
3.3.1. Spatial Grid
3.3.2. Population Density
3.3.3. Demand Analysis
3.3.4. Competition Intensity and Service Pressure
3.3.5. Feasible Point Selection
3.4. Methods
- 1.
- Data fusion and environment modeling.
- 2.
- Reinforcement learning decision mechanism.
- 3.
- Layered training and balanced optimization.
- 4.
- Decision support and empirical verification.
3.5. Problem Model Elaboration
- (i)
- Geospatial Optimization Location Selection Model
- (ii)
- Reward function design
3.6. Introduction to DRL Methods
3.6.1. Overall Architecture: Actor-Critic Spatial Decision System
- (i)
- Input Representation
- (ii)
- Output representation
- (iii)
- Objective function
3.6.2. Decision-Making Process
- Step 1:
- Spatial feature extraction—convolutional coding
- Step 2:
- State Action Space Generation
- where areas overlapping with the existing supermarket location are excluded from the candidate actions to avoid duplicate construction.
- Step 3:
- Strategy Sampling and Reward Calculation
- Step 4:
- GeoPPO-Clip strategy update
- Step 5:
- Curriculum Learning Optimization
4. Results
4.1. Analysis of Spatial Distribution Pattern
4.2. Location Selection Experiment

5. Discussion
5.1. Technical Context
5.2. Interpretation of Empirical Findings
5.3. Output Results
5.4. Limitations
6. Conclusions
6.1. Summary of Contributions
6.2. Outlook and Future Directions
- 1.
- Advancing Predictive Modeling with Dynamic Urban Data
- 2.
- Extending Applications to Larger Regions and Different Facility Types
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- State Council of the People’s Republic of China. Approval of Xi’an Territorial Spatial Master Plan (2021–2035). State Council Document No. 13. 2025. Available online: https://www.gov.cn/zhengce/zhengceku/202501/content_7000441.htm (accessed on 1 February 2025).
- Caixin Global. In Depth: E-Commerce Giants Shun Supermarkets as Sales Slump. 16 January 2025. Available online: https://www.caixinglobal.com/2025-01-16/in-depth-e-commerce-giants-shun-supermarkets-as-sales-slump-102279780.html (accessed on 1 February 2025).
- Larson, J.S.; Bradlow, E.T.; Fader, P.S. An exploratory look at supermarket shopping paths. Int. J. Res. Mark. 2005, 22, 395–414. [Google Scholar] [CrossRef]
- Harris, B.; Batty, M. Locational models, geographic information and planning support systems. J. Plan. Educ. Res. 1993, 12, 184–198. [Google Scholar] [CrossRef]
- Oliveira, V.; Pinho, P. Evaluation in urban planning: Advances and prospects. J. Plan. Lit. 2010, 24, 343–361. [Google Scholar] [CrossRef]
- Arulkumaran, K.; Deisenroth, M.P.; Brundage, M.; Bharath, A.A. Deep reinforcement learning: A brief survey. IEEE Signal Process. Mag. 2017, 34, 26–38. [Google Scholar] [CrossRef]
- Liang, H.; Wang, S.; Li, H.; Zhou, L.; Chen, H.; Zhang, X.; Chen, X. Sponet: Solve spatial optimization problem using deep reinforcement learning for urban spatial decision analysis. Int. J. Digit. Earth 2024, 17, 2299211. [Google Scholar] [CrossRef]
- Liang, H.; Wang, S.; Li, H.; Pan, J.; Li, X.; Su, C.; Liu, B. AIAM: Adaptive interactive attention model for solving p-Median problem via deep reinforcement learning. Int. J. Appl. Earth Obs. Geoinf. 2025, 138, 104454. [Google Scholar] [CrossRef]
- Panait, L.; Luke, S. Cooperative multi-agent learning: The state of the art. Auton. Agents Multi-Agent Syst. 2005, 11, 387–434. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł; Polosukhin, I. Attention is all you need. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30, pp. 5998–6008. [Google Scholar]
- Beheshti, Z.; Shamsuddin, S.M.H. A review of population-based meta-heuristic algorithms. Int. J. Adv. Soft Comput. Appl. 2013, 5, 1–35. [Google Scholar]
- Baviera-Puig, A.; Buitrago-Vera, J.; Escriba-Perez, C. Geomarketing models in supermarket location strategies. J. Bus. Econ. Manag. 2016, 17, 1205–1221. [Google Scholar] [CrossRef]
- Cui, Y.; Yu, Y.; Cai, Z.; Wang, D. Optimizing road network density considering automobile traffic efficiency: Theoretical approach. J. Urban Plan. Dev. 2022, 148, 04021062. [Google Scholar] [CrossRef]
- Liu, K.; Yin, L.; Lu, F.; Mou, N. Visualizing and exploring POI configurations of urban regions on POI-type semantic space. Cities 2020, 99, 102610. [Google Scholar] [CrossRef]
- Psyllidis, A.; Gao, S.; Hu, Y.; Kim, E.K.; McKenzie, G.; Purves, R.; Yuan, M.; Andris, C. Points of Interest (POI): A commentary on the state of the art, challenges, and prospects for the future. Comput. Urban Sci. 2022, 2, 20. [Google Scholar] [CrossRef]
- Lin, P.C.; Cheng, T.C.E.; Hsu, C.H. Retail location modeling of supermarket chains in Taipei city. Appl. Geogr. 2023, 161, 103126. [Google Scholar] [CrossRef]
- Engstrom, L.; Ilyas, A.; Santurkar, S.; Tsipras, D.; Janoos, F.; Rudolph, L.; Madry, A. Implementation matters in deep policy gradients: A case study on ppo and trpo. arXiv 2020, arXiv:2005.12729. [Google Scholar] [CrossRef]
- Jaderberg, M.; Simonyan, K.; Zisserman, A.; Kavukcuoglu, K. Spatial transformer networks. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2015; Volume 28, pp. 2017–2025. [Google Scholar] [CrossRef]
- Alnahhal, M.; Noche, B. A genetic algorithm for supermarket location problem. Assem. Autom. 2015, 35, 122–127. [Google Scholar] [CrossRef]
- Guo, W.; Xu, Y.; Jin, Y. Swap-based deep reinforcement learning for facility location problems in networks. arXiv 2023, arXiv:2312.15658. [Google Scholar] [CrossRef]
- Wei, H.; Liu, X.; Ying, L. Triple-q: A model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation. In International Conference on Artificial Intelligence and Statistics (AISTATS 2022); PMLR: New York, NY, USA, 2022; Volume 151, pp. 3274–3307. [Google Scholar]
- Su, H.; Zheng, Y.; Ding, J.; Jin, D.; Li, Y. Large-scale Urban Facility Location Selection with Knowledge-informed Reinforcement Learning. In Proceedings of the 32nd ACM International Conference on Advances in Geographic Information Systems (ACM SIGSPATIAL 2024), 29 October–1 November 2024; ACM: New York, NY, USA, 2024; pp. 553–556. [Google Scholar]
- Wang, S.; Xu, D.; Zhou, J.; Su, C.; Li, X.; Liang, X.; Cao, C.; Liu, C.; Zhong, Y. Convenience stores geospatial location optimization analytics using deep reinforcement learning. In Proceedings of the 3rd ACM SIGSPATIAL International Workshop on Spatial Big Data and AI for Industrial Applications; ACM: New York, NY, USA, 2024; pp. 16–23. [Google Scholar]
- Pflieger, G.; Rozenblat, C. Introduction. Urban networks and network theory: The city as the connector of multiple networks. Urban Stud. 2010, 47, 2723–2735. [Google Scholar] [CrossRef]
- Smith, H. Supermarket choice and supermarket competition in market equilibrium. Rev. Econ. Stud. 2004, 71, 235–263. [Google Scholar] [CrossRef]
- Noworól, A.; Kopyciński, P.; Hałat, P.; Salamon, J.; Hołuj, A. The 15-minute city—The geographical proximity of services in Krakow. Sustainability 2022, 14, 7103. [Google Scholar] [CrossRef]
- Yue, Y.; Zhuang, Y.; Yeh, A.G.; Xie, J.Y.; Ma, C.L.; Li, Q.Q. Measurements of POI-based mixed use and their relationships with neighbourhood vibrancy. Int. J. Geogr. Inf. Sci. 2017, 31, 658–675. [Google Scholar] [CrossRef]
- Unwin, D.J. GIS, spatial analysis and spatial statistics. Prog. Hum. Geogr. 1996, 20, 540–551. [Google Scholar] [CrossRef]
- Turk, T.; Kitapci, O.; Dortyol, I.T. The usage of Geographical Information Systems (GIS) in the marketing decision making process: A case study for determining supermarket locations. Procedia-Soc. Behav. Sci. 2014, 148, 227–235. [Google Scholar] [CrossRef]
- Oak Ridge National Laboratory. LandScan Global 2023: Silver Edition [Data Set]. 2024. Available online: https://landscan.ornl.gov/ (accessed on 1 February 2025).
- Haklay, M.; Weber, P. OpenStreetMap: User-generated street maps. IEEE Pervasive Comput. 2008, 7, 12–18. [Google Scholar] [CrossRef]
- Weng, Y.; Xu, M.; Chen, X.; Peng, C.; Xiang, H.; Xie, P.; Yin, H. An Efficient Algorithm for Extracting Railway Tracks Based on Spatial-Channel Graph Convolutional Network and Deep Neural Residual Network. ISPRS Int. J. Geo-Inf. 2024, 13, 309. [Google Scholar] [CrossRef]
- Xi’an Municipal Bureau of Statistics. Communiqué on Major Data of the Seventh National Population Census of Xi’an City. 31 May 2021. Available online: https://tjj.xa.gov.cn/ (accessed on 1 February 2025).
- Zhao, P.; Luo, A.; Liu, Y.; Xu, J.; Li, Z.; Zhuang, F.; Sheng, V.S.; Zhou, X. Where to go next: A spatio-temporal gated network for next poi recommendation. IEEE Trans. Knowl. Data Eng. 2020, 34, 2512–2524. [Google Scholar] [CrossRef]
- Xue, J.; Jiang, N.; Liang, S.; Pang, Q.; Yabe, T.; Ukkusuri, S.V.; Ma, J. Quantifying the spatial homogeneity of urban road networks via graph neural networks. Nat. Mach. Intell. 2022, 4, 246–257. [Google Scholar] [CrossRef]
- LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef] [PubMed]
- Zhao, L.; Fan, J.; Zhang, C.; Shen, W.; Zhuang, J. A DRL-based reactive scheduling policy for flexible job shops with random job arrivals. IEEE Trans. Autom. Sci. Eng. 2023, 21, 2912–2923. [Google Scholar] [CrossRef]
- Lin, C.H.; Yumer, E.; Wang, O.; Shechtman, E.; Lucey, S. St-gan: Spatial transformer generative adversarial networks for image compolocation selection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2018; pp. 9455–9464. [Google Scholar]
- Fotheringham, A.S.; Brunsdon, C. Local forms of spatial analysis. Geogr. Anal. 1999, 31, 340–358. [Google Scholar] [CrossRef]
- Glaeser, E.L.; Ponzetto, G.A.M.; Zou, Y. Urban networks: Connecting markets, people, and ideas. Pap. Reg. Sci. 2016, 95, 17–60. [Google Scholar] [CrossRef]
- Zheng, Y.; Lin, Y.; Zhao, L.; Wu, T.; Jin, D.; Li, Y. Spatial planning of urban communities via deep reinforcement learning. Nat. Comput. Sci. 2023, 3, 748–762. [Google Scholar] [CrossRef] [PubMed]
- Yin, H.; Wang, W.; Wang, H.; Chen, L.; Zhou, X. Spatial-aware hierarchical collaborative deep learning for POI recommendation. IEEE Trans. Knowl. Data Eng. 2017, 29, 2537–2551. [Google Scholar] [CrossRef]
- Ellickson, P.B.; Misra, S. Supermarket pricing strategies. Mark. Sci. 2008, 27, 811–828. [Google Scholar] [CrossRef]







| Symbol | Description |
|---|---|
| I | Set of all demand points (grid cells) |
| Binary decision variable: whether to build a facility at location j | |
| Binary decision variable: whether demand point i is covered | |
| Demand intensity of cell i: | |
| Grid distance between demand point i and facility j | |
| R | Coverage radius |
| Reward function at time step t: | |
| Scaling factor for coverage gain (default: 10); aligns reward magnitudes for stable gradient updates | |
| Population density of grid cell i | |
| Consumption level (purchasing power) of grid cell i | |
| Set of candidate facilities that can cover cell i: | |
| , | Weight hyperparameters for demand intensity and competition penalty (default: 0.8, 0.02) |
| Service Pressure at candidate j: |
| POI Name (EN) | POI Type (EN) | Count | POI Name (EN) | POI Type (EN) | Count |
|---|---|---|---|---|---|
| Cake Shop | Shopping Service | 511 | Post Office | Life Service | 472 |
| Bank | Life Service | 436 | Bubble Tea Shop | Shopping Service | 410 |
| Fruit Shop | Shopping Service | 399 | Fast Food Restaurant | Food & Beverage Service | 396 |
| Vegetable Market | Essential Shopping Service | 390 | Restaurant | Food & Beverage Service | 383 |
| Supermarket | Essential Shopping Service | 381 | Courier Point | Life Service | 367 |
| Plaza | Life Service | 364 | Hot Pot Restaurant | Food & Beverage Service | 364 |
| Coffee Shop | Food & Beverage Service | 354 | Convenience Store | Essential Shopping Service | 353 |
| Snack Shop | Shopping Service | 342 | Grain & Oil Store | Essential Shopping Service | 339 |
| Tobacco & Alcohol Store | Shopping Service | 333 | Butcher Shop | Essential Shopping Service | 321 |
| Department Store | Shopping Service | 319 | Retail Store | Shopping Service | 315 |
| Apartment | Residential Area | 315 | Small Grocery Store | Shopping Service | 294 |
| Building | Comprehensive Service | 258 | Residential Compound | Residential Area | 253 |
| Fresh Food Store | Shopping Service | 249 | Food Store | Shopping Service | 242 |
| Community | Residential Area | 207 | Grocery Store | Shopping Service | 182 |
| Garden | Life Service | 175 | Office Building | Comprehensive Service | 157 |
| Villa Area | Residential Area | 85 | Shopping Mall | Shopping Service | 79 |
| Urban Village | Residential Area | 67 | Commercial Building | Comprehensive Service | 55 |
| Configuration Parameter | Value |
|---|---|
| Hardware | GPU: NVIDIA GeForce RTX 4060 (NVIDIA Corporation, Santa Clara, CA, USA); CPU: Intel Core i5-13th Gen, RAM: 16 GB (Intel Corporation, Santa Clara, CA, USA) |
| Software | Python 3.10 (Python Software Foundation, Wilmington, DE, USA); PyTorch 2.0 (Meta Platforms, Menlo Park, CA, USA); Folium (open source) |
| Problem Scale | 427 candidate locations, 487 demand points |
| Coverage radius: km, Number of facilities: |
| Metric | GeoPPO | Vanilla PPO | GA | PSO |
|---|---|---|---|---|
| Demand Coverage (%) | 75.0 | N/A (infeasible) | 72.3 | 69.5 |
| Competitive Intensity Index | 1.21 | N/A | 5.88 | 6.24 |
| Average Response Time | 8.5 min | >45 min | 10.6 min | 15.2 min |
| Training/Convergence | 1000 epochs | No convergence | 1700+ iter. | 2400+ iter. |
| Solution Quality | High | Infeasible | Medium | Medium–Low |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Published by MDPI on behalf of the International Society for Photogrammetry and Remote Sensing. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Hu, Y.; Qin, K.; Wang, S. GeoPPO—A Location-Allocation Method of Superstores Based on Deep Reinforcement Learning—A Case Study of Xi’an. ISPRS Int. J. Geo-Inf. 2026, 15, 114. https://doi.org/10.3390/ijgi15030114
Hu Y, Qin K, Wang S. GeoPPO—A Location-Allocation Method of Superstores Based on Deep Reinforcement Learning—A Case Study of Xi’an. ISPRS International Journal of Geo-Information. 2026; 15(3):114. https://doi.org/10.3390/ijgi15030114
Chicago/Turabian StyleHu, Yuxuan, Kun Qin, and Shaohua Wang. 2026. "GeoPPO—A Location-Allocation Method of Superstores Based on Deep Reinforcement Learning—A Case Study of Xi’an" ISPRS International Journal of Geo-Information 15, no. 3: 114. https://doi.org/10.3390/ijgi15030114
APA StyleHu, Y., Qin, K., & Wang, S. (2026). GeoPPO—A Location-Allocation Method of Superstores Based on Deep Reinforcement Learning—A Case Study of Xi’an. ISPRS International Journal of Geo-Information, 15(3), 114. https://doi.org/10.3390/ijgi15030114

