TerriScan: An Incident-Evaluated, Doctrine-Governed Multi-Agent LLM System for Recalculable Urban Indicator Production in the Global South
Highlights
- A written constitution, progressively code-enforced over the four-month development period, governed agent-assisted urban indicator production; no unsupported numeric value detected by the recorded controls remained in the engraved matrix.
- During the instrumented period, six lots required at least one substantive interception before acceptance (lot population reported stratified by evidentiary status); the dated, hash-referenced taxonomy includes defects introduced by the arbitration agent itself.
- Agent-assisted urban data production can be evaluated through dated incident records with replayable archived components and explicitly reported operational counts, while clearly separating observed performance from untested generalization.
- The architecture formalizes “we did not find it” as a checkable object—named producers, a stated administrative tier, and a written audit trail; the submitted corpus did not yet satisfy that rule, and the gap between stating and enforcing it is reported as a principal finding.
Abstract
1. Introduction
2. Background
2.1. Urban Indicators and the Southern Data Gap
2.2. LLM Agents: From Decision Support to Governed Production
2.3. Provenance
3. Materials and Methods
3.1. Cells and States
3.2. The Panel
3.3. Roles
3.4. Threat Model
3.5. The Display Layer
3.6. The Doctrine
3.7. Evaluation Design and Units of Analysis
4. Results: The Incident Record
4.1. Denominators
4.2. What Happened, in Order
4.3. The Autonomous Overnight Demonstration
5. Results: The Production Record
5.1. Three Production Regimes
5.2. The Calculability Frontier, Read as Publication Regimes

5.3. The Interface with Scoring
6. Discussion
7. Conclusions
8. Access Outcomes and Endpoint Persistence (Bounded Measurements)
9. Developments Since Submission (Dated Record)
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A. Spatial Methods
References
- Acuto, M.; Parnell, S.; Seto, K.C. Building a global urban science. Nat. Sustain. 2018, 1, 2–4. [Google Scholar] [CrossRef] [Scilit]
- Watson, V. African urban fantasies: Dreams or nightmares? Environ. Urban. 2014, 26, 215–231. [Google Scholar] [CrossRef] [Scilit]
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. ReAct: Synergizing reasoning and acting in language models. In Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Wu, Q.; Bansal, G.; Zhang, J.; Wu, Y.; Li, B.; Zhu, E.; Jiang, L.; Zhang, X.; Zhang, S.; Awadallah, A.H.; et al. AutoGen: Enabling next-gen LLM applications via multi-agent conversation. arXiv 2023, arXiv:2308.08155. [Google Scholar]
- Kalyuzhnaya, A.; Mityagin, S.; Lutsenko, E.; Getmanov, A.; Aksenkin, Y.; Fatkhiev, K.; Fedorin, K.; Nikitin, N.O.; Chichkova, N.; Vorona, V.; et al. LLM agents for smart city management: Enhancing decision support through multi-agent AI systems. Smart Cities 2025, 8, 19. [Google Scholar] [CrossRef] [Scilit]
- Jiang, F.; Ma, J.; Jin, Y. Unleashing the potential of large language models in urban data analytics: A review of emerging innovations and future research. Smart Cities 2025, 8, 201. [Google Scholar] [CrossRef] [Scilit]
- Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.J.; Madotto, A.; Fung, P. Survey of hallucination in natural language generation. ACM Comput. Surv. 2023, 55, 1–38. [Google Scholar] [CrossRef] [Scilit]
- Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. Constitutional AI: Harmlessness from AI feedback. arXiv 2022, arXiv:2212.08073. [Google Scholar]
- Pesaresi, M.; Freire, S. GHS Settlement Grid; JRC, European Commission: Ispra, Italy, 2016. [Google Scholar]
- Schiavina, M.; Melchiorri, M.; Pesaresi, M.; Politis, P.; Carneiro Freire, S.M.; Maffenini, L.; Florio, P.; Ehrlich, D.; Goch, K.; Carioli, A.; et al. GHSL Data Package 2023; JRC, European Commission: Ispra, Italy, 2023. [Google Scholar]
- Tatem, A.J. WorldPop, open data for spatial demography. Sci. Data 2017, 4, 170004. [Google Scholar] [CrossRef] [Scilit]
- Barrington-Leigh, C.; Millard-Ball, A. The world’s user-generated road map is more than 80% complete. PLoS ONE 2017, 12, e0180698. [Google Scholar] [CrossRef] [Scilit]
- Herfort, B.; Lautenbach, S.; Porto de Albuquerque, J.; Anderson, J.; Zipf, A. A spatio-temporal analysis investigating completeness and inequalities of global urban building data in OpenStreetMap. Nat. Commun. 2023, 14, 3985. [Google Scholar] [CrossRef] [Scilit]
- ISO 37120:2018; Sustainable Cities and Communities—Indicators for City Services and Quality of Life. ISO: Geneva, Switzerland, 2018.
- Kitchin, R.; Lauriault, T.P.; McArdle, G. Knowing and governing cities through urban indicators, city benchmarking and real-time dashboards. Reg. Stud. Reg. Sci. 2015, 2, 6–28. [Google Scholar] [CrossRef] [Scilit]
- Robinson, J. Comparative urbanism: New geographies and cultures of theorizing the urban. Int. J. Urban Reg. Res. 2016, 40, 187–199. [Google Scholar] [CrossRef] [Scilit]
- Roy, A. The 21st-century metropolis: New geographies of theory. Reg. Stud. 2009, 43, 819–830. [Google Scholar] [CrossRef] [Scilit]
- Taylor, L. What is data justice? The case for connecting digital rights and freedoms globally. Big Data Soc. 2017, 4, 2053951717736335. [Google Scholar] [CrossRef] [Scilit]
- D’Ignazio, C.; Klein, L.F. Data Feminism; MIT Press: Cambridge, MA, USA, 2020. [Google Scholar]
- Milan, S.; Treré, E. Big data from the South(s): Beyond data universalism. Telev. New Media 2019, 20, 319–335. [Google Scholar] [CrossRef] [Scilit]
- Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Zettlemoyer, L.; Cancedda, N.; Scialom, T. Toolformer: Language models can teach themselves to use tools. In Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 10–16 December 2023. [Google Scholar]
- Hong, S.; Zhuge, M.; Chen, J.; Zheng, X.; Cheng, Y.; Zhang, C.; Wang, J.; Wang, Z.; Yau, S.K.S.; Lin, Z.; et al. MetaGPT: Meta programming for a multi-agent collaborative framework. In Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Park, J.S.; O’Brien, J.C.; Cai, C.J.; Morris, M.R.; Liang, P.; Bernstein, M.S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST), San Francisco, CA, USA, 29 October–1 November 2023. [Google Scholar]
- Gao, C.; Lan, X.; Li, N.; Yuan, Y.; Ding, J.; Zhou, Z.; Xu, F.; Li, Y. Large language models empowered agent-based modeling and simulation: A survey and perspectives. Humanit. Soc. Sci. Commun. 2024, 11, 1259. [Google Scholar] [CrossRef] [Scilit]
- Badreddine, O.; Radoine, H.; Hajji, R. TwinCity: An urban digital twin framework for data-scarce environments—A case study of Benguerir, Morocco. Smart Cities 2026, 9, 23. [Google Scholar] [CrossRef] [Scilit]
- Gregori, L.; Lazzaro, P.L.; Lazzaro, M.; Missier, P.; Torlone, R. An LLM-guided platform for multi-granular collection and management of data provenance. J. Big Data 2025, 12, 187. [Google Scholar] [CrossRef] [Scilit]
- Schelter, S.; Lange, D.; Schmidt, P.; Celikel, M.; Biessmann, F.; Grafberger, A. Automating large-scale data quality verification. Proc. VLDB Endow. 2018, 11, 1781–1794. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Shao, P.; Qin, W.; Liu, F.; Yang, Y.; Hong, R. Debate over Mixed-knowledge: A Robust Multi-Agent Reasoning Framework for Incomplete Knowledge Graph Question Answering. arXiv 2025, arXiv:2511.12208. [Google Scholar]
- Shao, P.; Chen, L.; Liu, F.; Yang, Y.; Yang, X.; Wang, M. Multi-Agent Debate based Concept Augmentation for Enhanced Cognitive Diagnosis. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1 (KDD ’26), New York, NY, USA, 9–13 August 2026; pp. 1287–1296. [Google Scholar]
- Zhu, F.; Ng, X.Y.; Liu, Z.; Liu, C.; Zeng, X.; Wang, C.; Tan, T.; Yao, X.; Shao, P.; Xu, M.; et al. FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis. arXiv 2025, arXiv:2510.13936. [Google Scholar]
- Moreau, L.; Missier, P. PROV-DM: The PROV Data Model; W3C Recommendation: Wakefield, MA, USA, 2013. [Google Scholar]
- Wilkinson, M.D.; Dumontier, M.; Aalbersberg, I.J.; Appleton, G.; Axton, M.; Baak, A.; Blomberg, N.; Boiten, J.-W.; da Silva Santos, L.B.; Bourne, P.E.; et al. The FAIR guiding principles for scientific data management and stewardship. Sci. Data 2016, 3, 160018. [Google Scholar] [CrossRef] [Scilit]
- Gebru, T.; Morgenstern, J.; Vecchione, B.; Vaughan, J.W.; Wallach, H.; Daumé, H., III; Crawford, K. Datasheets for datasets. Commun. ACM 2021, 64, 86–92. [Google Scholar] [CrossRef] [Scilit]
- OECD/JRC. Handbook on Constructing Composite Indicators: Methodology and User Guide; OECD Publishing: Paris, France, 2008. [Google Scholar]










| System/Family | Code-Enforced Constitution | Independent Same-Diff Reviewer | Deterministic Validators | Human Decision Classes | Refusal as Success Metric | Incident-Based Longitudinal Evaluation | City-Level Data Production |
|---|---|---|---|---|---|---|---|
| Tool-using single agents [3,21] | — | — | — | — | — | — | — |
| Multi-agent frameworks [4,22,23,24] | — | partial (role separation) | — | — | — | — | — |
| Urban decision-support agents [5] | — | — | — | — | — | — | query answering |
| Urban digital twins, data-scarce [25] | — | — | — | — | — | — | model enrichment |
| LLM provenance platforms [26] | — | — | partial | — | — | — | — |
| Declarative data-quality checks [27] | — | — | ✓ | — | — | — | — |
| TerriScan (this work) | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Quantity | Value |
|---|---|
| Lot population (evidentiary strata) | stratified population—see the caption note and the interception row; the submitted single figure of 16 is withdrawn (not decomposable into an archived named list); 2 purely operational attempts counted apart |
| Work orders attempted (incl. re-fires and aborts) | 18 |
| Executed end-to-end (PASS) | 11 |
| Graver-agent spawns (incl. re-fires) | 16—a distinct quantity from the withdrawn lot-population figure; the two class-C orders never spawn a graver, so 18 = 16 + 2. Both counts are read from the retained night logs and therefore exclude lots attested only by an execution commit or by artifacts never committed; the strata are disjoint and are not summed into a single denominator (four lots ran on 18 July with execution commits but unretained logs; two ran late on 20 July with pipeline artifacts never committed) |
| Class-C orders correctly refused and routed to the human decision queue | 2 (executed later under recorded decisions) |
| Orders deliberately left empty (no evidentiary documents; invention refused) | 1 |
| Lots with ≥1 substantive defect intercepted before acceptance | six, drawn from two evidentiary strata—three attested by committed logs and artifacts (0009, 0010, 0013), which belong to the fourteen lots with retained logs, and three attested by the committed denominator record with execution commits but unretained logs (0000–0002), which lie outside those fourteen; no substantive ratio computed; the population is reported stratified by evidentiary status; purely operational failures (machine sleep, stream timeout) counted apart |
| Defect origins across the full record (dated, hashed taxonomy in Supplementary S11) | 11 detected defect occurrences across four origin categories—ten in the submitted taxonomy plus the lot 0010 under-scoped order, restored during revision from the committed defect record of 20 July; eight broader defect classes in the development record (Supplementary S11) |
| Reviewer verdicts | 14 verdicts: 12 concordant (one rendered 152 s before a deterministic validator halted the same lot); 1 substantive discordant (lot 0010—under-scoped repair caught by the independent reviewer before acceptance, corrected, re-run to PASS); 1 parser false error re-read as concordant |
| Integrity freezes | 1; remote repository protected (incident 8) |
| Autonomous overnight demonstration | 2 lots completed end-to-end; ≈6 min active machine time; 0 detected integrity violations |
| Longest supervised lot | 8 min 31 s (geometry repair; 0 cells changed; checksum propagated to 13 files); wall-clock including the refused first attempt: 18 min 20 s |
| # | Doctrine | Rule | Origin |
|---|---|---|---|
| 1 | Recalculability | No provenance-complete (GRAVE) state without a complete provenance block | design |
| 2 | Single writer | One process per tree, lock-enforced | incident 4 |
| 3 | Fail-closed | “I don’t know” → isolate the cell, continue; “it is broken” → freeze for autopsy | design |
| 4 | Sense verification | The executor re-derives the canonical definition and stops on divergence—including divergence from the arbiter’s own order | incidents 2–3 |
| 5 | Hash or nothing | Any claimed commit is verified against the log, including in the system’s own reports | incident 1 |
| 6 | Nobody guesses | Mappings via candidate tables only; tiers proposed, never decided, by hunters | design |
| 7 | Threshold ≠ bound | Out-of-band values pass with documented justification; in-band values without sources do not | design |
| 8 | Proven absence is a deliverable | ≥3 named sources plus an audit trail, at a stated tier—and a technical barrier is never an absence | design |
| 9 | Gauges tell the truth | No unverified counter is displayed; any unexplained gauge movement requires a named diff | incidents 5–6 |
| 10 | Repository guards | No force-push (hook), push after every accepted lot, rebase refused by default | incident 4 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Attarassi, Y.; Al Karkouri, J. TerriScan: An Incident-Evaluated, Doctrine-Governed Multi-Agent LLM System for Recalculable Urban Indicator Production in the Global South. Smart Cities 2026, 9, 154. https://doi.org/10.3390/smartcities9090154
Attarassi Y, Al Karkouri J. TerriScan: An Incident-Evaluated, Doctrine-Governed Multi-Agent LLM System for Recalculable Urban Indicator Production in the Global South. Smart Cities. 2026; 9(9):154. https://doi.org/10.3390/smartcities9090154
Chicago/Turabian StyleAttarassi, Yassine, and Jamal Al Karkouri. 2026. "TerriScan: An Incident-Evaluated, Doctrine-Governed Multi-Agent LLM System for Recalculable Urban Indicator Production in the Global South" Smart Cities 9, no. 9: 154. https://doi.org/10.3390/smartcities9090154
APA StyleAttarassi, Y., & Al Karkouri, J. (2026). TerriScan: An Incident-Evaluated, Doctrine-Governed Multi-Agent LLM System for Recalculable Urban Indicator Production in the Global South. Smart Cities, 9(9), 154. https://doi.org/10.3390/smartcities9090154
