Next Article in Journal
Identity-Private P2P Energy Trading for Virtual Power Plants
Previous Article in Journal
CoMa: Contextual Massing Generation with Vision-Language Models
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

TerriScan: An Incident-Evaluated, Doctrine-Governed Multi-Agent LLM System for Recalculable Urban Indicator Production in the Global South

by
Yassine Attarassi
* and
Jamal Al Karkouri
*
Laboratory of Territorial Planning, Geo-Environment and Development, Faculty of Humanities and Social Sciences, Ibn Tofail University, Kenitra 14000, Morocco
*
Authors to whom correspondence should be addressed.
Smart Cities 2026, 9(9), 154; https://doi.org/10.3390/smartcities9090154 (registering DOI)
Submission received: 23 July 2026 / Revised: 22 August 2026 / Accepted: 26 August 2026 / Published: 17 September 2026

Highlights

What are the main findings?
  • A written constitution, progressively code-enforced over the four-month development period, governed agent-assisted urban indicator production; no unsupported numeric value detected by the recorded controls remained in the engraved matrix.
  • During the instrumented period, six lots required at least one substantive interception before acceptance (lot population reported stratified by evidentiary status); the dated, hash-referenced taxonomy includes defects introduced by the arbitration agent itself.
What are the implications of the main findings?
  • Agent-assisted urban data production can be evaluated through dated incident records with replayable archived components and explicitly reported operational counts, while clearly separating observed performance from untested generalization.
  • The architecture formalizes “we did not find it” as a checkable object—named producers, a stated administrative tier, and a written audit trail; the submitted corpus did not yet satisfy that rule, and the gap between stating and enforcing it is reported as a principal finding.

Abstract

City-level indicators are difficult to ground across heterogeneous statistical systems in the Global South, where large language model (LLM) agents accelerate multilingual source discovery but risk unsupported values and fabricated execution reports. We present TerriScan, a doctrine-governed multi-agent system built while producing a 142-indicator matrix for ten emerging centralities in eight countries. A versioned charter separates production, model review, deterministic validation and non-delegable human decisions. We evaluate it as a four-month longitudinal design case study with one instrumented four-day period; its lot evidence is stratified, and six lots required substantive interception. During 15–19 July 2026, eight defect classes were registered, each with an identifiable corrective and no recorded intra-class recurrence; exposure denominators were published where countable; and seven earlier qualifying corrections predating the register are reported. The arbiter origin recurred across classes; the exhaustive hash-resolution control remained planned. No unsupported value detected by recorded controls remained in the engraved matrix. The system stopped when work required unrecorded human decisions, but our audit found its absence rule unenforced—181 of 210 absence-state cells named no consulted source. We claim no minimal or universal architecture, but show how incident records, deterministic controls and decision boundaries make urban data production auditable and capable of principled refusal.

Graphical Abstract

1. Introduction

Anyone who has tried to fill an indicator table for a secondary African city knows where the real work is. It is not the scoring model. It is the afternoon spent hunting a municipal budget that three institutions should publish and that no documented search of ours could retrieve; the census table whose perimeter matches nothing on the map; the value you wrote down in March and can no longer defend in November. Global gridded products help—up to the exact point where institutions, not satellites, hold the data [1,2].
Large language models (LLMs) appear to offer an escape. Agentic systems read statistical yearbooks in six languages, query OpenStreetMap, run zonal statistics on population rasters, and draft documentation at machine speed [3,4]. Multi-agent LLM systems have already reached urban practice—supporting decision-making, routing citizen queries and answering planners in natural language [5,6]. But that work sits downstream, where data already exist. The collection stage—where a value is found, sourced and defended—remains comparatively under-governed [6]. In that stage, an LLM may return a fluent city-level value without adequate evidential support [7]. We recorded a more operational version of the problem. On 15 July 2026, an agent session reported a completed processing lot, twenty-eight raster tiles and commit 44e30c3 as proof. The commit did not exist. The report was coherent and confident, but unsupported by repository state. Once production was instrumented, six lots triggered at least one substantive defect interception before acceptance (population stratified by evidentiary status); each occurrence was typed, dated and archived (Section 4). The risk is therefore not only an incorrect answer, but an apparently complete production record that cannot be reconciled with the underlying artifacts.
This paper is about the governance architecture built in response, rather than a new collection interface. Its first commitment is a written constitution—a versioned charter loaded into every agent context and enforced by code rather than by prompt phrasing alone. The idea transposes constitution-guided model alignment [8] to a different object: the production of scientific data. Its second commitment is a set of machine-readable decision classes that bound what agents may execute without human authorization. Its third is the evaluation regime itself: an incident-based record in which the four principal operational counts (work-order attempts; lots executed end-to-end; graver-agent spawns; reviewer verdicts) are exactly recomputable from the archived logs, on the population those logs retain, rather than a synthetic benchmark, together with an explicit account of what such evidence cannot establish (Section 6). Methodologically, this is a longitudinal design case study with an embedded incident-based evaluation. It is not a controlled comparison of alternative agent architectures and does not establish that the TerriScan components are individually necessary or jointly minimal.
The empirical object is a 1420-cell matrix—142 indicators by 10 cities—built for a doctoral comparison of emerging centralities in Morocco, Egypt, Senegal, Ghana, Ivory Coast, Mozambique, Vietnam and Ecuador. The composite scoring layer sits under a pre-defense embargo and is not evaluated here; Section 5.3 states the contract at the interface. The production system uses the internal state label GRAVE (from the French verb graver, to engrave) for a provenance-complete, reviewed and committed cell. In the manuscript, we therefore use “provenance-complete (GRAVE)” at first mention and retain the internal code where it is necessary to describe the implementation. The participles graved and engraved, used interchangeably below, both denote a cell in the GRAVE state. We report two completion measures: the provenance-complete rate and terminal closure, defined as the share of cells in any terminal state—GRAVE, qualified, or an absence state (valid only when its evidentiary conditions are demonstrated; Section 9).

2. Background

2.1. Urban Indicators and the Southern Data Gap

Global products give any city on Earth a population surface and a built-up footprint: the Global Human Settlement Layer (GHSL) [9,10], WorldPop [11]. OpenStreetMap adds networks and amenities, with a completeness that is measurable and very uneven [12,13]. What none of this covers is the institutional layer—budgets, governance, and services—which lives in national sources of wildly different depth. Standardized schemes such as ISO 37120 [14] quietly assume municipal statistical publication that much of our study area does not make retrievable [1,15]. Comparative urbanism has spent two decades arguing that Southern cities need methods which neither import Northern data assumptions nor give up on comparison [2,16,17]. Data-justice scholarship adds a warning we take as a design constraint: absence of published data is not the absence of knowledge, and measuring the former must not slide into a deficit story about Southern institutions [18,19,20]. Our contribution to both agendas is narrow and practical. Every comparative cell carries its evidential status on its face—including absence, an evidential state that is valid only when its tier-specific conditions are demonstrated.

2.2. LLM Agents: From Decision Support to Governed Production

Tool-using agents [3,21] and multi-agent frameworks [4,22,23], surveyed with their simulation uses in [24], demonstrate capability, and the same literature documents the coordination failures and hallucination cascades that come with it [7]. In this journal, multi-agent LLM systems have been shown to route urban queries with high accuracy and to improve response quality substantially over standalone models for city-management decision support [5], and natural-language interfaces are opening urban data to non-specialists. A recent systematic review in this journal, covering 178 studies across the urban analytics pipeline, finds LLM applications fragmented across disciplines and concentrated downstream of data collection [6]. Collection is the stage this paper governs. The gap is visible in this journal from the application side as well: urban digital twins are now being built for data-scarce Global South contexts—including Morocco—with LLM-based enrichment of city models identified as an emerging path [25]; such systems presuppose exactly what TerriScan governs, the accountable production of the city-level values themselves.
Table 1 positions this combination against the adjacent strands. Each row is a published system or system family from the works cited above; columns are the dimensions of comparison adopted for this study; the table is descriptive and does not establish superiority, necessity or exhaustiveness.
The accountability question is being asked, mostly from the systems side. Platforms for multi-granular collection and management of data provenance in LLM workflows are emerging [26]; observability tooling records agent executions for inspection; and governance proposals increasingly route high-impact agent actions to human oversight. The combination examined here is narrower and operational: a versioned normative document enforced across agents; separation between the agent that produces a change, the agent that reviews it and deterministic code that validates it; refusal treated as an observable production outcome; and a longitudinal incident record from scientific work rather than a simulation. We do not claim that no prior system contains any of these components. The contribution is the integrated architecture and its incident-based evaluation in multilingual urban indicator production.
Three recent lines sharpen this positioning. Debate over Mixed-knowledge [28] argues that benchmark incompleteness is typically simulated by randomly removing triples, “failing to capture the irregular and unpredictable nature of real-world knowledge incompleteness”; the corpus studied here inhabits exactly the regime that critique names—absence is structural and unpredictable, and this architecture documents it rather than imputing it. MACA-CD [29] is the closest use of structured multi-agent disagreement to improve a measured outcome: debate is aggregated toward consensus under a judge and evaluated against measured diagnostic outcomes on real-world datasets. The contrast with the present system is operational: where debate systems aggregate disagreement into better answers, this system converts disagreement into a stopped lot—the lot 0010 discordance is the measured instance (Section 4). FinDeepResearch [30] grades deep-research agents against a hierarchical rubric of 15,808 items across 64 companies, eight markets and four languages; it is the sharpest measure of what this study cannot do—no equivalent published ground-truth rubric exists for this panel, where target values are, in a substantial share of cells, not published at the required tier, and this study did not construct one. We state that contrast as a boundary of this study’s evaluation.

2.3. Provenance

Our recalculability contract adapts provenance standards [31] and FAIR-style documentation norms [32,33] to a setting where the chain must survive not just human error but machine confabulation. Sources are pinned by content hash and commit; exhaustive mechanical hash-resolution checking (V11) was not active at submission and remains planned (Section 3.3 and Section 9, Supplementary S1). The validator suite resembles declarative data-quality verification from production data engineering [27], with one shift of emphasis: our upstream producer can invent coherently, so provenance and scope checks matter more than outlier statistics.

3. Materials and Methods

3.1. Cells and States

The unit of production is the cell (city × indicator). Six implementation states are used. GRAVE denotes a provenance-complete, accepted and committed cell. QUALIFIE denotes an accepted cell with a declared limitation—for instance, a documented completeness floor not met. EN_ATTENTE_G2 denotes a cell waiting on an official denominator. PUBLIC_NOT_FOUND was defined by doctrine as a proven publication absence, requiring at least three named sources and a written audit trail; the post-submission audit found this definition unenforced on the reported corpus (Section 9), and enforcement became an executable gate on 8 August 2026. NFE means not found in this execution. A_FAIRE means open. A GRAVE cell must carry raw values, an explicit formula, a hash-pinned source and a perimeter identifier such that a third party can recompute the value from the referenced artifacts alone (Supplementary S8). PUBLIC_NOT_FOUND asserts, when its conditions are demonstrated, a property of the publication regime at a stated administrative tier, not a property of local knowledge; for the submitted corpus, those conditions were largely undemonstrated (Section 9). Bouake’s budget may exist within Bouake’s administration even when it is not publicly retrievable. Figure 1 shows the decision tree that assigns these states.

3.2. The Panel

Ten cities are chosen as a typological panel rather than a sample. One pivot: the Kenitra–Mehdiya–Sidi Taibi corridor system in Morocco, with Temara as an intra-national contrast. State-decreed new towns at three levels of state capacity: Diamniadio in Senegal, New Cairo and Sheikh Zayed in Egypt. Secondary and industrial centralities: Bouake in Ivory Coast, and Thu Dau Mot (Binh Duong) in Vietnam. Coastal intermediates: Manta in Ecuador, Pemba in Mozambique, and Sekondi-Takoradi in Ghana. An earlier fourteen-city benchmark informed the doctoral framework; the production panel was cut to ten for calculability, and what it does not represent is discussed there. For this paper, the panel’s job is to stress one production system across eight statistical regimes and six working languages. Notation was used throughout. Cities carry four-letter codes in the figures: KMST (the Kenitra–Mehdiya–Sidi Taibi corridor system), TEMA (Temara), NCAI (New Cairo), SHZA (Sheikh Zayed), DIAM (Diamniadio), BOUA (Bouake), SEKO (Sekondi-Takoradi), PEMB (Pemba), BINH (Thu Dau Mot, Binh Duong) and MANT (Manta). The 142 indicators are distributed over three axes, identified by prefix: RA for regional governance (74 indicators, 740 cells), CUE for economic centrality (37 indicators, 370 cells) and ULV for local urbanism (31 indicators, 310 cells). An indicator code combines its axis prefix with a group identifier and a rank within that group, as in RA-GT1; the twenty-two groups are the row units of the coverage figures in Section 5. Deterministic validators are numbered V1 to V10, with V11 planned; their checks are tabulated in Supplementary S1. Production lots carry a four-digit sequential identifier, as in lot 0010. Supplementary sections are cited as S1–S13, listed at the end of this article. Figure 2 locates the panel; the frozen envelopes themselves are not mapped there and are versioned with their checksums in Appendix A. Perimeters are functional urban envelopes—the objects these centralities are, which administrative boundaries mis-capture—frozen in a versioned registry with checksum-pinned geometries (V1.1 at submission). Appendix A covers the spatial choices, including what we do and do not claim about the modifiable areal unit problem.

3.3. Roles

The platform is an Electron/Node.js desktop application (Electron 28.3.3, Node 20.x). The pipeline is plain Node around a git repository—some 1850 lines in total for runner, validators, reviewer and display data. Agents are headless invocations of a commercial LLM CLI, and no model identifier is attested in the execution records of the instrumented window: harness-version logging began on 20 July at 02:43 UTC, per-role model identification exists from 24 July onward, per-call identification with the configured-versus-served distinction from 27 July, and since 9 August, every commit carries a mandatory agent-identity trailer. We state this instrumentation trajectory rather than an unattested identifier.
Exactly one process may write the canonical data tree at a time. Locks enforce this at repository and worktree level, and any git rebase or merge state triggers refusal to operate—a rule whose origin story is in Section 4.
The runner is a loop that consumes work orders. Each order carries a YAML header: identifier, campaign, wave, decision class, the exhaustive list of files it may touch, the exact number of output lines and commits expected, and a timeout (Supplementary S2 shows a real one). The runner spawns a graver agent headlessly with a deliberately short tool list—edit, write, git add and commit, node, and python. Never push. Never permission-bypass. The diff then goes to review and validation, and the runner writes a heartbeat to a state file at every phase. Snapshots of that state, the per-run logs and the night event journal were committed during the instrumented window (three event-journal commits on 19 July; a hashed audit bundle containing the ten run logs and the terminal runner state, committed 20 July at 04:56 local), and the four principal operational counts of Table 2 can be recomputed exactly from the bundled logs by the single search command stated in the denominator record committed on 20 July at 14:32 local. The state file itself was not committed at every phase; the two commits it received during the window both capture an idle state.
For class-A and class-B lots, a second agent reviews the work: freshly spawned, read-only tools, no shared conversational state with the graver. It receives the order, the git diff, the canonical indicator definition, and the cited decision. It answers CONCORDANT or DISCORDANT, ten lines of reasons maximum. Discordance isolates the lot; the queue moves on.
Then deterministic code evaluates the diff. Ten validators, with no LLM in the validation path, check counter coherence between ledger and exports; an invariant sentinel cell; physical bounds; diff scope against the declared file list; output schema; source-key existence in the canonical registry; raster recalculability against a frozen zonal table; export regeneration; commit count; and tree hygiene (Supplementary S1). Each ships with failing fixtures (suites L1–L6 and D1–D7), run before every campaign. Two controls were not active during the reported period and are therefore not counted among the ten: V11, mechanical resolution of every cited external content hash, and a multi-cell sentinel set across the three axes. They are scheduled prerequisites for the final matrix. The design rule is simple: the model that produced a diff never validates it.
Hunters live at the edge. Browser-based and scheduled agents, restricted to the public web, produce typed fiches—VALUE, LEAD, NOT-FOUND, FLAG—under a strict schema (Supplementary S3). A hunter proposes a value, its source, and its tier. It decides none of them. Mapping a hunted theme to one of the 142 codes happens upstream, through generated candidate tables and recorded arbitration. Two daily scheduled hunters run statelessly from self-contained prompts with deterministic rotations—one across national statistical offices, and one across the 22 indicator groups; language routing covers French, English, Arabic, Portuguese, Spanish and Vietnamese, and anything outside those is a declared blind spot. A doctrine refined in production: a technical access barrier—client-side rendering, an invalid certificate, and a paywall—is never a proven absence; it is a lead, with the exact document located.
Above the pipeline sits an arbiter: an LLM instance that holds the charter, the decision register and the live state, renders logged class-B arbitrations, adjudicates candidate tables, and drafts work orders. It cannot write data. It cannot launch the runner. The arbiter was itself the source of four defects during production, and the pipeline corrected it each time (Section 4). Figure 3 shows how these roles are arranged, from the 142 indicator-level agents up through group and axis levels to the arbiter. And above everything is one human. Class C—panel, definitions, perimeters, freezes, and publication—belongs to the human decider alone, and the GO that starts any live run is a physical action outside any chat.
Figure 4 shows the pipeline in time, with the measured durations of a single lot. Figure 5 shows the same pipeline as a state machine: order parsed, class routing, graver, reviewer, validators, then push, isolate, or freeze, with the fail-closed transitions marked.

3.4. Threat Model

Three adversarial threats extend beyond ordinary error. First, prompt injection through hunted content: a fetched page can contain instructions addressed to the reading agent. Containment is structural. Hunters cannot write the canonical data tree; their only output channel is a typed fiche that is mechanically validated, while tier and mapping decisions remain upstream. A malicious or misleading page may still produce a malformed or false fiche, which candidate-table arbitration and source checks are intended to reject. Second, source poisoning: provenance guarantees recalculability from a pinned source, not the truth of that source, so source criticism remains human work. Third, correlated judges: the graver and reviewer used the same model family during the reported period, making correlated errors plausible. Every substantive interception observed to date came from one of six named mechanisms—deterministic validators, the executor’s mandatory sense check, the runner pre-flight, the post hoc hash audit, the refusal-to-operate checks with reflog forensics, and the independent reviewer (Supplementary S11). Only one is attributable to the reviewer: lot 0010, an under-scoped repair caught by the independent reviewer before acceptance, corrected and re-run to PASS—recorded in the committed defect record but dropped from the submitted table, and restored here (Table 2). The reviewer is retained as a control against well-formed but semantically incorrect diffs, but its marginal operational value was unproven at submission, and a cross-model replay on archived lots was planned. During revision, that replay and the pre-registered ablation arms were executed; their outcomes are reported in Section 9, and independent human auditing remains future work. Repository locks and hooks address environment integrity; the remote-plus-push regime duplicates history but does not make it tamper-evident, so signed commits or periodic external anchoring of HEAD remain planned extensions. No adversarial experiment was run against these safeguards, before or during revision: the containment described is structural intent, the record contains naturally occurring adversarial events (Section 4 and Section 9), and bounded adversarial testing remains future work.

3.5. The Display Layer

We render the matrix as a “galaxy” (Figure 6): 142 star-agents in 22 clusters on three spiral arms, radiance proportional to measured closure, and event particles emitted only by real pipeline events. The reason is not aesthetics. The first dashboard we built lied to us—hard-coded figures from an earlier project era, displayed as if live. We quarantined it behind a non-dismissible banner and rebuilt the display around one rule: gauges are cross-checked against the ledger verifier at generation time, and on divergence, the interface refuses to render. It shows the error instead of stale light. The same verification standard applies to pixels as to data.

3.6. The Doctrine

Prompt-level instructions are vulnerable to truncation, context loss and local override. TerriScan therefore stores its normative rules in a versioned charter loaded by every agent and amendable only through a ratified commit. Each rule is linked to its design rationale or originating incident so that later optimization does not detach the rule from the failure it addresses. Table 3 lists the ten doctrines and their origins. The genesis record is partial and reported as such: for five of the ten doctrines the repository preserves a dated triggering incident committed before submission, and three of those five carry it in the doctrine’s own text; of the five design-attributed doctrines, two are stated design positions and three carry a plausible but not established genealogy. We report that split rather than claiming a genealogy for all ten.
Decision classes complete the constitution. A class-A order applies a dated, recorded human decision, which it must cite; it runs alone, under graver, reviewer and validators. Class B requires the arbiter’s committed decision file to exist first. Class C never runs: the runner moves it, unexecuted, to a pending-decision queue that batches everything the human must decide—context, options, a recommendation, and the lots waiting behind each choice. One more doctrine arrived late, forced by an audit (Section 4, incident 6): a national-level value may sit in a cell only where the indicator’s registry explicitly accepts the national tier as the measure. Otherwise, it is a proxy, to be requalified—usually into a proven absence or a documented lead at city tier. National indices that map to no city-tier indicator go to a context annex. Not into cells.

3.7. Evaluation Design and Units of Analysis

The evaluation distinguishes five units. A work order is a formally specified production request submitted to the runner. A lot is the execution unit associated with one work order, including its review and validation artifacts. A defect occurrence is one observed deviation from the applicable doctrine, registry definition, execution contract or integrity rule. A defect class groups occurrences sharing an underlying failure mechanism. An incident is a dated operational episode that may contain one or more defect occurrences and affect one or more lots. A substantive interception is a detected defect that required correction, requalification, isolation or a human decision before acceptance; machine sleep, transport interruption and stream timeout are classified separately as operational failures.
The main quantitative record is the instrumented-period lot population, reported stratified by evidentiary status in Table 2; the submitted version stated a single denominator of 16 evaluable lots, which the revision audit could not decompose into an archived named list. Two purely operational attempts were excluded from the substantive interception denominator. The outcome categories reported in Table 2 describe both attempt-level events and final lot dispositions and therefore should not be summed as mutually exclusive branches. The broader four-month record is used as design-rationale evidence because controls evolved over that interval. That evolution is itself dated: pre-repository work through early May 2026 is evidenced by dated documents only and is marked as such; version-controlled production began on 7 May (first proxy values, 12 June; the indicator registry, 13 June), evidenced by commit hashes; systematic hunting began on 15 July, the day of the founding incident, after which controls hardened into the instrumented window of 17–20 July. A dated chronology of these milestones, distinguishing machine-evidenced from document-evidenced entries, accompanies the audit materials. Claims about recurrence are limited to the monitored re-exposure available after each corrective action. Classification and attribution of the defect record were performed by the system’s designer and operator, without independent adjudication; the archived artifacts are intended to make independent re-adjudication possible.
This design does not estimate a general model error rate, false-negative rate or causal contribution of individual controls. Independent human auditing of accepted cells, retrospective component ablation and cross-model review are required to estimate those quantities. At submission, these were identified as follow-up evaluations; during revision, the retrospective ablation and the cross-model review were pre-registered and are reported in Section 9 and Supplementary S12, while independent human auditing of accepted cells remains future work.

4. Results: The Incident Record

4.1. Denominators

Table 2 reports operational totals reconstructed from the record of the instrumented period (17–20 July 2026). Because the submitted denominator of 16 could not be reconstructed as an archived named population, the revision reports evidentiary strata and does not compute a substantive interception ratio. Figures derive from state files, validation reports and the git log reproduced in the audit bundle. Earlier months used progressively weaker controls; incidents from that period are therefore presented as design-rationale evidence, not as results from a stable controlled protocol.
The interception record—six lots, reported within stratified evidentiary counts rather than as a ratio—should not be interpreted as an estimate of general LLM reliability. It is a descriptive property of this production record, under this task mix and control architecture. Its value lies in the explicit evidentiary strata and in the availability of dated artifacts connecting each intercepted defect to its origin, consequence and corrective mechanism. Figure 7 combines the eleven recorded defect occurrences in the registered taxonomy with the separately reported execution strata of the instrumented window; the two populations are shown together descriptively and are not used to form a common ratio. The same occurrences are re-presented stage-wise, with per-defect interception profiles, in Supplementary S11-c.
Two limitations are immediate. The instrumented window is short, and the reviewer intercepted one substantive defect (lot 0010); whether the deterministic validators would also have caught it cannot be determined from the record, because the chain stopped at review and the validators never ran on the refused diff. Its marginal contribution is therefore one interception of undetermined uniqueness; the pre-registered retrospective ablation reported in Section 9 addresses exactly this question.

4.2. What Happened, in Order

Figure 8 tabulates the eight incidents in date order, each with its interceptor and the doctrine it created.
The fabricated report (15 July). Described in the introduction. What it changed: “hash or nothing” became doctrine number five, and every escalation report now embeds the actual git log, so a narrative can no longer drift from repository state without the drift being visible on the same page.
The arbiter’s errors (16 and 17 July). Twice, work orders drafted by the arbitration agent misstated an indicator. Once, a growth rate was ordered as a population count. Once, a registry quantification contradicted its own unit and the error propagated into the order. Both times, the executing agent did something we had specified but not yet seen: it refused. It re-derived the canonical definition, found the mismatch, quoted the registry back, and stopped. We had written the sense-verification step to catch sloppy sources. It caught the arbiter instead.
The rogue rebases (17–18 July). Four times across two days, the runner’s worktree ended up midway through an interactive rebase that no agent had launched—conflicts on export files, HEAD detached, and our own fixes reverted. The fourth one cost us most of a Saturday. The reflog finally named the culprit, and it was nobody’s agent: the code editor’s background synchronization was quietly running pull-with-rebase against the worktree. The fixes were mechanical (refusal-to-operate on any rebase state, a worktree lock, a pre-rebase hook), but the lesson we wrote into the charter is broader: any automation not declared in the architecture is an illegitimate writer, even when it ships with the tooling.
The silent floor bypass (18 July). A perimeter-correction lot re-graved accessibility cells and quietly skipped a documented amenity-count floor. Four cells that should not have claimed GRAVE did. Nothing in the lot’s own output looked wrong; what caught it was doctrine 9—two totals disagreed by ten, the rule demands a named diff for any unexplained gauge movement, and the diff exposed the four. They were requalified the same day.
The proxy erratum and the registry contamination (19–20 July). A three-cell gauge drop, investigated under the same rule, turned out to be the opposite of a defect: hunter fiches had correctly requalified national statistical proxies—water access, electricity access, and school enrolment, all for one city—into city-tier absence states under the doctrine then in force—labels whose evidentiary conditions the post-submission audit later found largely undemonstrated (Section 9). A quality gain, recorded by the counters as a loss. The audit that legitimated it produced the tier doctrine, then a panel-wide audit of every remaining national-context cell—which found the indicator registry itself partially contaminated by copy-pasted tier fields from an earlier project era. Fixing the normative layer took a class-C decision, and the doctrine was then applied at scale: fifty national-proxy cells were requalified out of GRAVE in documented sub-lots, one national index was graved on the single indicator whose registry admits the national tier, and a crime-rate indicator lost four national-proxy cells to documented city-tier leads. The pipeline audited its own constitution; the constitution needed it; and every gauge movement was explained.
The push-path flaw (19 July). The graver could push a lot that validation later rejected. Closed the same day: only the orchestrator pushes, and only after concordance and validation.
The export freeze (19 July, night). A requalification lot updated thirteen cells correctly but the work order—drafted by the arbiter—omitted regeneration of one export that the counter validator reads. V1 saw a divergence, froze everything, and refused the push. The remote stayed clean. The local tree was held for autopsy, the missing regeneration was added, and the lot completed; the export requirement was then generalized to every status-touching order. This is the one integrity freeze in Table 2, and it is the pipeline working exactly as written: the data were right, the paperwork was not, and nothing left the building until both agreed.
Across the development record, eight defect classes were observed—all recorded from the register’s institution on 15 July—and associated with eight corrective responses. No recurrence was recorded during the subsequent monitored exposure available for each corrected class. Because exposure duration and opportunities for re-exposure differ across classes, this is an observed production record rather than an estimate of recurrence probability.

4.3. The Autonomous Overnight Demonstration

On the night of 18–19 July, the orchestrator received a standing order to continue until no further authorized actions remained. Four agent processes ran—two gravers and two reviewers—under the orchestrator. Two lots completed the full path from production to review, validation and push, and their target-log entries were verified the following morning. Two class-C orders were refused and routed to the human decision queue with briefs. One lot remained empty because the repository contained insufficient evidence to substantiate the targeted absences. The run terminated with the recorded status “legitimate acts exhausted.” With only two completed lots, this is a demonstration rather than an evaluation. Its evidential value is limited to the observed refusal and stopping behavior under the recorded mandate.

5. Results: The Production Record

5.1. Three Production Regimes

At submission, 221 of 1420 cells (15.6%) were provenance-complete (GRAVE); by axis, the corresponding shares were 14% for regional governance (RA, 102 of 740 cells), 16% for economic centrality (CUE, 59 of 370) and 19% for local urbanism (ULV, 60 of 310); the three axis denominators sum to the 1420 cells of the matrix and the three numerators to the 221 graved cells, and both are readable in the cockpit status bar of Figure 6. Terminal closure was 33.0% under the labels then assigned. These figures were lower than three days earlier because the tier audit described in Section 4.2 requalified fifty national-proxy cells out of the accepted count, each through a documented sub-lot. Figure 9 records this decrease as a traceable quality correction rather than treating completion as necessarily monotonic. Figure 10 maps terminal closure by indicator group and city.
Three production regimes showed different marginal costs. First, panel telemetry: after perimeters were frozen and a zonal table was pinned by commit, raster-derived indicators could be processed across all ten cities at low additional computational cost. The pivot city’s frozen geometry was later found to be self-intersecting; the repair ran as a reviewed lot under an identical-output constraint. The zonal output remained unchanged, the area delta was 0.000 km2, the new checksum propagated to thirteen files and the previous checksum remained in the perimeter register. Second, institutional hunting, where source verification dominated the effort. Measured OSM road-network completeness ranged from 12.0 to 62.9 km of mapped road per square kilometer of built-up surface, against a floor of 5.28 calibrated on the pivot city. This single-city calibration is declared; the pre-announced sensitivity treatment was executed during revision: across the 40 cells carrying the ratio, over eight of the ten entities, none passes the floor by less than a factor of two—the closest exceeds it by 2.27—so the flagged set is empty under the pre-declared criterion, which bounds the sensitivity exposure of this record without establishing that single-city calibration is generally sound (Appendix A). One walk-graph extraction for a 119 km2 perimeter produced a 12,508-node network in 77 s. Third, doctrine-blocked cells were batched for human decision. Two-pass review approximately doubled model invocations per lot, and validators added execution time. The present record shows that the reviewer produced one substantive interception (lot 0010), while the other recorded interceptions are distributed across the five further mechanisms named in Supplementary S11; whether the reviewer’s interception was uniquely necessary is unresolved in the historical record and is addressed by the pre-registered retrospective ablation. Per-call token accounting (model identifier, input/output tokens, cache and cost) exists for the hunting engine from 27 July 2026 onward (post-submission instrumentation); no token accounting exists for the runner lots of the instrumented window, and we report that gap rather than an estimate.

5.2. The Calculability Frontier, Read as Publication Regimes

As doctrine, the absence rule required justification at a stated tier; the post-submission audit reported in Section 9 found it unenforced on the reported corpus (181 of 210 absence-state cells named no consulted source). What follows is therefore a provisional description of the submitted state, whose absence labels were subsequently invalidated—not a validated map of publication regimes. Two interpretations must remain separate. When its evidentiary conditions are satisfied, the record is intended to measure publication conditions—whether an institution publishes a value at the required tier, in a retrievable source and within the searched languages. It does not measure what local administrations know [18,19]. At 33.0% terminal closure under the labels then assigned—a non-homogeneous series, restated as 18.3% after the reclassification of 28 July and 17.6% under the doctrine of 11 August—the emerging pattern is provisional: municipal-budget and participatory-governance indicators were not publicly grounded at city tier across much of the African subsample; climate-vulnerability indices were available only at national tier for all ten cities, of which eight entered the graved count under the registry that admits that tier and were routed to a context annex except where the registry explicitly admitted that tier; and amenity-dependent accessibility fell below documented completeness floors in several smaller cities. Figure 11 aggregates terminal closure by indicator group and country. If these patterns persist as closure increases, the record may support a comparative analysis of publication regimes and help identify where field collection or institutional partnership is required.
Figure 11. Submitted-state terminal closure under the labels then assigned, aggregated by indicator group and country. Because the post-submission audit invalidated most absence labels, the figure is retained as a provisional historical description, not a validated map of publication regimes; it concerns retrievable institutional publication, not what local administrations know.
Figure 11. Submitted-state terminal closure under the labels then assigned, aggregated by indicator group and country. Because the post-submission audit invalidated most absence labels, the figure is retained as a provisional historical description, not a validated map of publication regimes; it concerns retrievable institutional publication, not what local administrations know.
Smartcities 09 00154 g011

5.3. The Interface with Scoring

The scoring layer is embargoed until the thesis defense, but the contract it inherits belongs here, because the production states create it. National-context cells (tier doctrine) carry a macro-context scoring role: they enter no local composite score and count toward no local-data floor. Qualified cells below a completeness floor are excluded from the indicator’s city score, with limitation declared and nothing imputed. Proven absences feed the frontier analysis and never the scores. Whether one definition measures the same construct in Kenitra and in Pemba—cross-city measurement invariance—is not established by this pipeline; it is treated at the scoring stage following composite-indicator practice [34]. What the pipeline contributes is blunter: it makes the evidential heterogeneity explicit enough to be modeled instead of hidden. Deployment cost is reported as a record, not an estimate: human arbitration time and doctrine maintenance effort were not measured, and the dated stream of correctives, doctrine amendments and gates in the audit bundle constitutes the maintenance record. We do not present the architecture as economically efficient—a system with a non-delegable human gate is expensive by construction, and this study establishes neither economic efficiency nor end-to-end scalability.

6. Discussion

The incident record establishes that each observed defect class was associated with a nameable interception mechanism and that no recurrence was recorded within the subsequent monitored exposure available for each corrected class. It does not establish that the mechanisms are individually necessary, jointly minimal or causally responsible for all avoided errors. A retrospective ablation—replaying archived lots while disabling selected components—was executed during revision under a pre-registered protocol; Section 9 reports which controls caught which archived defects. Independent human auditing is also required to estimate false negatives among accepted cells. Until that audit is complete, the claim is limited to design rationale supported by incident and replay evidence.
The human is the system boundary in more than one sense. Class C concentrates judgment risk in one person rather than eliminating it. The architecture partially exposes that risk because the arbiter is subject to the same sense checks and because every class-C decision is stored as a dated, versioned file available for audit. The registry-contamination episode illustrates both sides: the pipeline detected a problem in its normative layer, but that layer still required human correction. The system therefore improves traceability of judgment without converting judgment into an automated fact.
Several architectural components may transfer to other domains where agentic speed is useful and unsupported assertions are costly: versioned rules, explicit decision classes, role separation, deterministic validation and a human authorization boundary. Transferability is not demonstrated here, however. Urban indicator production presents specific requirements—misalignment between functional and administrative perimeters, heterogeneous publication tiers, multilingual sources and frequent reliance on national proxies for local phenomena—that shape the present implementation. Applications in public health, heritage or environmental compliance would require domain-specific registries, evidence rules and validation tests.
The manuscript was drafted with LLM assistance under author review, as declared below. That use is reported for transparency but is not itself evidence for the effectiveness of the TerriScan architecture. The scientific claims rely on the archived production artifacts, incident records and validation outputs described in the Methods and Data Availability Statement.

7. Conclusions

TerriScan emerged from a practical problem: producing a 1420-cell comparative urban indicator matrix without allowing agentic speed to obscure evidential uncertainty. The resulting architecture combines a versioned doctrine, machine-readable decision classes, role separation, deterministic validators, provenance requirements and a human boundary for non-delegable decisions. The evidence reported here is an incident-evaluated longitudinal case study, not a controlled validation of a universally sufficient architecture. During the instrumented period, substantive defects were intercepted in six lots before acceptance; the lot population is reported stratified by evidentiary status in Table 2. Across the broader development record, eight observed defect classes were associated with identifiable corrective responses, with no recurrence recorded within the subsequent monitored exposure available for each corrected class. The system stopped on contradictory orders and froze an internally inconsistent export state; its absence rule, by contrast, was stated but not enforced on the reported corpus (Section 9) before the corpus reached the remote repository. These observations support a limited but consequential claim: agent-assisted scientific data production can be made more auditable when authority, evidence, validation and refusal are represented explicitly in the production architecture. They do not establish that TerriScan is minimal, optimal or transferable without adaptation. Retrospective component ablation and cross-model review were executed during revision under pre-registered protocols and are reported in Section 9; independent human auditing of accepted cells remains necessary to estimate the marginal contribution of each control. The audit artifacts are intended to make that audit possible.

8. Access Outcomes and Endpoint Persistence (Bounded Measurements)

Two bounded measurements are now available; their limits are as informative as their values. First, endpoint persistence. On 29 July 2026, the assessment covered the source endpoints associated with 222 engraved, source-named cells (60 on the local-urbanism axis, 103 on the regional-governance axis, 59 on the economic-centrality axis, using the axis names of Section 5.1). The recorded outcome counts were 193 resolved, 19 access-control responses, 7 endpoint non-resolutions and 3 indeterminate connection resets. Because multiple cells may cite the same endpoint, these counts are not interpreted here as 222 distinct URLs or as independent observations. Of the seven non-resolutions, one was recovered through the archival fallback; for the remaining six, the fallback was attempted with a retry, and no archived copy was found. Non-resolution at an endpoint is not disappearance of the underlying document—the resource was not searched for at alternative paths—but the intervals are short: the affected cells had been engraved 16 to 23 days earlier. All six unrecovered non-resolutions and all three indeterminate outcomes fall on the governance axis, which carries 103 of the 222 cells. Because distinct cells may cite the same endpoint, these outcomes are not independent across cells, and the asymmetry may reflect a small number of endpoints rather than a property of the axis; we report it as a described asymmetry, not a tested hypothesis. Second, HTTP outcomes of automated acquisition attempts. Among 277 logged download attempts carrying timestamp and duration, 16 (5.8%) initially returned one of the pre-specified non-success status codes: ten 403s across eight distinct domains, two 404s, two 429s, one 401, and one 302 followed to the end of its redirect budget without converging. Thirteen of the sixteen are access-control responses to an automated client and describe the requester’s status rather than the publisher’s practice; two established endpoint-level non-resolution (0.7% of attempts). Every one of the sixteen triggered the designed human-handover mechanism, and none was subsequently resolved in the logged record. These events postdate submission and characterize the re-hunting campaign, not the corpus behind the submitted matrix. The wider proposition—that in weakly institutionalized publication settings, a dated, content-addressed snapshot preserves verifiable evidence of what the system retrieved at a stated time—is offered as methodological, not measured; the Supplementary Materials specifies the per-URL instrumentation (status, timestamp, duration, redirect chain, archival outcome, joined to cell identifiers) that would make it measurable.

9. Developments Since Submission (Dated Record)

Nothing in this section is used to defend the submitted corpus; the separation of states is itself part of the method.
Between 26 July and 14 August 2026, the production system and its published record were audited against the repository under the evidentiary rules this paper describes. Four findings correct the submitted manuscript and are integrated above: the replayability sentence in the Methods (the archived bundle supports exact recomputation of the four principal operational counts, and no more); the lot-population recount of Table 2 (the submitted 16 was not decomposable into an archived named list); the restored reviewer interception on lot 0010; and the re-scoping of the incident register to its institution date of 15 July 2026, with seven earlier qualifying corrections now listed and hash-referenced in Supplementary S11-b, and one of them additionally dated.
The principal post-submission finding is the enforcement audit of the absence rule. Of the 210 cells carrying the absence state at submission, 181 named no consulted source; of the remaining 29, 28 cited at least one source but failed another evidentiary condition, and a single cell satisfied all of them. The 209 under-investigated cells were reclassified as non-investigated on 28 July 2026 under a signed named attestation—a verdict that also created the stricter reading now in force, three distinct producers rather than three addresses; the rule became an executable seventeen-condition gate on 8 August. Since that date, an independent adversarial second pass is a production precondition for any absence-state transition; the engraving tool refuses the six legacy absence labels; and the validator re-checks every declared condition rather than trusting it—declared booleans are not evidence. The record includes one real refusal in which the gate named the capability it lacked. The doctrine itself predates submission by five days (committed 18 July, 18:03 local); its enforcement did not. We report the mechanism and the failure to apply it as one result, not two. A second post-submission finding concerns the acceptance contract itself. One acceptance falls outside the eleven reported PASS outcomes: lot 0018 committed eight ND-GAIN cells at 02:02:43Z between its second launch and its second escalation, and that commit entered the accepted tree through the push of the following lot. The revision audit establishes that the commit is an ancestor of the submitted state, that no later commit touched those cells before submission, and that lot 0018 is the only lot of the fourteen with no recorded review event. The mechanism is nameable: the acceptance gate is placed on the push and evaluates a lot’s diff against its predecessor’s commit, so an escalated lot’s commit becomes the baseline of the next validation rather than its object, and its cells are never examined. We report this as a coverage gap of the acceptance contract—identified by our own audit, with its mechanism—and not as a detection failure: the controls did not miss these cells, as they did not run on them. A third post-submission finding corrects the last term of the closure series reported in Section 5.2. Recomputed on 19 August 2026 against the tree that produced it, terminal closure under the doctrine of 11 August is 250 of 1420 cells (17.6%), not the 240 (16.9%) previously obtained. That rule admits three classes and no others—provenance-complete and source-named, qualified, and a third class holding a single cell—and, unlike the submitted-state definition given in Section 1, it does not admit absence labels; the recount moves only the first class, from 222 cells to 232, leaving 17 and 1 unchanged. The ten recovered cells are the complete row of one regional-governance indicator (RA-GT1) across all ten cities. The cause is a traversal defect in the coverage counter, not a change of rule: the counter reads one cell field where that indicator file writes another, and the guard meant to suppress an empty result tests an object that JavaScript evaluates as true when the object is empty, so the cells were computed and never reported. The deterministic suite saw what the counter did not: validator V2, whose sentinel invariant is exactly that cell—RA-GT1 for the pivot corridor system, expected value 50.18 (Supplementary S1)—returns a passing verdict when replayed on the same tree. The value against which the scoring engine is checked for drift was, in the same tree and at the same second, filed by the coverage counter as not yet investigated. No cell value was altered and no clause of the closure rule was changed in the recount; the correction is three lines, and it is confined to the traversal. A stricter reading exists and is recorded rather than adopted: two of the ten cells carry a not-comparable flag at their parent indicator that the serving engine honors and the classifier does not test, and honoring it would return 248 cells (17.5%). We do not apply it here because it would change the closure rule and not merely its traversal. The series is therefore non-homogeneous in two respects at once, and we state both. Each term was computed under a different closure doctrine; the code that computes closure also changed between the three dates—the coverage routine did not exist at the date of the first term, and at the date of the second, the qualified class that the count aggregates had not yet been defined. Applied with the corrected counter, the rule of 11 August returns 250 on the record as it stood on 20 July, on 28 July and on 11 August alike. The three published terms differ by rule and by counting code, not by any change in the record, and the series is reported as a chronology of the record’s own definitions.
Pre-registered evaluation results (14–15 August 2026). The protocol was committed before execution (with the publication rule signed in advance); every result below carries the commit of the replayed system and the commit of the replay configuration, and none is used to qualify the submitted corpus.
Replay fidelity (same family). On the 13 admissible archived replay inputs, the same-family reviewer reproduced the archived accepted-state verdicts in 13/13 cases, against a historical baseline whose model identity is not attested in the July artifacts—the gate therefore also provides a bounded check for temporal drift against a model-unknown baseline; none was detected at the verdict level. Because the admissible population is all-concordant (the one historical discordant verdict survives in the logs, but its input—the refused diff—is unrecoverable and its review file was overwritten, so that verdict is inadmissible for replay), this gate certifies replay validity, not reviewer discrimination.
Cross-family replay (blinded). The same 13 frozen dossiers were re-reviewed by a second model family (GPT-5 family; the configured model identifier was later recovered from session records at the forensic audit of 15 August; the served identifier was not exposed by the client and is recorded as such), one fresh context per lot, with verdicts sealed before disclosure. Outcome: 10/13 concordant, and 3 discordant. Adjudication against each lot’s own order and the doctrine in force at its date—performed by the operator (the system’s designer) and recorded in a dated adjudication record, a self-adjudication stated as such: of the three discordant verdicts, one contained a confirmed defect and two were adjudicated as review-standard differences, with zero lot-level false positives and zero indeterminate verdicts; within the two latter verdicts, two subsidiary factual allegations were refuted, which we report as the measured cost of diversity.
The confirmed defect (lot 0010) is a process finding, not a detection finding. The fact—a registry line asserting a zero grep count for a superseded checksum when the count in the accepted tree is two—was mechanically confirmed, was covered by the lot’s own order, and was seen by all three reviewers: the historical reviewer recorded it verbatim as a non-blocking reviewer reservation (internal label: réserve) requesting correction before push; the same-family replay recorded the same reservation; and the cross-family reviewer classified it blocking. The run log records the push within the PASS event itself and reaches its terminal state 2.9 s later; the push is not separately timestamped. Non-blocking reviewer reservations had no follow-up channel—the reservation was seen, then lost for 27 days, and recovered by cross-family replay. No deterministic validator checks documentary assertions against the tree (the nearest control, V11, was planned at submission): a coverage gap, not a validator failure. Family diversity acted here as a severity re-calibrator and recovered a previously unclosed reviewer reservation, consistent with—but not proving—family-correlated severity norms. The register gains two entries: the missing reservation channel—closed by dated rule on 16 August 2026: non-blocking reviewer reservations now suspend the push until explicitly acknowledged—and the platform-degradation family (a shell-escaping quirk silently emptied eleven frozen diff artifacts during the campaign; detected by the first consuming agent, repaired before any verdict).
The two review-standard differences localize the normative drift. The cross-family reviewer required exact source URLs and stricter tier justifications; the July orders specified neither (one root-domain URL was the order’s verbatim specification). Model substitution shifted the review standard, not the facts: review norms are partly family-borne, and freezing inputs does not freeze standards.
Gates under substitution. Launched autonomously overnight, the prospective demonstration stopped at the human gate under the repository’s own doctrine—twice, once per model family: 19 cells—ten in one presented group, nine in the other, as presented at that run; the production pilot below ran groups of ten and eight—were halted before any network access and before any source, value or absence decision. Model substitution did not bypass the model-independent execution gates. No integrity regression was detected by the specified post-run checks (targeted 15/15; full suite including 282/282 retrieval-chain tests and 9/9 hooks). The human queue stood at six entries the following morning, timestamped—a queue-depth measurement produced by the operator’s absence. The revision audit also produced a census of non-delegable human interventions across the record: 24 traced to dated documents in the paper window, of which 21 are concentrated in 17–21 July 2026, with the instrumentation boundary stated.
Deterministic-validator replay and ablation arms (16 August 2026). Replayed on an era-faithful bench at the epoch commit, nine of the ten deterministic validators reproduce the archived validation outcomes on all 13 admissible lots, status by status—including the sole historical FAIL, reproduced with verbatim detail (lot 0013, V1). The tenth (V8, the export check) diverges on all 13 replays from a single structural cause discovered by the replay itself: the export tool hard-codes the production tree path, so V8 never validated the target tree—masked in July because target and production coincided; registered as the fourth member of the silent-platform-degradation family (after CRLF line-ending parsing, the cmd.exe caret escaping quirk, and prompt truncation not surfaced to the operator interface), with a dated corrective. Ablation arm C (reviewer without validators, fresh context): the reviewer again passed lot 0013—the defect is a global-state property invisible from the diff, outside the reviewer’s mandate, and was caught in all three exercises only by validator V1. Lot 0013 therefore provides the complementary case to lot 0010: the reviewer passed a global-state defect that only validator V1 intercepted, while on lot 0010, the reviewer intercepted the scope defect—the uniqueness of that historical interception remaining undetermined, since the validators never ran on the refused diff. In arm B, on lot 0010’s corrected diff and within the declared limit of the pre-registration, the deterministic suite ran without the reviewer and raised nothing on the residual documentary assertion—no deterministic validator is mandated to check documentary assertions against the tree, which is the coverage gap, mechanically confirmed; the assertion itself has since been corrected by a dated record, and the reservation follow-up channel instituted by dated rule.
Prospective bi-family production pilot (protocol pre-registered 14 August 2026, executed 15–16 August, closed under a closure rule pre-registered before the final run). Of the two pre-registered prospective production groups, one (CUE-MC5) was grounded in full at the variable freeze—none of its ten cells carried a pre-existing territorial tier, and the pre-registered rule excluded rather than created mandates. In the second group (RA-CS, eight cells), the first-family leg completed under the standing gates: two values engraved after human validation with citation-grain provenance and six cells left in documented barrier or lead states (2 + 6 = 8); one of the six stopped at the enforced-absence gate as designed. The substituted family’s legs self-invalidated three times on sealing-topology and gate-formalization defects, each self-declared before any comparison, and the pre-registered closure rule ended the comparative arm: the bi-family production comparison was not achieved within the revision window and is reported as such. The pilot’s yield is process evidence—the freeze intercepted a stale registry assertion and grounded an unmandated group before any network access; strict execution surfaced five specification gaps, registered for the campaign protocol; and no leg, in any pilot run, fabricated execution evidence—every invalidation was self-declared (Supplementary S12).

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/smartcities9090154/s1, S1: The deterministic validator suite (V1–V10, V11 planned); S2: Machine-readable work-order schema + real instance; S3: Hunter fiche schema (typed, mechanically validated); S4: Two-pass review protocol (class A); S5: Real production state (runner_state.json, field names translated, lot 0010, 19/07/2026); S6: Incident registry (from its institution on 15 July 2026—the paper’s Section 4 in table form); S7: The autonomous night (metrics, 18–19/07/2026); S8: Provenance block (grandeur_brute—French field name, “raw quantity”)—Section 3.1 recalculability contract; S9: Code and data availability (statement for submission); S10: Figures (as published); S11: Measured defect-interception taxonomy (registered record; execution strata of the instrumented period, 17–20 July 2026); S11-b: Pre-register corrections (6 June–14 July 2026); S11-c: Stage-wise re-presentation of the defect record; S12: Pre-registered retrospective evaluation (delivered, with results); S13: Use of generative AI in manuscript preparation (MDPI disclosure).

Author Contributions

Conceptualization, Y.A.; methodology, Y.A.; software, Y.A.; investigation, Y.A.; data curation, Y.A.; writing—original draft preparation, Y.A.; visualization, Y.A.; supervision, J.A.K.; validation, J.A.K.; writing—review and editing, J.A.K. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

An anonymized audit bundle is deposited on Zenodo (concept DOI: https://doi.org/10.5281/zenodo.21463056, resolving to the latest version; revision bundle: version v3, DOI: https://doi.org/10.5281/zenodo.21966264, extended on 24 August 2026 by version v4, DOI: https://doi.org/10.5281/zenodo.22085786): the charter, validator suite, runner and reviewer code, three complete lots (orders, diffs, reviewer verdicts, validation reports and provenance blocks), incident artifacts from Section 4, the measured defect taxonomy and operational logs. The governance-layer code and audit artifacts required to evaluate the claims are available to editors and reviewers during peer review; the indicator matrix and composite scoring layer remain under doctoral embargo until January 2027, with a documented release plan. Per-cell provenance blocks will be released with the matrix. Confidential repository access can be provided through the corresponding author where required.

Acknowledgments

This work was conducted under the doctoral supervision of Jamal Al Karkouri, Ibn Tofail University. The authors thank the Laboratory of Territorial Planning, Geo-Environment and Development for the doctoral framework in which this system was built. During the preparation of this work the authors used a large language model assistant (Anthropic Claude; configured models claude-fable-5 and claude-opus-5, and Claude Code 2.1.178; full disclosure in Supplementary S13) for drafting and editing text, under the review procedures described in this paper. The authors reviewed and edited all content and take full responsibility for the content of the published article. For the revision, the tools, versions and roles—including one tool tested and disqualified—are detailed in Supplementary S13.

Conflicts of Interest

The authors declare no competing financial interests. The first author (Y.A.) practices as a licensed architect in the pivot city’s region; no client relationship involves the studied sites’ data. This position is also declared as a methodological resource: field verification of the pivot system draws on professional knowledge of the local planning regime.

Appendix A. Spatial Methods

Built-up surface and population derive from the GHSL R2023 package (GHS-BUILT-S, GHS-POP, GHS-SMOD for settlement classification) at epochs 2015 and 2020. The 2025 and 2030 epochs are model projections in the source package; we excluded them on provenance grounds. Zonal statistics run on the native Mollweide grid (ESRI:54009). Vector perimeters are projected to the raster CRS, never the reverse; boundary pixels are weighted by intersection fraction; and the full zonal output is frozen as one commit-pinned table that every raster-derived cell must reference (validator V7). We use functional urban envelopes rather than administrative boundaries because corridor systems and new towns are exactly the objects administrative boundaries mis-capture; each envelope is versioned with a checksum, and the one geometry repair during production was accepted only under the identical-output constraint reported in Section 5.1. This does not solve the modifiable areal unit problem. It holds it fixed and declares it: every comparison holds for these envelopes, and the registry records the envelope version against every raster-derived value. The OSM completeness floor (5.28, calibrated on the pivot city) is a declared single-city calibration; cells passing it by less than a factor of two are flagged in the registry for sensitivity treatment at the scoring stage. The treatment was executed during revision: across the 40 ratio-carrying cells (eight of the ten entities), the flagged set is empty—the closest cell exceeds the floor by a factor of 2.27. This bounds the sensitivity exposure of the current record; it does not validate the single-city calibration itself.

References

  1. Acuto, M.; Parnell, S.; Seto, K.C. Building a global urban science. Nat. Sustain. 2018, 1, 2–4. [Google Scholar] [CrossRef] [Scilit]
  2. Watson, V. African urban fantasies: Dreams or nightmares? Environ. Urban. 2014, 26, 215–231. [Google Scholar] [CrossRef] [Scilit]
  3. Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. ReAct: Synergizing reasoning and acting in language models. In Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
  4. Wu, Q.; Bansal, G.; Zhang, J.; Wu, Y.; Li, B.; Zhu, E.; Jiang, L.; Zhang, X.; Zhang, S.; Awadallah, A.H.; et al. AutoGen: Enabling next-gen LLM applications via multi-agent conversation. arXiv 2023, arXiv:2308.08155. [Google Scholar]
  5. Kalyuzhnaya, A.; Mityagin, S.; Lutsenko, E.; Getmanov, A.; Aksenkin, Y.; Fatkhiev, K.; Fedorin, K.; Nikitin, N.O.; Chichkova, N.; Vorona, V.; et al. LLM agents for smart city management: Enhancing decision support through multi-agent AI systems. Smart Cities 2025, 8, 19. [Google Scholar] [CrossRef] [Scilit]
  6. Jiang, F.; Ma, J.; Jin, Y. Unleashing the potential of large language models in urban data analytics: A review of emerging innovations and future research. Smart Cities 2025, 8, 201. [Google Scholar] [CrossRef] [Scilit]
  7. Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.J.; Madotto, A.; Fung, P. Survey of hallucination in natural language generation. ACM Comput. Surv. 2023, 55, 1–38. [Google Scholar] [CrossRef] [Scilit]
  8. Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. Constitutional AI: Harmlessness from AI feedback. arXiv 2022, arXiv:2212.08073. [Google Scholar]
  9. Pesaresi, M.; Freire, S. GHS Settlement Grid; JRC, European Commission: Ispra, Italy, 2016. [Google Scholar]
  10. Schiavina, M.; Melchiorri, M.; Pesaresi, M.; Politis, P.; Carneiro Freire, S.M.; Maffenini, L.; Florio, P.; Ehrlich, D.; Goch, K.; Carioli, A.; et al. GHSL Data Package 2023; JRC, European Commission: Ispra, Italy, 2023. [Google Scholar]
  11. Tatem, A.J. WorldPop, open data for spatial demography. Sci. Data 2017, 4, 170004. [Google Scholar] [CrossRef] [Scilit]
  12. Barrington-Leigh, C.; Millard-Ball, A. The world’s user-generated road map is more than 80% complete. PLoS ONE 2017, 12, e0180698. [Google Scholar] [CrossRef] [Scilit]
  13. Herfort, B.; Lautenbach, S.; Porto de Albuquerque, J.; Anderson, J.; Zipf, A. A spatio-temporal analysis investigating completeness and inequalities of global urban building data in OpenStreetMap. Nat. Commun. 2023, 14, 3985. [Google Scholar] [CrossRef] [Scilit]
  14. ISO 37120:2018; Sustainable Cities and Communities—Indicators for City Services and Quality of Life. ISO: Geneva, Switzerland, 2018.
  15. Kitchin, R.; Lauriault, T.P.; McArdle, G. Knowing and governing cities through urban indicators, city benchmarking and real-time dashboards. Reg. Stud. Reg. Sci. 2015, 2, 6–28. [Google Scholar] [CrossRef] [Scilit]
  16. Robinson, J. Comparative urbanism: New geographies and cultures of theorizing the urban. Int. J. Urban Reg. Res. 2016, 40, 187–199. [Google Scholar] [CrossRef] [Scilit]
  17. Roy, A. The 21st-century metropolis: New geographies of theory. Reg. Stud. 2009, 43, 819–830. [Google Scholar] [CrossRef] [Scilit]
  18. Taylor, L. What is data justice? The case for connecting digital rights and freedoms globally. Big Data Soc. 2017, 4, 2053951717736335. [Google Scholar] [CrossRef] [Scilit]
  19. D’Ignazio, C.; Klein, L.F. Data Feminism; MIT Press: Cambridge, MA, USA, 2020. [Google Scholar]
  20. Milan, S.; Treré, E. Big data from the South(s): Beyond data universalism. Telev. New Media 2019, 20, 319–335. [Google Scholar] [CrossRef] [Scilit]
  21. Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Zettlemoyer, L.; Cancedda, N.; Scialom, T. Toolformer: Language models can teach themselves to use tools. In Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 10–16 December 2023. [Google Scholar]
  22. Hong, S.; Zhuge, M.; Chen, J.; Zheng, X.; Cheng, Y.; Zhang, C.; Wang, J.; Wang, Z.; Yau, S.K.S.; Lin, Z.; et al. MetaGPT: Meta programming for a multi-agent collaborative framework. In Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024. [Google Scholar]
  23. Park, J.S.; O’Brien, J.C.; Cai, C.J.; Morris, M.R.; Liang, P.; Bernstein, M.S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST), San Francisco, CA, USA, 29 October–1 November 2023. [Google Scholar]
  24. Gao, C.; Lan, X.; Li, N.; Yuan, Y.; Ding, J.; Zhou, Z.; Xu, F.; Li, Y. Large language models empowered agent-based modeling and simulation: A survey and perspectives. Humanit. Soc. Sci. Commun. 2024, 11, 1259. [Google Scholar] [CrossRef] [Scilit]
  25. Badreddine, O.; Radoine, H.; Hajji, R. TwinCity: An urban digital twin framework for data-scarce environments—A case study of Benguerir, Morocco. Smart Cities 2026, 9, 23. [Google Scholar] [CrossRef] [Scilit]
  26. Gregori, L.; Lazzaro, P.L.; Lazzaro, M.; Missier, P.; Torlone, R. An LLM-guided platform for multi-granular collection and management of data provenance. J. Big Data 2025, 12, 187. [Google Scholar] [CrossRef] [Scilit]
  27. Schelter, S.; Lange, D.; Schmidt, P.; Celikel, M.; Biessmann, F.; Grafberger, A. Automating large-scale data quality verification. Proc. VLDB Endow. 2018, 11, 1781–1794. [Google Scholar] [CrossRef] [Scilit]
  28. Liu, J.; Shao, P.; Qin, W.; Liu, F.; Yang, Y.; Hong, R. Debate over Mixed-knowledge: A Robust Multi-Agent Reasoning Framework for Incomplete Knowledge Graph Question Answering. arXiv 2025, arXiv:2511.12208. [Google Scholar]
  29. Shao, P.; Chen, L.; Liu, F.; Yang, Y.; Yang, X.; Wang, M. Multi-Agent Debate based Concept Augmentation for Enhanced Cognitive Diagnosis. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1 (KDD ’26), New York, NY, USA, 9–13 August 2026; pp. 1287–1296. [Google Scholar]
  30. Zhu, F.; Ng, X.Y.; Liu, Z.; Liu, C.; Zeng, X.; Wang, C.; Tan, T.; Yao, X.; Shao, P.; Xu, M.; et al. FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis. arXiv 2025, arXiv:2510.13936. [Google Scholar]
  31. Moreau, L.; Missier, P. PROV-DM: The PROV Data Model; W3C Recommendation: Wakefield, MA, USA, 2013. [Google Scholar]
  32. Wilkinson, M.D.; Dumontier, M.; Aalbersberg, I.J.; Appleton, G.; Axton, M.; Baak, A.; Blomberg, N.; Boiten, J.-W.; da Silva Santos, L.B.; Bourne, P.E.; et al. The FAIR guiding principles for scientific data management and stewardship. Sci. Data 2016, 3, 160018. [Google Scholar] [CrossRef] [Scilit]
  33. Gebru, T.; Morgenstern, J.; Vecchione, B.; Vaughan, J.W.; Wallach, H.; Daumé, H., III; Crawford, K. Datasheets for datasets. Commun. ACM 2021, 64, 86–92. [Google Scholar] [CrossRef] [Scilit]
  34. OECD/JRC. Handbook on Constructing Composite Indicators: Methodology and User Guide; OECD Publishing: Paris, France, 2008. [Google Scholar]
Figure 1. How a cell earns its epistemic state. A technical access barrier is never a proven absence: it yields a lead with the exact document located. PUBLIC_NOT_FOUND requires at least three named sources and a written audit trail at a stated tier; GRAVE requires a complete provenance block. The figure states the rule in force for the submitted corpus; the stricter, executable post-submission gate is described in Section 9.
Figure 1. How a cell earns its epistemic state. A technical access barrier is never a proven absence: it yields a lead with the exact document located. PUBLIC_NOT_FOUND requires at least three named sources and a written audit trail at a stated tier; GRAVE requires a complete provenance block. The figure states the rule in force for the submitted corpus; the stricter, executable post-submission gate is described in Section 9.
Smartcities 09 00154 g001
Figure 2. The study panel: ten emerging centralities in eight countries. Panel countries shaded; the pivot is the Kenitra–Mehdiya–Sidi Taibi corridor system (Morocco), complemented by an intra-national contrast (Temara), state-decreed new towns at three levels of state capacity (New Cairo, Sheikh Zayed, Diamniadio), secondary and industrial centralities (Bouake, Thu Dau Mot), and coastal intermediates (Sekondi-Takoradi, Pemba, Manta). City codes on the map, in the order of the list above: KMST, TEMA, NCAI, SHZA, DIAM, BOUA, SEKO, PEMB, BINH, MANT (Section 3.2). Two pairs lie less than one plotted pixel apart and carry a single symbol with a compound label: the KMST corridor system and Temara (0.47°), New Cairo and Sheikh Zayed (0.53°). Basemap: Natural Earth (1:110 m); Plate Carrée projection. Country polygons follow the basemap and carry no jurisdictional endorsement.
Figure 2. The study panel: ten emerging centralities in eight countries. Panel countries shaded; the pivot is the Kenitra–Mehdiya–Sidi Taibi corridor system (Morocco), complemented by an intra-national contrast (Temara), state-decreed new towns at three levels of state capacity (New Cairo, Sheikh Zayed, Diamniadio), secondary and industrial centralities (Bouake, Thu Dau Mot), and coastal intermediates (Sekondi-Takoradi, Pemba, Manta). City codes on the map, in the order of the list above: KMST, TEMA, NCAI, SHZA, DIAM, BOUA, SEKO, PEMB, BINH, MANT (Section 3.2). Two pairs lie less than one plotted pixel apart and carry a single symbol with a compound label: the KMST corridor system and Temara (0.47°), New Cairo and Sheikh Zayed (0.53°). Basemap: Natural Earth (1:110 m); Plate Carrée projection. Country polygons follow the basemap and carry no jurisdictional endorsement.
Smartcities 09 00154 g002
Figure 3. The agent hierarchy mirrors the matrix. Escalation carries states, conformity reports and hunt requests upward; it is never a write path. Hunters hold web access but no write access; any agent holds a logged right of objection directly to the human decider.
Figure 3. The agent hierarchy mirrors the matrix. Escalation carries states, conformity reports and hunt requests upward; it is never a write path. Hunters hold web access but no write access; any agent holds a logged right of objection directly to the human decider.
Smartcities 09 00154 g003
Figure 4. One lot through the pipeline, in time. Segment order follows the execution sequence; segment length is indicative and not a metrological reading. Timings measured on lot 0010, a frozen-geometry repair accepted only under the constraint of an identical zonal output; 8 min 31 s to validation PASS.
Figure 4. One lot through the pipeline, in time. Segment order follows the execution sequence; segment length is indicative and not a metrological reading. Timings measured on lot 0010, a frozen-geometry repair accepted only under the constraint of an identical zonal output; 8 min 31 s to validation PASS.
Smartcities 09 00154 g004
Figure 5. The production pipeline as a state machine. Pre-flight checks refuse to operate on a foreign lock or a rebase state; class-C orders are never executed but batched for the human decider; fail-closed transitions (red) isolate a lot on reviewer discordance and freeze the tree on any validator failure, with the remote protected.
Figure 5. The production pipeline as a state machine. Pre-flight checks refuse to operate on a foreign lock or a rebase state; class-C orders are never executed but batched for the human decider; fail-closed transitions (red) isolate a lot on reviewer discordance and freeze the tree on any validator failure, with the remote protected.
Smartcities 09 00154 g005
Figure 6. The cockpit on 20 July 2026, display refreshed at 23:42: the galaxy display (radiance proportional to measured closure; gauges cross-checked against the ledger at render time) above the agent-conversation dock, showing lot 0021—the Manta homicide-rate cell—passing validation and being pushed (conversation-dock timestamps 21:39). Screenshot of the production application; interface in French (see the note on language preceding Supplementary S1).
Figure 6. The cockpit on 20 July 2026, display refreshed at 23:42: the galaxy display (radiance proportional to measured closure; gauges cross-checked against the ledger at render time) above the agent-conversation dock, showing lot 0021—the Manta homicide-rate cell—passing validation and being pushed (conversation-dock timestamps 21:39). Screenshot of the production application; interface in French (see the note on language preceding Supplementary S1).
Smartcities 09 00154 g006
Figure 7. Taxonomy of recorded intercepted defects by origin, shown alongside instrumented-period execution strata: eleven defect occurrences across four origin categories in the registered record—including four introduced by the arbitration agent, with one of them restored during revision from the committed defect record (lot 0010)—traced to the mechanisms that intercepted them. The figure reports detected occurrences and does not estimate undetected defects.
Figure 7. Taxonomy of recorded intercepted defects by origin, shown alongside instrumented-period execution strata: eleven defect occurrences across four origin categories in the registered record—including four introduced by the arbitration agent, with one of them restored during revision from the committed defect record (lot 0010)—traced to the mechanisms that intercepted them. The figure reports detected occurrences and does not estimate undetected defects.
Smartcities 09 00154 g007
Figure 8. The recorded incident history: eight defect classes over the development period—all recorded from the incident register’s institution on 15 July 2026—the mechanism associated with each interception and the corrective doctrine that followed. No recurrence was recorded; exposure denominators are published in Supplementary S11 for the classes whose post-corrective re-exposure is countable and reported as a bound for the others. Artifacts are dated and hash-referenced in Supplementary S6 and S11.
Figure 8. The recorded incident history: eight defect classes over the development period—all recorded from the incident register’s institution on 15 July 2026—the mechanism associated with each interception and the corrective doctrine that followed. No recurrence was recorded; exposure denominators are published in Supplementary S11 for the classes whose post-corrective re-exposure is countable and reported as a bound for the others. Artifacts are dated and hash-referenced in Supplementary S6 and S11.
Smartcities 09 00154 g008
Figure 9. The graved count, 17–20 July 2026: a matrix that can shrink for documented reasons. The tier audit removed fifty national proxies; the single indicator whose registry admits the national tier (the ND-GAIN climate-vulnerability index) gained eight cells in the graved count; a crime indicator moved to city tier lost four proxies to documented leads—one of which a demand-driven hunt converted into a sourced city-tier value within hours.
Figure 9. The graved count, 17–20 July 2026: a matrix that can shrink for documented reasons. The tier audit removed fifty national proxies; the single indicator whose registry admits the national tier (the ND-GAIN climate-vulnerability index) gained eight cells in the graved count; a crime indicator moved to city tier lost four proxies to documented leads—one of which a demand-driven hunt converted into a sourced city-tier value within hours.
Smartcities 09 00154 g009
Figure 10. Submitted-state terminal closure by indicator group and city on 21 July 2026—graved 221 of 1420 cells; terminal closure 33.0% under the labels then assigned. Twenty-two group rows (RA/CUE/ULV prefixes) by ten cities; horizontal rules separate the three axes. Closure here aggregates graved, qualified and the then-assigned absence labels that the post-submission audit later invalidated (Section 9); the figure is retained as a provisional historical description. Figure 11 aggregates the same data by country.
Figure 10. Submitted-state terminal closure by indicator group and city on 21 July 2026—graved 221 of 1420 cells; terminal closure 33.0% under the labels then assigned. Twenty-two group rows (RA/CUE/ULV prefixes) by ten cities; horizontal rules separate the three axes. Closure here aggregates graved, qualified and the then-assigned absence labels that the post-submission audit later invalidated (Section 9); the figure is retained as a provisional historical description. Figure 11 aggregates the same data by country.
Smartcities 09 00154 g010
Table 1. Positioning against adjacent systems (✓ documented in the cited work; — not a design goal there).
Table 1. Positioning against adjacent systems (✓ documented in the cited work; — not a design goal there).
System/FamilyCode-Enforced ConstitutionIndependent Same-Diff ReviewerDeterministic ValidatorsHuman Decision ClassesRefusal as Success MetricIncident-Based Longitudinal EvaluationCity-Level Data Production
Tool-using single agents [3,21]
Multi-agent frameworks [4,22,23,24]partial (role separation)
Urban decision-support agents [5]query answering
Urban digital twins, data-scarce [25]model enrichment
LLM provenance platforms [26]partial
Declarative data-quality checks [27]
TerriScan (this work)
Table 2. Instrumented period totals (17–20 July 2026). Rows report attempt-level events and final lot dispositions on overlapping populations and are not summable as mutually exclusive branches.
Table 2. Instrumented period totals (17–20 July 2026). Rows report attempt-level events and final lot dispositions on overlapping populations and are not summable as mutually exclusive branches.
QuantityValue
Lot population (evidentiary strata)stratified population—see the caption note and the interception row; the submitted single figure of 16 is withdrawn (not decomposable into an archived named list); 2 purely operational attempts counted apart
Work orders attempted (incl. re-fires and aborts)18
Executed end-to-end (PASS)11
Graver-agent spawns (incl. re-fires)16—a distinct quantity from the withdrawn lot-population figure; the two class-C orders never spawn a graver, so 18 = 16 + 2. Both counts are read from the retained night logs and therefore exclude lots attested only by an execution commit or by artifacts never committed; the strata are disjoint and are not summed into a single denominator (four lots ran on 18 July with execution commits but unretained logs; two ran late on 20 July with pipeline artifacts never committed)
Class-C orders correctly refused and routed to the human decision queue2 (executed later under recorded decisions)
Orders deliberately left empty (no evidentiary documents; invention refused)1
Lots with ≥1 substantive defect intercepted before acceptancesix, drawn from two evidentiary strata—three attested by committed logs and artifacts (0009, 0010, 0013), which belong to the fourteen lots with retained logs, and three attested by the committed denominator record with execution commits but unretained logs (0000–0002), which lie outside those fourteen; no substantive ratio computed; the population is reported stratified by evidentiary status; purely operational failures (machine sleep, stream timeout) counted apart
Defect origins across the full record (dated, hashed taxonomy in Supplementary S11)11 detected defect occurrences across four origin categories—ten in the submitted taxonomy plus the lot 0010 under-scoped order, restored during revision from the committed defect record of 20 July; eight broader defect classes in the development record (Supplementary S11)
Reviewer verdicts14 verdicts: 12 concordant (one rendered 152 s before a deterministic validator halted the same lot); 1 substantive discordant (lot 0010—under-scoped repair caught by the independent reviewer before acceptance, corrected, re-run to PASS); 1 parser false error re-read as concordant
Integrity freezes1; remote repository protected (incident 8)
Autonomous overnight demonstration2 lots completed end-to-end; ≈6 min active machine time; 0 detected integrity violations
Longest supervised lot8 min 31 s (geometry repair; 0 cells changed; checksum propagated to 13 files); wall-clock including the refused first attempt: 18 min 20 s
Table 3. The doctrines. Origin distinguishes doctrines with a preserved dated triggering incident from design-attributed doctrines; the evidentiary split of the design attributions is stated in the text.
Table 3. The doctrines. Origin distinguishes doctrines with a preserved dated triggering incident from design-attributed doctrines; the evidentiary split of the design attributions is stated in the text.
#DoctrineRuleOrigin
1RecalculabilityNo provenance-complete (GRAVE) state without a complete provenance blockdesign
2Single writerOne process per tree, lock-enforcedincident 4
3Fail-closed“I don’t know” → isolate the cell, continue; “it is broken” → freeze for autopsydesign
4Sense verificationThe executor re-derives the canonical definition and stops on divergence—including divergence from the arbiter’s own orderincidents 2–3
5Hash or nothingAny claimed commit is verified against the log, including in the system’s own reportsincident 1
6Nobody guessesMappings via candidate tables only; tiers proposed, never decided, by huntersdesign
7Threshold ≠ boundOut-of-band values pass with documented justification; in-band values without sources do notdesign
8Proven absence is a deliverable≥3 named sources plus an audit trail, at a stated tier—and a technical barrier is never an absencedesign
9Gauges tell the truthNo unverified counter is displayed; any unexplained gauge movement requires a named diffincidents 5–6
10Repository guardsNo force-push (hook), push after every accepted lot, rebase refused by defaultincident 4
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Attarassi, Y.; Al Karkouri, J. TerriScan: An Incident-Evaluated, Doctrine-Governed Multi-Agent LLM System for Recalculable Urban Indicator Production in the Global South. Smart Cities 2026, 9, 154. https://doi.org/10.3390/smartcities9090154

AMA Style

Attarassi Y, Al Karkouri J. TerriScan: An Incident-Evaluated, Doctrine-Governed Multi-Agent LLM System for Recalculable Urban Indicator Production in the Global South. Smart Cities. 2026; 9(9):154. https://doi.org/10.3390/smartcities9090154

Chicago/Turabian Style

Attarassi, Yassine, and Jamal Al Karkouri. 2026. "TerriScan: An Incident-Evaluated, Doctrine-Governed Multi-Agent LLM System for Recalculable Urban Indicator Production in the Global South" Smart Cities 9, no. 9: 154. https://doi.org/10.3390/smartcities9090154

APA Style

Attarassi, Y., & Al Karkouri, J. (2026). TerriScan: An Incident-Evaluated, Doctrine-Governed Multi-Agent LLM System for Recalculable Urban Indicator Production in the Global South. Smart Cities, 9(9), 154. https://doi.org/10.3390/smartcities9090154

Article Metrics

Back to TopTop