Abstract
Census-style address records contain heterogeneous structures that must be decomposed into fine-grained fields before they can support linkage, geocoding, and administrative processing. This paper describes the design and implementation of a three-stage Planner–Manager–Worker pipeline using a locally deployed open-weight language model. The prototype includes address-specific extraction, same-model semantic consistency review, formatting, record-preserving chunking, and an exploratory bounded retry pathway. Although the prototype also implements workers for names, phone numbers, email addresses, Social Security numbers, and dates of birth, the quantitative evaluation is limited to address-component parsing. We conduct a controlled pilot using 700 author-generated synthetic records distributed across seven address categories. Each record was first created as a structured object and then rendered into text; the original structured object served as the evaluation reference. Across three repeated executions on the same fixed benchmark, the complete hierarchical configuration obtained a mean micro-averaged address-component exact-match score of 95.7%, compared with 80.9% for the evaluated monolithic LLM reference; the deterministic rule-based reference obtained 52.3%. The fixed prompts were not matched, and realized inference calls, dynamic token allocation, original batching, and reported-run segmentation could not be verified from retained artifacts. No component ablations were conducted. Consequently, the numerical difference cannot be attributed uniquely to planning, task coordination, specialization, semantic review, feedback, chunking, or multi-agent organization. The reference labels were not independently annotated, the benchmark does not estimate performance on naturally occurring records, and the retry pathway was invoked only three times. The contribution is therefore the specification and controlled pilot characterization of an on-premise address-parsing pipeline, rather than evidence of architectural causality, production readiness, broad PII-extraction accuracy, or real-record robustness.
1. Introduction
Every census begins with records about people and households: names, residential addresses, contact information, dates of birth, and demographic or administrative details collected during population enumeration. At the national scale, these ordinary pieces of information become one of the largest and most sensitive datasets maintained by government agencies. The United States counted 331.4 million people in the 2020 Census [1], illustrating the scale at which such records are collected, protected, and processed. Before census data can support population mapping, record linkage, geocoding, identity verification, duplicate detection, or policy planning, raw responses must be converted into consistent structured fields [2,3,4].
This conversion is difficult because census records do not follow one simple format. Addresses may appear as conventional street addresses, apartment or unit addresses, university department addresses, rural routes, highway addresses, military APO/FPO/DPO addresses, or attention-line records that include organizations and routing instructions. A human reader can usually interpret these variations from context, but an automated system must assign each token to the correct component, such as street number, street name, unit, city, state, ZIP code, organization, or attention line. Errors at this stage can propagate into downstream systems and affect matching, geocoding, household detection, and statistical processing.
Traditional parsing methods have addressed this problem with varying degrees of success. Rule-based systems are fast, transparent, and inexpensive, but they depend on patterns anticipated by their designers. Prior work has shown that the usaddress library performs well on common street addresses but struggles with military, highway, and attention-line formats, where overall accuracy dropped to 47.55% [5]. Statistical approaches such as Conditional Random Fields and Hidden Markov Models improved generalization by learning from labeled examples [6,7], but their performance depends heavily on the coverage and quality of training data. Neural methods, including BiLSTM-CRF models, transformers, and BERT-based Named Entity Recognition systems, further improved entity recognition [8,9]. However, many such systems identify address spans at a coarse level rather than decomposing them into the fine-grained fields required for census processing.
Large Language Models introduced a new possibility for structured extraction. Since GPT-3 [10], instruction-following LLMs have shown that complex information extraction tasks can be performed with little or no task-specific training data. In address parsing, an LLM can use natural language instructions to infer that an APO or FPO value functions as a city, that AE or AP functions as a military state code, or that an attention line should be separated from the physical address. Prior cloud-hosted LLM work reported strong address-parsing results on its own evaluated corpus [11]. Because that study used a different dataset, model, deployment setting, and metric, its results are treated as related-work context rather than a direct benchmark.
Despite this progress, two major challenges remain. The first is extraction reliability when input scope and output structure become complex. Prior research has documented position-dependent degradation in some long-context language-model tasks [12]. This literature motivated limiting the amount of text presented in individual calls. However, the present study did not conduct a position-controlled long-context evaluation. The aggregate evaluation summary reports that the monolithic reference did not produce usable output for 47 of 700 generated records, but the causes of those omissions were not isolated and the retained artifacts do not provide the original prompt-position mapping. The omissions are therefore not interpreted as evidence of a “lost in the middle” effect. The second challenge is privacy. Census data contains highly sensitive personally identifiable information, and sending such data to cloud-hosted LLM services can create governance and compliance concerns. Any operational system must therefore address secure deployment as well as extraction quality.
Open-weight language models make a different deployment model possible. In 2025, OpenAI released gpt-oss-20b, a 21-billion-parameter open-weight language model under the Apache 2.0 license [13]. Because the model can be deployed on local infrastructure, sensitive records can be processed without transmitting data to an external service. This changes the practical design space for census extraction. Instead of choosing between brittle local rule-based systems and powerful cloud-hosted LLMs, it becomes possible to investigate whether a locally deployed open-weight model can provide useful extraction quality while keeping data within the organization’s secure environment.
This paper presents a hierarchical multi-agent architecture for fine-grained parsing of census-style addresses using locally deployed gpt-oss-20b. The system divides processing into three stages. A Planner Agent analyzes the input document and formulates an extraction strategy. A Manager Agent converts that strategy into a dependency-aware task graph. A fleet of specialized Worker Agents performs extraction, semantic consistency review, and formatting. The prototype also implements workers for other PII types, but their task-level accuracy is not evaluated in this study. The design keeps inference within locally controlled infrastructure while providing structured interfaces for orchestration.
We conduct a controlled pilot on 700 author-generated synthetic records across seven predefined address categories: standard, individual, university, military, rural route, highway, and attention-line addresses. Across three repeated executions on the same fixed benchmark, the complete hierarchical configuration obtained a mean micro-averaged address-component exact-match score of 95.7%, compared with 80.9% for the monolithic reference; the deterministic rule-based parser obtained 52.3%. Category-level results describe the controlled benchmark only. The three-execution design does not support inferential claims about which address categories exhibit larger performance differences. These values measure agreement with generator-derived fields under the benchmark’s controlled conventions and are not estimates of performance on naturally occurring census records. We also discuss two published census parsing systems—a pattern-based active learning system [5] and a cloud-based LLM system [11]—only as contextual examples with different datasets, models, human-involvement assumptions, and deployment constraints.
The study is guided by the following research questions:
- RQ1: What address-component exact match, record coverage, and structural completeness does the complete hierarchical configuration obtain on the controlled author-generated synthetic benchmark?
- RQ2: How do the rule-based, monolithic LLM, and hierarchical configurations differ descriptively in benchmark performance and reported runtime, and what resource-accounting differences can be verified from the retained artifacts?
- RQ3: What qualitative error patterns, validation limitations, and operational trade-offs are observed in the evaluated configurations?
The importance of this work lies in the intersection of three requirements that are often treated separately: structured extraction, operational autonomy, and locally controlled deployment. Rule-based systems can run locally but require ongoing pattern engineering. Cloud-hosted LLM systems may be difficult to use with sensitive census data. This paper specifies an integrated multi-agent configuration that distributes extraction across specialized agents while keeping inference on local infrastructure, and it characterizes the complete configuration in a controlled synthetic pilot.
The paper makes three main contributions:
- Domain-specific system design. We specify and implement a Planner–Manager–Worker pipeline for fine-grained census-style address parsing, with structured interfaces, record-preserving chunking, same-model semantic consistency review, formatting, and execution logging. The prototype also contains an exploratory bounded retry pathway; its reliability is not evaluated as a contribution of the present study.
- On-premise implementation. We deploy gpt-oss-20b on local infrastructure with structured schemas, orchestration contracts, and execution logging so that prototype inference does not require sending input records to an external model provider. Local execution is an operational property, not by itself a complete privacy or security guarantee.
- A controlled synthetic address-pilot characterization. We compare three complete configurations descriptively on 700 author-generated records spanning seven predefined address categories. The evaluation reports address-component exact match, coverage, completeness, fixed prompt-template accounting, reported runtime, and limitations while identifying the run-level artifacts that were not retained. The implementation and reproducibility materials are publicly available at https://github.com/MohdMuzakkiruddinAhmed/census-pii-agents (accessed on 22 July 2026).
This work is a domain-specific system-design contribution accompanied by a controlled synthetic address-parsing pilot, not a production-readiness evaluation. All evaluation records are author-generated and synthetic, and no real census records or real personal information were used. The reported percentages characterize the generator-induced address benchmark only; they are not estimates of operational census accuracy, real-record robustness, broad PII-extraction accuracy, or any component’s independent effect.
The remainder of the paper is organized as follows. Section 2 reviews prior work on address parsing, name parsing, autonomous agent systems, and census-specific parsing tools. Section 3 describes the proposed architecture in detail. Section 4 presents the implementation, including local model deployment and synthetic data generation. Section 5 reports the evaluation methodology and results. Section 6 discusses the findings, limitations, and trade-offs. Section 7 concludes the paper and outlines future work.
2. Related Work
Structured knowledge extraction from census-style records draws on several related research areas: rule-based and statistical parsing, neural information extraction, Large Language Models, autonomous multi-agent systems, and privacy-preserving deployment. This section reviews these areas with emphasis on the specific requirements of the present work: fine-grained decomposition of personal records, robustness across diverse address formats, reduced human intervention, and secure local processing.
2.1. Parsing Personal Information: From Rules to Neural Networks
Early approaches to structuring personal information relied on human-crafted rules and domain-specific pattern libraries. For names, early work showed that personal naming conventions vary substantially across cultural and linguistic groups, making simple dictionary or position-based rules difficult to generalize [14]. For addresses, postal standardization systems such as the United States Postal Service Coding Accuracy Support System (CASS) established practical standards for conventional formats, while open-source tools such as usaddress and libpostal extended address parsing and normalization to broader use cases [15]. Rule-based NLP frameworks such as GATE further demonstrated the value of modular processing pipelines for information extraction [16]. These systems are attractive because they are fast, deterministic, transparent, and easy to deploy locally. Their main limitation is brittleness: handcrafted rules can only recognize patterns that have already been anticipated, which becomes problematic when records include uncommon, incomplete, or non-standard formats [2].
Statistical sequence models were developed to reduce this dependence on manually specified rules. Conditional Random Fields, Hidden Markov Models, and transformation-based learning allowed systems to learn token-level patterns from labeled data rather than relying entirely on handcrafted templates [6,7,17]. These methods improved generalization and were later extended to address parsing across multiple countries and address conventions [18]. However, their performance remains strongly tied to the coverage and quality of the training corpus. If the labeled examples are dominated by standard street addresses, the resulting model may struggle with military addresses, rural routes, highway references, university departments, or attention-line records.
Neural models further improved sequence labeling and entity recognition. BiLSTM-CRF architectures showed that recurrent neural networks could capture dependencies across address tokens [19], while transformers and BERT-based models changed the broader landscape of Named Entity Recognition [8,9,20]. These methods are powerful for detecting entity spans, but many systems operate at a coarse level, identifying that a span is an address without decomposing it into the fine-grained components needed for census processing. This distinction is important because downstream applications such as geocoding, record linkage, and identity resolution depend not only on detecting an address, but also on correctly assigning street number, street name, unit, city, state, ZIP code, organization, attention line, and other components [4].
2.2. Large Language Models for Information Extraction
Large Language Models introduced a different approach to structured extraction. GPT-3 showed that sufficiently large language models could perform many tasks with zero or few examples, including information extraction tasks that previously required task-specific training data [10]. For census-style parsing, this capability is important because an LLM can use natural language instructions and contextual reasoning to interpret address formats that may not be explicitly represented in a rule base or training set.
Several lines of LLM research directly inform the present architecture. Chain-of-Thought prompting has been shown to improve performance on tasks requiring multi-step reasoning [21], which motivates the use of a Planner Agent for document analysis and strategy formation. Work on constrained decoding and structured generation has shown that schema adherence is a central challenge for LLM-based extraction [22,23]. Research on feedback, retrieval, and factual correction has also shown that model outputs can be improved when generation is combined with structured review or external signals [24]. Together, these findings suggest that LLM extraction systems should not rely only on a single unconstrained prompt, especially when outputs must be machine-readable and auditable.
Task decomposition is another important theme. Prompt programming frameworks such as DSPy show that complex LLM tasks can often be improved by decomposing them into smaller composable steps [25]. Work on code-oriented and format-aware LLMs similarly suggests that models can perform well when prompts clearly specify structured input and output formats [26]. These findings motivate the separation of planning, task assignment, extraction, validation, and formatting in the proposed system.
At the same time, LLM-based extraction has known limitations. Prior research has documented position-dependent degradation in some long-context tasks [12]. Models can also be sensitive to prompt ordering [27], and self-assessed confidence scores are useful but imperfect indicators of correctness [28]. These findings provide general motivation for controlling input scope and using structured prompts. The present study did not randomize record position, compare beginning, middle, and end positions, or otherwise test position-dependent degradation; it therefore does not claim that the architecture mitigates a “lost in the middle” effect.
2.3. Autonomous Multi-Agent Systems
The idea that complex tasks can be improved by distributing work among specialized agents has a long history in artificial intelligence. Classical agent theory defines intelligent agents in terms of autonomy, reactivity, proactivity, and social ability [29], while broader AI frameworks distinguish deliberative, reactive, and hybrid agent architectures [30]. Multi-agent systems further show that teams of specialized agents can solve problems that are difficult for a single monolithic system to handle efficiently [31]. This principle is directly relevant to census extraction, where document analysis, address parsing, name extraction, validation, and formatting require different kinds of reasoning.
Recent LLM-based agent frameworks have made these ideas practical. ReAct demonstrated that reasoning and action can be interleaved to improve agent behavior [32]. AutoGen supported collaborative conversations among agents with different roles [33]. CrewAI introduced role-based hierarchical agent teams [34], and LangGraph provided graph-based orchestration for complex agent workflows [35]. These systems show that LLM-powered agents can be coordinated through roles, dependencies, and structured workflows rather than being used only as isolated prompt-response models.
Prior work on entity-level tasks has also shown that specialization can improve performance. A LangGraph-based entity resolution framework with multiple task-specific agents achieved strong accuracy on name variation matching while reducing API calls compared with a single-LLM baseline [36]. This supports the design decision in the present work to assign separate responsibilities to address extraction, name extraction, validation, and formatting agents rather than asking one general-purpose prompt to perform all tasks simultaneously.
Self-correction and feedback mechanisms provide background for an exploratory feature of the prototype. Reflexion studied verbal self-reflection [37], and Self-Refine studied iterative revision using structured feedback [38]. Motivated by this work, the prototype can route the selected downstream error context to the Planner before a bounded revised attempt. The present study does not evaluate whether this pathway is more reliable than an ordinary retry or single-agent refinement.
2.4. Census-Specific Parsing and Privacy-Preserving Deployment
The work most directly related to the present study concerns census-specific parsing systems and data-local AI deployment. Prior pattern-based active-learning work demonstrated the potential of persistent expert mappings on its own evaluation corpus [5]. That system tokenized input strings, created token masks, and mapped those masks to functional categories through user-defined mappings. When an input did not match an existing mapping, it was routed to a domain expert for correction. Its principal design distinction is that expert corrections permanently enrich the pattern base. Novel formats, however, require manual analysis and rule extension, creating operational overhead as new patterns appear.
The semantic component definitions from that line of work remain useful for the present study. What changes is the assignment mechanism. Rather than relying on stored token-mask mappings, the proposed architecture delegates token assignment to contextual reasoning performed by specialized worker agents. This shift is intended to reduce the need for format-specific manual rule engineering while preserving fine-grained component decomposition.
Cloud-hosted LLM systems demonstrate a different design and deployment choice. A prompt-driven, validation-centered census address parsing framework used Claude 4.0 Sonnet through AWS Bedrock and reported strong results on its own evaluated corpus [11]. Because that study used a different corpus, metric, model, and serving environment, its results are not a numerical benchmark for the present pilot. Cloud-hosted inference also raises governance concerns for sensitive census data, especially when records contain personally identifiable information protected by strict legal and institutional requirements. In addition, the prior system used a single-pipeline design rather than the coordinated task decomposition implemented here.
Related work from entity resolution, household movement detection, multilingual record linkage, and multi-model extraction further supports the use of LLM reasoning for personal and organizational information processing. LLM-based systems have been used to detect co-residence patterns from unstandardized names and addresses [39], compare multiple LLMs for household movement discovery [40], perform cross-lingual entity resolution [41], and coordinate multiple models for industrial part specification extraction [42]. These studies indicate that LLMs can support complex knowledge extraction tasks, but they do not directly resolve the combined requirements of fine-grained census address decomposition, local deployment, traceable pipeline execution, and comparison against census-specific baselines.
Locally controlled deployment is therefore central to this paper. Open-weight language models, including the LLaMA family, Mistral mixture-of-experts models, and gpt-oss-20b, have made it increasingly feasible to run capable models on local infrastructure [13,43]. Other privacy-oriented methods, including federated learning and differential privacy, provide formal mechanisms for distributed or statistical privacy protection [44,45]. Policy-aware generative AI controllers have also shown how LLM reasoning can be constrained by explicit governance rules in regulated environments [46]. The present prototype performs model inference on-premise so that benchmark inputs are not transmitted to an external model provider; this deployment property is not a formal privacy guarantee or a complete security model.
2.5. Summary and Research Gaps
The literature shows that existing approaches address different parts of the census extraction problem, but few address the full set of requirements together. Rule-based and pattern-based systems can be locally deployable but may require continuing expert involvement when new formats appear. Cloud-hosted LLM systems provide prompt-driven extraction but introduce governance concerns for sensitive records. Single-prompt local LLM systems keep inference data on local infrastructure but may face challenges with long, heterogeneous documents. Multi-agent LLM architectures provide a promising way to distribute complex extraction tasks, but their use for fine-grained census-style address decomposition remains underexplored.
Table 1 summarizes how the main research streams relate to the proposed system.
Table 1.
Positioning of the proposed system relative to prior approaches.
To the best of our knowledge, the literature has not yet provided a systematic evaluation of a locally deployed open-weight multi-agent LLM architecture for fine-grained census-style address parsing across diverse address categories. Four gaps motivate the proposed system:
- Fine-grained census-style decomposition. Existing systems do not fully address the problem of decomposing diverse address formats into detailed components across standard, individual, university, military, rural route, highway, and attention-line categories.
- Joint treatment of privacy and extraction quality. Prior work often treats local deployment and high-quality extraction as separate goals rather than evaluating them together in one architecture.
- Reduced dependence on manual pattern engineering. Pattern-based systems can achieve high accuracy, but new formats may require expert intervention. More autonomous alternatives remain needed.
- Practical trade-off analysis. The accuracy, latency, cost, privacy, and human-involvement trade-offs among rule-based parsers, single-prompt LLM baselines, cloud-hosted LLM systems, and multi-agent local LLM systems have not been sufficiently characterized.
The system presented in this paper is designed to address these gaps by combining fine-grained structured extraction, local open-weight model deployment, specialized multi-agent decomposition, and descriptive comparison against general and census-specific references.
3. Methodology and System Architecture
This section describes the methodological foundations, design principles, and system architecture of the proposed multi-agent framework, with emphasis on the fine-grained census-style address-parsing pathway evaluated in this study. The system receives a raw document containing one or more personal records and produces structured outputs containing decomposed address components, additional implemented PII fields, semantic-review results, and an execution log. The architecture separates processing into three stages: planning, coordination, and execution. The prototype also contains an exploratory bounded retry pathway, specified in Section 3.5, whose reliability and causal benefit are not evaluated. Only the address-component outputs are quantitatively evaluated in this paper.
3.1. Methodological Foundations
The proposed architecture is built on four methodological choices: multi-agent decomposition, local LLM-based reasoning, planning with Chain-of-Thought prompting, and schema-constrained structured output. Each choice addresses a specific challenge in census-style extraction.
First, the system uses multi-agent decomposition. A Multi-Agent System (MAS) consists of multiple interacting agents that coordinate to solve problems that may be difficult for a single agent to handle alone [29,31]. Census-style extraction is naturally decomposable because address parsing, name extraction, phone number recognition, validation, and formatting require different forms of reasoning. Assigning all of these responsibilities to a single prompt increases prompt scope and structural complexity. Prior work on cognitive load and LLM task decomposition motivates the use of smaller, specialized tasks rather than a single monolithic prompt [25,47]. The present evaluation did not isolate prompt complexity or record position as a cause of output omissions. The proposed system operationalizes task decomposition through a three-level hierarchy: a Planner Agent for strategic reasoning, a Manager Agent for task coordination, and a fleet of Worker Agents for extraction, validation, and formatting [30].
Second, the system uses Large Language Models as local reasoning engines. All LLM-based components run on gpt-oss-20b, a 21-billion-parameter Mixture-of-Experts model with 3.6 billion active parameters per inference call [13]. This model is used because it supports instruction-following, structured reasoning over text, and on-premise deployment. These properties are important for census-style records, where the system must interpret heterogeneous formats while avoiding transmission of sensitive records to an external model provider. The model’s open-weight availability, efficient inference profile, and long context window make it suitable for local deployment in the proposed pipeline.
Third, the Planner Agent uses Chain-of-Thought prompting to support document-level analysis. Chain-of-Thought prompting has been shown to improve performance on tasks requiring multi-step reasoning [21]. In this system, the Planner must identify the document format, estimate record structure, detect relevant PII types, recognize address categories, and formulate extraction guidance. These operations require broader reasoning than the worker-level extraction tasks, so Chain-of-Thought prompting is applied only at the planning stage.
Fourth, all LLM agents use schema-constrained prompt engineering. Each agent receives a role specification, task instructions, and an explicit JSON output example. This design encourages outputs that conform to predefined schemas and supports deterministic downstream parsing [23,24]. Because LLMs may still produce extraneous text, markdown wrappers, or malformed JSON, schema prompting is supplemented with post-hoc JSON extraction and fallback logic, described in Section 4.5.
3.2. Design Principles
Four design principles govern the architecture. Each principle addresses a practical risk associated with LLM-based structured extraction.
P1: Separation of Concerns.Strategic reasoning, task coordination, operational extraction, validation, and formatting are handled by distinct components. This prevents a single agent from simultaneously analyzing document structure, extracting multiple PII types, validating outputs, and formatting final records. The design directly addresses the quality degradation observed when long heterogeneous extraction tasks are handled by a single prompt [12].
P2: Bounded Cognition. Each Worker Agent is assigned a narrow responsibility. For example, the address worker handles address decomposition, the name worker handles name extraction, and the formatter handles output assembly. Limiting each worker’s scope reduces prompt complexity and helps mitigate cognitive overload [12,47].
P3: Graceful Degradation. Every LLM-dependent component is paired with deterministic fallback behavior. If the Planner Agent fails to produce a valid plan, a default plan is substituted. If the Manager Agent fails to generate a valid task graph, a minimum viable graph is constructed from detected PII types. If a Worker Agent fails after exhausting configured retry attempts, the failure is logged and the pipeline continues with partial results. This ensures that a single failure does not halt the entire extraction process.
P4: Observability. All major actions are logged with timestamps, component identifiers, status codes, and execution metadata. The execution log supports debugging, reproducibility, runtime analysis, and future audit requirements for regulated census-style processing. In production settings, this logging must be implemented with appropriate safeguards so that sensitive information is not exposed through logs.
3.3. Architectural Overview
The system consists of five major components: a Document Processor, a Planner Agent, a Manager Agent, a Worker Agent Fleet, and an Orchestrator. The Planner, Manager, and Worker Agents are LLM-based. The Document Processor, Orchestrator, JSON extraction logic, fallback logic, and logging are deterministic. The current review stage is an LLM-based semantic consistency check using the same model as the extractors; it is not independent validation. Production deployments should add deterministic controls such as typed-schema enforcement, regular-expression format checks, finite-domain checks, state–ZIP consistency rules, and postal-reference checks.
Figure 1 illustrates the architecture and data flow. A raw input document is first processed into text, format metadata, and chunks. The Planner Agent analyzes the document and produces an extraction plan. The Manager Agent converts that plan into a dependency-aware task graph. Worker Agents then execute extraction tasks over the chunks. Semantic-review and formatting workers consolidate the results into final structured records. If a worker fails or produces output that violates a configured check, the error context can be sent back to the Planner Agent for bounded re-planning.
Figure 1.
System architecture. Solid arrows indicate data flow; the dashed retry-pathway arrow indicates the recursive re-planning path. The Orchestrator is a deterministic controller, not an LLM-powered agent.
Table 2 summarizes the role, type, input, and output of each major component.
Table 2.
Summary of major system components.
The system is described as autonomous in the sense of reactive autonomy [29]. It executes the complete pipeline from raw input to structured output without human intervention during a run, including error logging, fallback handling, and configured retry routing. This describes automated execution, not validated error-recovery effectiveness. The system follows a fixed processing architecture rather than performing unrestricted goal-directed planning over arbitrary actions. All LLM inference occurs on a locally deployed gpt-oss-20b instance, so input records are not transmitted to an external model provider.
- Stage 1: Planner Agent.
The Planner receives the raw document text sample, limited to the first 3000 characters, together with the detected file format. The 3000-character sample provides enough context to infer document structure while keeping the planning prompt compact and consistent across runs. Using Chain-of-Thought prompting, the Planner produces a structured JSON plan with three main sections: (i) an analysis section containing detected format, estimated record count, identified PII types, and identified address categories; (ii) a roadmap section specifying phased execution and inter-phase dependencies; and (iii) an extraction strategy section providing guidance for the address categories and PII types detected in the document. If the Planner fails to produce valid JSON, a default plan is used. When the exploratory retry pathway is invoked, the Planner can also receive an error context and produce a revised plan.
- Stage 2: Manager Agent.
The Manager receives the Planner output and converts it into a dependency-aware task graph.
Definition 1 (Task Graph).
Let be a directed acyclic graph where each task is a tuple . The tuple specifies a unique task identifier, the PII type or operation to perform, task-specific instructions derived from the Planner strategy, dependencies that must complete first, and task priority. An edge indicates that task must complete before task begins.
In a typical run, the Manager produces six independent extraction tasks for addresses, names, phone numbers, email addresses, Social Security Numbers, and dates of birth. These are followed by a task labeled “validation” in the implementation that performs same-model semantic consistency review, and then by a formatting task. If the Manager fails to produce a valid task graph, the Orchestrator constructs a deterministic minimum viable graph from the PII types identified by the Planner.
- Stage 3: Worker Agent Fleet.
Eight specialized Worker Agents execute the tasks assigned by the Manager. These include workers for address extraction, name extraction, phone extraction, email extraction, Social Security Number extraction, date-of-birth extraction, validation, and formatting. All workers operate at temperature 0.1 for near-deterministic generation and use schema-constrained prompts. The following subsections describe the most architecturally significant workers.
3.3.1. Address Extraction Worker
The address extraction worker is the most complex worker in the fleet because census-style records include both standard postal formats and non-standard address structures. Its output schema contains fifteen fields, shown in Table 3. These fields were selected to cover conventional address components as well as census-relevant components such as organization names, attention lines, military city and state equivalents, and route classifications.
Table 3.
Address extraction output schema with fifteen components.
The worker prompt contains category-specific guidance covering the seven address categories evaluated in this study: standard street addresses, individual unit addresses, university department addresses, military APO/FPO/DPO addresses, rural route addresses, highway-referenced addresses, and attention-line records. This guidance specifies, for example, that APO, FPO, and DPO should be mapped to the city field; that AE, AP, and AA are military state equivalents; that route numbers in highway addresses form part of the street_name rather than the unit_number; and that attention-line content should be separated from organization names. Exact token accounting for the retained prompt templates is reported in Section 4.4. Because the worker relies on contextual reasoning rather than fixed token-mask rules, it can handle variations that are not explicitly enumerated in the prompt, although this capability is not a substitute for independent validation [5].
The confidence field contains a self-assessed value in . Such confidence scores are treated as soft signals rather than definitive correctness indicators because prior work has shown that model confidence is meaningful but imperfect [28]. In this system, confidence values are used to prioritize validation and to identify outputs that may require a configured retry attempt or later human review.
3.3.2. Name Extraction Worker
The name extraction worker produces a seven-field output: full_name, first_name, middle_name, last_name, prefix, suffix, and confidence. The prompt includes guidance for compound surnames, hyphenated names, multiple middle names, single-name individuals, professional titles, and generational suffixes. Because personal names vary across cultural and linguistic groups, dictionary-based or position-based parsing can be insufficient [14]. The LLM-based worker provides broader contextual coverage, although systematic bias evaluation across cultural groups remains necessary, as discussed in Section 6.4.
3.3.3. Semantic Consistency Review and Formatting
The component originally termed the validation worker performs an LLM-based semantic consistency review of upstream outputs. Its prompt asks the same gpt-oss-20b checkpoint used by the extraction workers to inspect apparent schema compliance, completeness, internal consistency, military state-code usage, state/ZIP alignment, and obvious field conflation. This stage is not an independent source of truth and may reproduce or accept errors correlated with the extraction workers. It is therefore described as semantic consistency review rather than independent validation. Table 4 summarizes the validation layers considered in the prototype and distinguishes implemented controls from deployment recommendations.
Table 4.
Layered validation controls and their status in the reported prototype evaluation.
The formatting worker merges outputs from the extraction and semantic-review workers into unified per-record structures. It preserves document order, resolves field-level conflicts when possible, and produces the final structured output expected by downstream processing.
The implemented review stage can flag apparent schema, completeness, and contextual inconsistencies, but it cannot certify factual correctness. Stronger deployment would require deterministic schema and field-type checks, regular-expression format checks, finite-domain code validation, cross-field consistency rules, postal-reference checks, secure logging, and human escalation for unresolved or consequential cases. These safeguards are deployment recommendations; they were not part of the reported evaluation and are not credited with improving its benchmark results.
3.3.4. Other PII Workers and Evaluation Scope
The prototype also implements workers for names, phone numbers, email addresses, Social Security Numbers, and dates of birth. Each follows the same schema-constrained prompt pattern: role declaration, output schema, example output, and extraction instructions tailored to the specific PII type. However, the quantitative evaluation in this study is limited to fine-grained address-component parsing. No task-level accuracy, robustness, or generalization conclusion is made for these additional PII workers.
3.4. Chunk-Level Processing
Documents exceeding 3000 characters are segmented into chunks at natural line boundaries. This strategy limits the amount of source text presented in each worker call and keeps individual processing units bounded. Whenever possible, chunk boundaries are selected so that individual records are not split across chunks. If a single record exceeds the chunk target, the system preserves the complete record as one chunk rather than splitting it internally.
Each Worker Agent processes chunks independently, and the outputs are merged in the original document order:
where w denotes the worker function, denotes chunk j, and ⊕ denotes ordered concatenation of structured result sets. This strategy limits the source text presented in each worker call and permits failures to be isolated to smaller processing units: a failure on one chunk can be logged, can enter a configured retry attempt, or can fall back to partial output without invalidating all other chunks. Prior long-context research motivated controlling input scope [12], but the present study did not evaluate position-dependent degradation and does not establish that chunking mitigated a “lost in the middle” effect.
3.5. Exploratory Bounded Retry Pathway
The prototype includes a bounded retry pathway intended to support recovery from selected execution failures. This subsection specifies the intended control flow of that pathway. The present evaluation does not establish its reliability, causal benefit, or robustness. Accordingly, the mechanism is treated as an exploratory design feature rather than as an empirically validated component.
When Worker Agent fails during execution, the system can activate the bounded retry pathway shown in Algorithm 1. In this context, failure includes invalid JSON, schema violation, missing required output structures, an explicit error object, or a configured semantic-review failure. Low-confidence output may also trigger the pathway when a configured check marks it as unreliable. The pathway is designed to permit a revised attempt within a run but does not permanently update the model, prompts, or stored knowledge.
The distinguishing design feature is that error context can be routed across agent boundaries. The failure is detected at the extraction or semantic-review stage, while the Planner can formulate a revised attempt using broader document and strategy context. Whether this produces more reliable recovery than an ordinary retry or single-agent refinement was not evaluated. No feedback-disabled comparison was conducted, and any changed post-retry output may reflect both the revised plan and ordinary stochastic generation variation. The pathway is non-persistent across runs; the same failure may recur unless prompts, deterministic checks, or fallback logic are updated.
| Algorithm 1 Exploratory Bounded Retry Pathway. |
|
3.6. Orchestrator
The Orchestrator is a deterministic Python controller, not an LLM-powered agent. It manages document processing, planning, task-graph construction, worker execution, configured retry-pathway routing, fallback behavior, and final output compilation. Algorithm 2 specifies the complete pipeline.
| Algorithm 2 Orchestrator Pipeline. |
|
The Orchestrator enforces the topological ordering of the task graph. Extraction tasks can execute before validation, validation must complete before formatting, and formatting produces the final structured records. The execution log records each component invocation, status, output summary, and wall-clock time. This supports debugging and runtime analysis while also establishing the basis for auditable operation in future regulated deployments.
3.7. Design Alternatives and Rationale
Several architectural alternatives were considered during system design. For inter-agent communication, JSON contracts were chosen over natural language messaging because JSON supports deterministic parsing, schema validation, and repeatable downstream processing [33]. The Manager was implemented as an LLM-powered agent rather than a hardcoded task graph so that task assignment could adapt to the PII types and address categories detected in the input document. Worker Agents were specialized by PII type rather than by address category because PII types are stable across datasets, while address categories may vary or overlap.
For the exploratory retry design, error context is routed to the Planner because it has broader document-level context than an individual worker [37]. This was an implementation choice, not an experimentally established advantage over ordinary retry or single-agent refinement. For structured output, post-hoc JSON extraction was selected instead of constrained decoding because it is model-agnostic and compatible with common local serving frameworks [22]. The retry pathway was bounded at to keep runtime finite and avoid unbounded recursion.
4. Implementation
This section describes the implementation of the proposed multi-agent extraction system, including the hardware and software stack, local model deployment, runtime configuration, prompt design, JSON recovery mechanism, and synthetic data generation process. The implementation was designed around three practical requirements: the system should run on local infrastructure, produce repeatable experimental outputs, and avoid the use of real census records or real personally identifiable information during evaluation.
4.1. Technology Stack and Hardware Configuration
Table 5 summarizes the technology stack and hardware configuration used throughout the experiments.
Table 5.
Technology stack and hardware configuration.
The hardware configuration which is shown in Table 6 was selected to test whether the full multi-agent pipeline could run on professional workstation hardware rather than requiring cloud-hosted models or data-center GPU infrastructure. Prior LLM-based census parsing work has relied on cloud-hosted commercial models [11]. In contrast, this implementation runs the complete pipeline on a local workstation equipped with a 24 GB GPU. The 24 GB GPU should be interpreted as the evaluated workstation-class reference configuration, not as a universal minimum or a sufficient national-scale production configuration.
Table 6.
Tested hardware configuration and production implications. The study evaluated one workstation configuration and did not conduct hardware-sizing experiments.
The model footprint alone does not determine deployment capacity. Memory used by the inference engine, temporary tensors, active contexts, batching, and the key–value cache affects the prompt length and concurrency that a server can support. Lower-memory GPUs, CPU-only execution, alternative quantization levels, and multi-replica configurations were not evaluated; no performance or minimum-memory claim is made for those settings.
The implementation and reproducibility materials are publicly available at https://github.com/MohdMuzakkiruddinAhmed/census-pii-agents (accessed on 22 July 2026). The repository provides the Planner, Manager, Worker, and orchestration code; agent prompt templates; synthetic-data generation script; evaluation scripts and baselines; package configuration; and execution instructions. It identifies the evaluated base model as openai/gpt-oss-20b and documents the fixed synthetic-data generation command using 700 records and random seed 42. The repository therefore permits the controlled benchmark and evaluation workflow to be regenerated without access to real census records or personal information.
4.2. On-Premise Model Deployment
The gpt-oss-20b model is deployed locally using a 4-bit GPTQ quantized configuration. Quantization reduces the approximate memory footprint from 42 GB at half precision (FP16) to approximately 12 GB, which fits within the 24 GB memory capacity of the NVIDIA RTX PRO 5000 Blackwell GPU. The remaining memory is used for inference overhead and key–value cache storage during generation. This configuration supports the prompt templates and document chunks used in the evaluation while preserving the model’s long-context capability.
The model is served through vLLM as a local HTTP endpoint bound to the host machine. The endpoint is not exposed to external traffic, and the experimental pipeline does not transmit input records to any external model provider. This deployment provides a strong operational privacy posture because inference occurs inside the local execution environment. However, local deployment alone is not a complete security model. Production use with real census data would also require access controls, encrypted storage, secure deletion policies, log redaction, endpoint hardening, audit procedures, and governance review.
4.3. Runtime Configuration
All LLM-based agents use the same local model endpoint. Inference calls are configured with a temperature of 0.1 to encourage near-deterministic outputs while still allowing the model to generate complete structured responses. The system applies exponential backoff retry logic with a maximum of three retries for transient inference failures, such as local timeout or temporary server unavailability.
The runtime pipeline is controlled by a deterministic Python Orchestrator. The Orchestrator manages document preprocessing, Planner invocation, Manager invocation, Worker Agent execution, configured retry-pathway invocations, fallback behavior, and final output compilation. It also records execution metadata, including component name, task identifier, status, retry count, and wall-clock time. Because execution logs may contain sensitive values in production settings, the prototype logging design should be paired with redaction or secure logging before use on real records.
4.4. Prompt Engineering
Each agent receives a structured system prompt designed to constrain the model’s role, scope, and output format. The prompts follow a common template, summarized in Table 7.
Table 7.
Common prompt structure used across LLM-based agents.
The retained address-worker implementation uses a dedicated address-extraction prompt containing the output schema and category-specific instructions. We counted the fixed prompt text with the tokenizer associated with the evaluated openai/gpt-oss-20b artifact. The address worker’s fixed system prompt contains 302 model tokens, including a 68-token schema substring and 201 tokens of category guidance. The retained address-only monolithic reference system prompt contains 176 tokens, including a 43-token schema substring and 99 tokens of category guidance. When the fixed user wrapper and chat-template framing are included with an empty dynamic input, the fixed per-call totals are 390 tokens for the address worker and 252 tokens for the monolithic reference. These prompt allocations were not matched. The configurations also differ in task scope, same-model semantic review, formatting, and orchestration. Dedicated address guidance is therefore treated as one bundled feature of the complete hierarchical configuration rather than as an independently evaluated architectural mechanism. Dynamic input tokens, aggregate prompt-token totals, and realized call totals for the reported runs were not retained and cannot be reconstructed without the original execution logs; we do not estimate them retrospectively.
The Planner Agent uses a different prompt pattern. In addition to the JSON schema, it receives Chain-of-Thought style instructions for document analysis. The Planner is asked to infer the document structure, identify likely PII types, recognize address categories, and produce an extraction strategy. Worker Agents do not receive the same broad reasoning instructions because their tasks are intentionally narrower.
4.5. Robust JSON Extraction
Although agents are instructed to produce JSON-only output, LLM responses are not always syntactically valid JSON. In preliminary runs, approximately 18% of responses included markdown code fences, explanatory preambles, trailing commentary, or other extraneous content. To prevent these responses from halting the pipeline, the implementation uses a multi-stage JSON extraction function, safe_json().
The safe_json() function applies the following recovery sequence:
- Remove markdown code fences, language identifiers, and leading or trailing whitespace.
- Attempt standard parsing with json.loads().
- If parsing fails, locate the first opening brace { and the last balanced closing brace }, extract the candidate JSON substring, and attempt parsing again.
- If parsing still fails, return a default dictionary matching the expected schema, with empty fields or explicit error indicators.
- Log the parsing failure, raw response summary, recovery step, and fallback status for later inspection.
This post-hoc extraction strategy is model-agnostic and does not require changes to the inference engine. It was therefore preferred over constrained decoding in the prototype implementation, although constrained decoding remains a useful alternative for future systems that require stronger output guarantees [22]. The fallback dictionary ensures graceful degradation: the pipeline can continue even when one agent response cannot be parsed, while the execution log preserves evidence of the failure.
4.6. Synthetic Data Generation
The evaluation dataset contains 700 author-generated synthetic census-style records. The dataset was created programmatically in Python using stratified sampling across seven predefined address categories, with 100 records per category. Each record was first generated as a structured object containing reference fields, then rendered into textual record formats with controlled variation in punctuation, abbreviation style, line breaks, field order, and optional components.
4.6.1. Benchmark Construction and Interpretation
For each synthetic evaluation item i, the generator first creates a structured record . A rendering function g, together with sampled formatting variation , converts the structured record into textual input:
The extraction system receives and produces a predicted structured record . Evaluation compares with the original structured object .
This construction provides complete and internally consistent reference fields relative to the generator, but the references are not independent annotations of naturally occurring records. The textual input and scoring target share the taxonomy, formatting assumptions, and component definitions encoded by the generator. The benchmark therefore evaluates whether a system can recover structured fields from records produced under this controlled rendering process. It does not estimate performance on naturally occurring census submissions or records produced by an independent data-generating process.
Author review was used to check consistency between the rendered text and its source structured object. This review should not be interpreted as independent annotation because the authors had access to, and were validating, the output of the same generation procedure. The benchmark may consequently be easier than real-record parsing and may favor systems whose instructions align with the generator’s component definitions.
Table 8 summarizes the seven address categories and the main variations included in each category.
Table 8.
Synthetic data categories and generated variations.
4.6.2. Synthetic Data Quality Assurance
The retained generator and repository support verification of dataset construction, but they do not contain a separate annotation log, formal QA report, documented review sample, reviewer-count record, discrepancy log, or adjudication record. The phrase “author review” in the original description referred to informal consistency checks in which selected rendered records were compared with their source structured objects. The number and selection of reviewed records, the number of participating authors, and any corrections made during that informal process were not recorded as part of a formal protocol. These checks are therefore described as informal author consistency review, not independent annotation.
Table 9 separates construction properties supported by the retained generator from checks whose scope was not documented.
Table 9.
Synthetic-data quality controls recoverable from retained artifacts and the evidentiary scope of each control.
Each synthetic record also includes a generated personal name, phone number, email address, Social Security Number, and date of birth. Name generation includes variation in prefixes, suffixes, middle names, compound names, and hyphenated forms. Phone numbers, email addresses, Social Security Numbers, and dates of birth are generated to resemble valid formats without corresponding to real individuals.
All data used in the evaluation is fictitious. No real census records, real addresses, real Social Security Numbers, or real personal information were used at any stage of this research. The benchmark does not adequately represent misspellings, OCR and transcription artifacts, handwriting errors, multilingual content, incomplete responses, duplicated or contradictory fields, uncontrolled field ordering, or formats outside the seven predefined categories. The generation script is publicly available at https://github.com/MohdMuzakkiruddinAhmed/census-pii-agents (accessed on 22 July 2026). The repository documents the command used to regenerate the 700-record synthetic benchmark with random seed 42; consequently, no data-governance review is required to access the released reproducibility materials.
5. Evaluation
This section evaluates fine-grained address-component parsing through benchmark exact match, record coverage, structural completeness, qualitative error analysis, reported runtime, and retained resource-accounting evidence. Although the prototype contains additional PII workers, their task-level accuracy is outside the empirical scope of this study. All experiments use author-generated synthetic address records and generator-derived reference fields. The results describe execution feasibility and address-benchmark behavior under this controlled process rather than production accuracy, broad PII-extraction performance, or real-record robustness.
5.1. Evaluation Goals and Scope
The evaluation has three goals. First, it measures address-component exact match, output coverage, and structural completeness on the controlled author-generated benchmark. Second, it compares the three complete configurations descriptively while reporting wall-clock time and the resource information recoverable from retained artifacts. Third, it characterizes qualitative error patterns, same-model review limitations, and operational trade-offs. The three observed retry-pathway activations are documented only as illustrative execution traces and are not treated as an evaluation goal or an answer to a research question.
Table 10 summarizes how the evaluation addresses each research question.
Table 10.
Research questions and corresponding evaluation evidence.
Several limitations should be acknowledged before presenting the results. First, the textual records and reference labels originate from the same author-controlled generation process. Author consistency review is not independent annotation and does not remove this shared-source dependency. Second, the benchmark omits many irregularities of naturally occurring records, so its percentages do not estimate real-record performance. Third, only one open-weight LLM, gpt-oss-20b, was evaluated. Fourth, each LLM configuration was executed only three times on the same fixed benchmark. These repetitions describe limited execution-to-execution variability but are not independently sampled datasets and do not support confidence intervals, formal significance testing, or category rankings. Fifth, the original raw predictions, prompts as submitted, execution logs, and record-position mappings were not retained with the evaluation repository. The supplied repository runs both LLM conditions one record at a time and therefore does not reproduce the previously reported batch-call totals. Fixed prompt-template counts can be recovered, but realized call totals, dynamic input tokens, aggregate prompt allocation, and the batching and segmentation used for the reported summaries cannot be verified. Sixth, the quantitative metrics score address components only; implemented name, phone, email, Social Security number, and date-of-birth workers were not evaluated separately. Finally, the semantic-review worker uses the same model checkpoint as the extraction workers and is not an independent verifier. The evaluation is therefore a descriptive address-parsing pilot comparison of aggregate results on the generator-induced benchmark, not a real-record validation or a component-level, position-controlled, prompt-controlled, compute-controlled, independently validated, broad-PII, or inferential comparison.
5.2. Evaluation Dataset and Experimental Conditions
The evaluation dataset is identical to the synthetic dataset described in Section 4.6. It contains 700 records distributed equally across seven address categories. The category selection follows the same taxonomy used in prior census parsing work [5], which supports comparison with established address-format categories.
Three systems were evaluated on the same 700-record dataset:
- Rule-Based Parser. For the reported address evaluation, a deterministic reference using the Python usaddress library to map address tokens into the evaluated component schema.
- Monolithic Reference Configuration. An address-only gpt-oss-20b reference. The retained implementation receives one census-style address line, requests a 15-field address schema, and includes explicit guidance for standard, military, highway, rural-route, and attention-line formats. It does not combine name, phone, email, Social Security number, and date-of-birth extraction in the same prompt. This configuration uses the same model as the proposed system but does not use planning, task decomposition, specialized workers, semantic-review workers, formatting workers, or the exploratory retry pathway.
- Multi-Agent System. The proposed architecture using the Planner Agent, Manager Agent, specialized Worker Agents, same-model semantic consistency review, formatting, and an exploratory bounded retry pathway whose reliability was not evaluated. Only its address-component outputs are scored quantitatively.
The three configurations were evaluated on the same 700-record address dataset, comprising seven address categories with 100 records per category, as summarized in Table 11.
Table 11.
Evaluation dataset: seven address categories with 100 records each.
The monolithic reference and hierarchical pipeline use the same underlying model but differ in prompt scope, instruction content, schema allocation, semantic-review stages, and orchestration. Table 12 summarizes these design and resource differences, including fixed prompt-token counts, output limits, input segmentation, and features that were not matched or could not be verified. The hierarchical configuration assigns dedicated prompts to individual extraction tasks, whereas the reference configuration uses the address-only scope described above. The prompts were not matched for fixed instruction tokens or total input tokens. Both retained configurations specify the same 4096-token maximum output limit per call. The supplied evaluation scripts execute the monolithic condition on one address record per LLM call and the hierarchical condition on one address record per pipeline run; the hierarchical path still routes that record through its internal chunker. The raw predictions and execution logs underlying the reported aggregate summaries were not retained, so identical reported-run segmentation, original batching, and realized call totals cannot be verified. The manuscript therefore does not claim a verified call-count ratio or compute-efficiency difference. Consequently, the comparison characterizes complete configurations and does not isolate the effect of multi-agent coordination, worker specialization, semantic review, or segmentation.
Table 12.
Prompt and resource accounting for the two LLM configurations. Fixed token counts use the tokenizer associated with the openai/gpt-oss-20b artifact and the retained prompt templates. Chat-framed totals include the fixed user wrapper and an empty dynamic-input placeholder. Run-level dynamic and aggregate token totals were not retained. The conditions were not prompt- or compute-matched.
Scope of Comparison and Component Attribution
The present evaluation was designed to compare the complete hierarchical pipeline with rule-based and monolithic configurations. It was not designed as a factorial ablation study. The hierarchical and monolithic reference conditions differ along several dimensions at the same time, including the presence of a Planner Agent, an LLM-generated task graph, specialized worker prompts, semantic-review and formatting stages, an exploratory retry pathway, and chunk-level execution. Realized call totals and aggregate inference allocation for the reported runs could not be verified.
Therefore, the observed difference between the hierarchical and monolithic reference conditions is an end-to-end configuration difference. It does not establish that planning, worker specialization, same-model semantic review, retry-pathway routing, chunking, or any other individual component independently caused the difference. It also does not determine whether the difference resulted primarily from architectural decomposition, additional task-specific instructions, reduced prompt scope, unverified inference allocation, original batching, or interactions among these factors. Table 13 summarizes the component-level evidence available in the present evaluation and clarifies that no independent effect was established for any individual architectural component.
Table 13.
Scope of component-level evidence in the present evaluation.
Component-level conclusions would require controlled variants in which one factor is changed while chunking, prompt-token allocation, inference budget, model configuration, and scoring procedures are held constant. Such experiments were not included in the present study. The quantitative claims in this paper are consequently restricted to the complete configurations that were evaluated.
5.3. Evaluation Metrics
Four metrics are used to evaluate performance:
- Address Component-Level Exact Match (EM). This is the primary metric. It measures the proportion of address component fields that exactly match the corresponding generator-derived reference value. The score is micro-averaged across records and address fields. Missing predictions are counted as incorrect for non-empty reference fields. Empty fields are counted as correct only when the corresponding reference field is also empty.
- Record-Level Recall. This measures the proportion of input records for which the system produces any output. A system that skips records receives a lower recall score even if it performs well on the records it processes.
- Structural Completeness. This measures the average proportion of applicable non-empty generator-derived reference fields that receive non-empty predicted values. This metric reduces inflation from fields that are empty by design.
- Pipeline Completion Rate. This measures the proportion of Worker Agents that complete execution without unrecoverable failure. This metric applies only to the multi-agent system.
Each LLM configuration was executed three times on the same fixed 700-record synthetic benchmark. These executions provide three repeated realizations of model output for each configuration; they do not constitute three independently sampled datasets. Arithmetic means and observed sample standard deviations are reported solely as descriptive summaries of these executions. With only three executions, the study does not support reliable confidence intervals, formal hypothesis testing, or statistically supported rankings among address categories. Accordingly, all category-level values and differences reported below are interpreted descriptively only. The maximum observed across-execution standard deviation was 1.2 percentage points, but this value should not be interpreted as evidence of statistical stability. The rule-based parser is deterministic for the fixed benchmark and is therefore represented by a single result.
5.4. Main Results
Interpretation of reported percentages.Unless otherwise stated, all percentages in this section refer exclusively to the 700-record author-generated synthetic benchmark. They are not estimates of expected performance on naturally occurring census records, postal records, OCR-derived documents, handwritten submissions, multilingual administrative data, or production workloads.
Table 14 reports descriptive address-component exact match by generated category for the rule-based parser, the address-only monolithic reference, and the complete hierarchical configuration.
Table 14.
Descriptive address-component exact-match results (%) on the fixed author-generated synthetic benchmark. LLM values are arithmetic means across three repeated executions on the same records. No confidence intervals or significance tests are reported, and category-level differences should not be interpreted as statistically established. The values measure agreement with generator-derived reference fields, and the two LLM configurations were not compute- or prompt-budget matched.
Table 14 presents the observed category profile of the three evaluated configurations. Across the three repeated LLM executions, the complete hierarchical configuration had numerically higher mean exact-match values than the monolithic reference in each of the seven generated categories. These values are descriptive outcomes on the fixed synthetic benchmark. Because only three executions were conducted, the study does not establish that the difference was statistically larger for one address category than for another, and differences such as 99% versus 100% are not interpreted as meaningful. In addition, the fixed prompt allocations differed, while realized inference allocation and original batching could not be verified from the retained artifacts. The numerical differences are therefore not attributed to a specific architectural component.
5.5. Overall Pipeline Metrics and Runtime Trade-Offs
Table 15 presents aggregate pipeline metrics across all 700 records.
Table 15.
Overall pipeline metrics. LLM values are arithmetic means across three repeated executions on the same fixed benchmark; the deterministic rule-based values are reported once.
The retained aggregate summary reports 700 usable outputs for the hierarchical configuration, 692 for the rule-based parser, and 653 for the monolithic reference. The record-level prediction files needed to reconstruct those counts were not retained with the supplied repository. The cause of the reported omissions was not isolated, and the study did not evaluate whether omissions depended on record position within a prompt. Potential contributors include generation omissions, response-length truncation, malformed structured output, parsing or record-alignment failures, input segmentation, and other model-output failures. None is presented as the established cause, and the reported coverage values do not estimate omission rates in operational workflows.
The complete hierarchical configuration also recorded the highest benchmark structural-completeness score, 0.94, compared with 0.78 for the monolithic reference and 0.58 for the rule-based parser. Descriptively, it filled a larger proportion of applicable non-empty generator-derived reference fields. The difference between 95.7% exact match and 0.94 structural completeness reflects cases where the system produced a value for a field but the value did not exactly match the benchmark reference.
The aggregate evaluation summary reports wall-clock ranges of 20 to 25 min for the complete hierarchical configuration and 3 to 4 min for the monolithic reference on the local GPU; the rule-based parser was reported to complete in less than 5 s. Run-level timing logs and realized LLM-call totals were not retained, and the per-record execution structure in the supplied repository does not reproduce the call totals stated in an earlier manuscript draft. Accordingly, the runtime ranges are treated as approximate reported metadata, no call-count ratio is reported, and no compute-efficiency conclusion is drawn. Operational latency requires evaluation with retained timing logs and representative workloads.
5.6. Illustrative Retry Traces
Across the three repeated executions, the logging system recorded three instances in which the bounded retry pathway was invoked. The cases involved generated Military, Attention-Line, and Highway records. The corresponding first-attempt and post-retry outputs are documented in Appendix B for transparency.
These three observations do not constitute a reliability evaluation. They were not sampled from a predefined stress set, and the study did not include an otherwise identical feedback-disabled condition. Consequently, the observations cannot be used to estimate the probability of successful recovery, the frequency of false or unnecessary triggers, the likelihood of harmful revisions, maximum-retry exhaustion, or the mechanism’s contribution to overall extraction accuracy.
In each logged case, the post-retry output satisfied the benchmark’s field-level criterion. This establishes only what occurred in those particular executions. It does not establish that re-planning caused the changed output, because an additional stochastic generation without revised planning might also have produced a different result. The three cases are therefore treated as illustrative execution traces rather than evidence of general robustness or an answer to a formal research question.
Reliability evaluation would require a pre-specified set of failure-inducing inputs and a matched comparison between feedback-enabled and feedback-disabled configurations. Such an evaluation was not conducted in the present study.
5.7. Qualitative Error Analysis
Table 16 summarizes the main error patterns observed during evaluation. This analysis complements the quantitative results by explaining why performance differs across systems.
Table 16.
Qualitative error patterns observed during evaluation.
The configurations differ substantially in task scope and prompt allocation. The address worker receives dedicated address instructions, while the address-only reference condition uses a different instructional structure. These differences provide plausible explanations for the observed error patterns, but the present evaluation does not determine whether the numerical difference resulted from narrower task scope, more detailed guidance, additional inference, same-model semantic review, orchestration, or their interaction. The bounded-cognition principle remains a design rationale; it was not independently validated by the present comparison.
Confidence scores also reflected category difficulty at a broad level. Standard and individual address records received higher mean confidence scores, while military, highway, and attention-line records received lower confidence scores. These scores should not be treated as proof of correctness for individual records [28], but they may be useful as a prioritization signal for validation or future human review.
5.8. Relationship to Prior Census Parsing Systems
Two previously published systems provide relevant design context, but their reported performance cannot be compared numerically with the present pilot. The studies used different datasets, evaluation metrics, models, supervision assumptions, and deployment environments.
The pattern-based active-learning system illustrates an approach in which expert corrections are converted into persistent mappings [5]. Its principal design distinction is the accumulation of reusable human knowledge when unfamiliar patterns are encountered. The present prototype does not provide equivalent persistent learning.
The cloud-hosted LLM system illustrates prompt-driven extraction using a commercial model and an external inference environment [11]. Its principal distinctions include the model and serving environment, its validation workflow, and the corpus on which it was evaluated.
The present study instead examines a locally deployed open-weight implementation evaluated on an author-generated synthetic benchmark. These systems therefore represent different design and deployment choices rather than entries in a common performance ranking. No conclusion about their relative accuracy follows from the results reported in their separate evaluations.
5.9. Threats to Validity
A principal threat is synthetic-benchmark and shared-source bias. The textual records and their reference labels originate from the same structured generation process. This produces complete and unambiguous labels relative to the generator, but it also couples the input distribution, field taxonomy, component definitions, and expected outputs. The retained repository contains no formal QA report or annotation log, and the sample, reviewer count, checklist, discrepancies, and correction history for the informal author consistency review were not recorded. That review does not remove the shared-source dependency and should not be interpreted as independent annotation.
The benchmark does not adequately represent the uncontrolled errors and variation found in naturally occurring records, including misspellings, OCR and transcription artifacts, handwriting errors, multilingual content, incomplete responses, duplicated or contradictory fields, unexpected abbreviations, uncontrolled field ordering, and formats outside the seven predefined categories. The reported 95.7% value is therefore a score on the generator-induced benchmark distribution, not an estimate of accuracy on real census submissions.
No conclusion about production performance, real-record robustness, demographic fairness, multilingual generalization, or transfer to other administrative domains follows from this evaluation. Establishing those properties requires independently collected or independently annotated records, documented annotation guidelines, multiple annotators where appropriate, and analysis of disagreement and annotation uncertainty.
A second threat is model coverage. Only gpt-oss-20b was tested. The results may not generalize to other open-weight models, larger local models, or different quantization schemes. A stronger evaluation would compare multiple open-weight models under the same architecture.
A third threat is limited repeated-run evidence. Each LLM condition was executed three times on the same fixed benchmark. These repetitions provide a limited description of execution-to-execution variability, but they do not represent independently sampled datasets and do not increase the number of unique records beyond 100 per category. With only three executions, estimates of across-run variance are unstable, and the study does not support confidence intervals, formal significance testing, or statistically reliable category rankings. Consequently, all per-category values and between-category differences are treated as descriptive. Future evaluation should use a larger number of repeated executions and a pre-specified paired or hierarchical uncertainty analysis.
A fourth threat is uncharacterized output omissions. The study did not evaluate whether reported baseline omissions depended on record position within a prompt. It also did not isolate output truncation, malformed structured output, parser or alignment failure, generation omission, input segmentation, or other possible causes. The original raw prompts, responses, and record-position mappings are unavailable, so a retrospective positional analysis could not be performed. References to long-context degradation are therefore treated as design motivation rather than an explanation established by this evaluation.
A fifth threat is resource-accounting uncertainty and bundled treatment variation. The supplied evaluation repository executes one record per monolithic call and one record per hierarchical pipeline run, but it does not contain the raw run artifacts underlying the aggregate results. It therefore does not reproduce the call totals stated in an earlier draft. Realized call counts, dynamic input tokens, aggregate prompt allocation, generation totals, and the original batching cannot be verified. The configurations also differ in planning, task coordination, fixed prompt allocation, same-model semantic review, formatting, and the presence of an exploratory retry pathway. The present study cannot separate the contribution of specialization, any other architectural component, unverified inference allocation, or segmentation. Future work should retain complete prompts, outputs, timings, and call logs and include factorial ablations and compute-, prompt-, and chunk-matched comparisons.
A final threat is uncharacterized retry-pathway behavior. The bounded retry pathway was invoked in only three naturally occurring cases, and no pre-specified stress corpus or feedback-disabled control was evaluated. The study therefore cannot estimate trigger sensitivity, trigger precision, recovery probability, harmful-retry frequency, maximum-retry exhaustion, or causal contribution to extraction accuracy. The three cases are reported only as illustrative execution traces. Evaluation of this mechanism requires a controlled stress study with matched inference budgets and independently scored pre- and post-retry outputs.
6. Discussion
This section interprets the controlled pilot, identifies limitations, and discusses the design properties and evaluation needed before operational conclusions can be made. All reported performance comparisons concern the complete evaluated configurations on the author-generated synthetic benchmark. Because the LLM conditions were not matched for prompt and inference resources and no component-level ablations were conducted, none of the observed differences is interpreted as evidence that hierarchical organization, task decomposition, worker specialization, planning, semantic review, or retry processing improved accuracy. The findings also document local execution of the pipeline but do not establish real-record performance.
6.1. Interpretation of the Research Questions
For RQ1, across the three repeated executions on the controlled author-generated synthetic benchmark, the hierarchical configuration obtained a mean of 95.7% micro-averaged address-component exact match, 100% output coverage, and 0.94 structural completeness. These values describe agreement with generator-derived fields under the benchmark’s predefined categories and rendering conventions. Category-level values are descriptive and are not used to rank category difficulty. The results demonstrate execution feasibility and benchmark behavior of the complete Planner–Manager–Worker configuration, not accuracy or robustness on naturally occurring records.
For RQ2, the aggregate summaries for the three directly evaluated configurations differed descriptively in address-component exact match, reported record coverage, structural completeness, and reported wall-clock time. Under the evaluated, resource-unmatched conditions, the observed overall mean for the complete hierarchical configuration was 14.8 percentage points higher than that of the monolithic reference on the fixed synthetic benchmark; the complete hierarchical configuration also recorded the longest reported runtime range, whereas the deterministic rule-based reference was fastest. This is an end-to-end descriptive difference between bundled configurations, not an estimate of an architectural or component-level effect. Category-level numerical differences are reported in Table 14, but the three-execution design does not support statistically reliable category rankings or conclusions that the difference was greater for particular address types. The retained fixed prompts differed in scope and instruction allocation, while realized call totals, dynamic token allocation, original batching, reported-run segmentation, and run-level timing logs could not be verified. The comparison does not establish a model-call or compute-efficiency ranking and does not attribute the numerical difference to worker specialization, planning, coordination, semantic review, or another individual component. A prompt- and compute-matched sequential control with complete run logging would be required for such attribution.
For RQ3, observed errors and operational trade-offs further limit interpretation. The same-model semantic consistency review is not an independent source of truth and may share upstream error modes. The three logged retry activations are illustrative execution traces, not evidence of a reliable recovery mechanism, because no feedback-disabled control or stress-test corpus was included. The reported latency, unverified resource allocation, synthetic reference labels, and absence of deterministic or human validation controls identify concrete requirements for future evaluation; they do not support claims of production readiness or real-record robustness.
The prototype’s retry pathway was invoked only three times. These cases are documented as execution traces but are not treated as evidence of reliability or as an evaluated source of the pipeline’s accuracy. No feedback-disabled control or stress-test corpus was included.
6.2. Trade-Offs Among the Directly Evaluated Configurations
Table 17 summarizes only measurements obtained for the three configurations evaluated on the same author-generated synthetic benchmark. Results from published systems are excluded because their datasets, metrics, models, and evaluation procedures differ.
Table 17.
Descriptive comparison of the three configurations evaluated on the same author-generated synthetic benchmark. Values should not be compared with results from external studies using different datasets or metrics.
These values are descriptive summaries from the common pilot benchmark. The fixed prompts were not matched, and the realized inference allocation, original batching, and run-level timing logs could not be verified from the retained artifacts. The table therefore does not establish a compute-matched comparison or attribute the differences to any individual component.
Deployment environment, human intervention, persistent learning, and validation workflows for prior systems are described separately from the literature in Section 5.8. Those literature-based properties are not combined with the present measurements or used to rank the approaches. Whether any configuration is useful in practice requires independent evaluation on representative real records.
6.3. Architectural Design and Limits of Component Attribution
The study documents the implementation and execution of the bundled hierarchical pipeline on local workstation infrastructure and describes its behavior on the controlled synthetic benchmark. Because the evaluated conditions differed simultaneously in prompt allocation, planning, coordination, task-specific worker prompts, same-model semantic review, formatting, chunking, retry processing, and unverified inference allocation, the results do not show that hierarchy, decomposition, specialization, or any other included component improved performance.
Separately from performance attribution, the implementation has directly observable engineering characteristics. Components have separate responsibilities and structured interfaces; worker prompts can be modified independently; execution dependencies are explicitly represented; failures and retries can be logged; and additional extraction types can be introduced through new worker interfaces. These characteristics describe the organization of the implemented system only and should not be interpreted as experimentally established accuracy, robustness, or efficiency benefits.
Determining the performance contribution of each component requires a controlled ablation design. In particular, worker specialization should be compared with a merged extraction worker under matched prompt-token and inference budgets. The retry pathway requires a separate stress evaluation with otherwise identical enabled and disabled configurations under matched inference budgets. Planner, Manager, and chunking ablations would provide additional information about component interactions. Until such experiments are conducted, the performance results should be interpreted only at the level of the complete configuration.
6.4. Limitations and Failure Modes
The main limitation is shared-source synthetic benchmark construction. The rendered inputs and reference fields originate from the same generator and share its taxonomy, formatting assumptions, and component definitions. The reported 95.7% value is a benchmark score under that induced distribution, not an estimate of performance on real census submissions. Author consistency review did not constitute independent annotation. Moreover, the scope, sample, reviewer count, checklist, and correction history of the informal review were not retained as a formal QA record.
The quantitative evaluation is also limited to fine-grained address components. Although the prototype implements workers for names, phone numbers, email addresses, Social Security numbers, and dates of birth, this study provides no task-level accuracy, robustness, or generalization evidence for those workers. The reported benchmark values must not be interpreted as broad PII-extraction performance.
A second limitation is that only one open-weight model was tested. The findings may not generalize to other local models, larger open-weight models, or different quantization strategies. Future work should compare multiple open-weight models under the same architecture to determine how much performance depends on the model versus the agent design.
A third limitation is the small number of repeated executions. Three executions on the same fixed records provide only a limited descriptive view of model variability; they do not support confidence intervals, formal tests, or category rankings.
A fourth limitation is the uncharacterized retry pathway. Three logged activations are insufficient to estimate trigger or recovery behavior, and no feedback-disabled control was evaluated. The pathway is therefore an exploratory feature rather than a validated advantage.
A fifth limitation is runtime and untested scalability. The current prototype executes tasks sequentially and has a reported runtime of 20 to 25 min for 700 records. It is slower than both the rule-based parser and monolithic reference, and the present benchmark does not establish whether this latency is operationally acceptable. Distributed scheduling, dynamic batching, multiple model replicas, lower-memory GPUs, CPU-only operation, and high-availability serving were not evaluated. The capacity values in Section 6.5.2 are first-order analytical projections, not measured distributed throughput.
A sixth limitation is open-set and subgroup coverage. The seven U.S.-oriented address categories are not exhaustive, and no out-of-taxonomy evaluation, abstention evaluation, international-address test, multilingual test, or subgroup fairness audit was conducted. The current first/middle/last name representation and state/ZIP/APO/FPO/DPO address fields reflect U.S.-oriented and partly Western structural assumptions. The study therefore provides no evidence of fairness across cultural naming conventions, scripts, languages, countries, or regional postal systems.
Finally, LLM-based systems introduce model-inherent risks. They may hallucinate plausible but incorrect field values, share error modes between extraction and semantic review, or perform unevenly across cultural naming conventions [14,48]. In the evaluated prototype, JSON recovery and fallback handling provide only partial syntactic protection; typed schema enforcement, deterministic address-format checks, independent postal-reference checks, and human escalation were not implemented as evaluated controls. Production deployment should therefore add these safeguards, along with secure logging and human review for low-confidence cases.
6.5. Deployment Guidance
The pipeline was designed for scenarios in which records must remain locally controlled, extraction schemas are heterogeneous, and batch processing is possible. The present evaluation does not establish that it is appropriate for any operational setting; such a determination requires representative data, security review, latency requirements, and comparison with alternative systems.
The prototype’s reported 20-to-25-min runtime range on the synthetic pilot indicates that optimization may be needed for low-latency use, but the original timing logs were not retained for verification. No recommendation among rule-based parsers, commercial LLMs, active-learning systems, or the present pipeline follows from this benchmark alone.
A hybrid workflow that routes low-confidence or validation-failed records to human review is a future design option. Its accuracy, workload, and governance effects have not been evaluated here.
6.5.1. Open-Set and Out-of-Taxonomy Records
The seven address categories used in the evaluation define the benchmark taxonomy; they are not an exhaustive address ontology. The current study does not measure performance on formats outside that taxonomy. Because the address worker is organized by PII type rather than by one address category, an unfamiliar record could still be routed to it, but reliable open-set generalization has not been established.
For production extension, the Planner should be permitted to assign unknown, mixed, or outside_taxonomy rather than forcing every record into a known category. The address pathway should preserve the complete source string, extract only components supported by the source, retain unparsed content, avoid fabricating missing values, and expose uncertainty. Records with an unknown format, unsupported components, unresolved schema conflicts, failed deterministic checks, or low confidence should be abstained from or routed to human review rather than silently accepted.
A recurring new arrangement of familiar fields may require versioned Planner or address-worker guidance and new deterministic rules. A format containing semantic components that the current schema cannot represent requires versioned schema and downstream-mapping changes as well. A new Worker Agent is not necessarily required because the existing worker is specialized for address extraction rather than for a single evaluated category. This policy is a proposed production design and was not implemented or evaluated in the present pilot.
6.5.2. Scalability and Capacity Planning
The reported 20-to-25-min runtime reflects a sequential prototype using one model replica on one 24 GB GPU. Independent extraction tasks, document chunks, independent input files, and separate dataset partitions provide potential parallelism. They could be scheduled concurrently or assigned to multiple local model-server replicas, and dynamic request batching may improve accelerator utilization. Same-model semantic review and final formatting remain synchronization points because they depend on preceding extraction outputs. None of these parallel or distributed configurations was evaluated.
A first-order capacity estimate can be derived from the reported runtime range:
where N is the number of records and G is the number of equivalent model replicas. This projection assumes workload characteristics identical to the pilot and ideal linear scaling. It excludes request scheduling, synchronization, storage, network, model-server saturation, batching effects, failures, retries, and recovery overhead. It is not a measured service-level guarantee. Table 18 summarizes the resulting first-order capacity estimates for representative workload and replica-count scenarios.
Table 18.
Analytical capacity-planning estimates derived from the reported 20-to-25-min runtime for 700 records. Distributed scaling was not experimentally evaluated.
Production benchmarking must measure throughput, tail latency, GPU utilization, memory pressure, queue depth, batching efficiency, synchronization cost, failure recovery, availability, and cost under the intended workload. Model inference remains a material computational cost even when orchestration is parallelized.
6.6. Bias, Cultural Naming, and Regional Address Generalization
No subgroup fairness analysis was conducted. The synthetic benchmark is U.S.-oriented: it uses U.S. states and ZIP codes, U.S. military address conventions, and English-language Faker data. The implemented first/middle/last name representation also reflects a Western structural assumption and was not quantitatively evaluated. The present study therefore supports no conclusion about fairness or performance for multilingual names, non-Latin scripts, international address systems, or regional formats outside the benchmark.
A future audit should use legally usable and independently annotated evaluation sets stratified by documented language, script, region, and naming or address structure. Name strata should include mononyms, family-name-first ordering, patronymic or matronymic systems, multiple surnames, particles, compound names, diacritics, non-Latin scripts, and transliterated forms. Address strata should include country-specific administrative hierarchies, locality-first ordering, non-U.S. postal codes, addresses without street numbers, building- or landmark-based systems, rural locality descriptions, and multilingual or mixed-script records.
Evaluation should report component exact match, omission rate, over-normalization rate, source-string preservation, abstention rate, confidence calibration, subgroup sample sizes, worst-group performance, and uncertainty. Group membership must come from documented data provenance, consented metadata, or controlled benchmark construction; no protected characteristic should be inferred from a person’s name.
Planned mitigations include preserving Unicode and original source strings, avoiding forced assignment to inapplicable name fields, supporting ordered or locale-specific name components, using jurisdiction-specific address schemas and validators, abstaining on unsupported structures, and revising guidance based on independently annotated error audits. These are future evaluation and design commitments, not evidence that the present system is unbiased. Table 19 summarizes the principal risks, proposed evaluation measures, and planned mitigations for future bias and regional-generalization audits.
Table 19.
Future bias and regional-generalization audit plan. These measures and mitigations were not evaluated in the present study.
6.7. Broader Implications
The architectural pattern has not been evaluated outside the controlled synthetic address benchmark. Adaptation to healthcare, finance, legal, insurance, or other administrative extraction remains a design hypothesis requiring domain-specific data, schemas, validation rules, and independent evaluation.
7. Conclusions and Future Work
This paper specified and implemented a hierarchical Planner–Manager–Worker pipeline for fine-grained census-style address parsing using a locally deployed open-weight language model. The complete address-parsing configuration includes specialized extraction, same-model semantic consistency review, formatting, record-preserving chunking, and execution logging. The term “specialized” describes the assignment of task-specific responsibilities in the implementation; it does not denote an experimentally established performance benefit. Other PII workers are present in the prototype but were not quantitatively evaluated.
In a controlled pilot using 700 author-generated synthetic records, across the three observed executions the complete configuration obtained a mean 95.7% micro-averaged address-component exact-match score. The monolithic LLM reference obtained a mean of 80.9%, and the deterministic rule-based reference obtained 52.3%. The category-level results provide a descriptive profile of the evaluated configurations on the fixed synthetic benchmark. Because only three repeated executions were conducted, the study does not establish statistically meaningful differences among address categories or identify categories for which the performance difference is reliably larger. These figures describe the evaluated generator-induced benchmark only. Because the textual records and reference labels originated from the same synthetic generation process, the evaluation does not estimate performance on naturally occurring census records. It also does not establish robustness to OCR artifacts, misspellings, multilingual content, missing fields, uncontrolled formatting, or formats outside the predefined categories. The LLM configurations differed simultaneously in prompt specialization, fixed prompt-token allocation, same-model semantic review, orchestration, and unverified inference allocation. Accordingly, their numerical difference is an observation about two bundled, resource-unmatched configurations and is not evidence that worker specialization, task decomposition, hierarchical orchestration, or any other individual component improved accuracy.
The prototype includes a bounded retry pathway for routing selected execution failures back to the Planner. Because the pathway was invoked only three times and no feedback-disabled comparison was conducted, the present study does not establish its reliability or contribution to overall extraction performance. The mechanism should therefore be regarded as an exploratory design feature requiring dedicated stress evaluation. Future work should use a pre-specified corpus containing malformed-output triggers, OCR substitutions and deletions, missing or contradictory fields, unseen address structures, unexpected line breaks, and prompt-like text embedded in records. Feedback-enabled and feedback-disabled configurations should use matched model, prompt, chunk, seed, and inference budgets and report trigger rate, false-trigger rate, post-retry exact match, recovery relative to the no-feedback condition, harmful-retry rate, retry-bound exhaustion, added token cost, and added latency.
The tested 24 GB workstation is one evaluated reference configuration, not a minimum specification or a production-scale deployment. The analytical capacity projection derived from the reported runtime suggests approximately 20–25 equivalent replicas for one million records in 24 h only under ideal linear scaling; this was not measured. Future deployment studies should benchmark multi-replica scheduling, dynamic batching, synchronization, tail latency, GPU utilization, failure recovery, availability, and cost. Lower-memory GPUs, CPU-only operation, and alternative quantization levels also require direct testing.
Future external-validity work should include explicit open-set evaluation and independently annotated cultural and regional audits. Unknown address formats should be evaluated for conservative extraction, raw-string preservation, abstention, false confidence, and review workload. Cultural and regional studies should use documented provenance rather than inferring protected characteristics from names, and should report performance for diverse naming structures, scripts, languages, and postal systems together with worst-group results and uncertainty. Schema, prompt, deterministic-rule, and downstream-mapping changes should be versioned when recurring novel formats introduce new semantic components.
The principal contribution is therefore the design, implementation, and controlled characterization of an integrated on-premise address-parsing pipeline, rather than a claim of validated real-world accuracy, broad PII-extraction performance, architectural causality, or production readiness. All reported performance differences apply only to the complete evaluated configurations and do not establish a causal benefit from hierarchy, decomposition, worker specialization, semantic review, retry processing, or any other individual component. The same-model semantic review does not replace independent validation. Evaluation using independently annotated real or census-derived records, documented annotation guidelines, multiple annotators where appropriate, and analysis of disagreement remains necessary before operational or cross-domain conclusions can be drawn.
Author Contributions
Conceptualization, A.T.; Methodology and study design, M.A.M.; Software, A.T.; Validation, A.T., M.A.M., S.A.M. and M.C.C.; Formal analysis, A.T. and M.C.C.; Investigation, A.T., M.A.M. and S.A.M.; Data curation, A.T.; Writing—original draft preparation, A.T.; Writing—review and editing, M.A.M., S.A.M. and M.C.C.; Visualization, A.T.; Supervision and mentorship, J.R.T.; Project guidance and academic mentoring, J.R.T. All authors have read and agreed to the published version of the manuscript.
Funding
This research was partially supported by the National Science Foundation under EPSCoR Award No. OIA-1946391.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The implementation and reproducibility materials are publicly available at https://github.com/MohdMuzakkiruddinAhmed/census-pii-agents (accessed on 22 July 2026). The repository includes the agent prompts, Planner–Manager–Worker and orchestration code, synthetic-data generation script, evaluation scripts and baselines, environment configuration, and model setup instructions. The 700-record synthetic benchmark can be regenerated using the documented random seed 42; no real census records or personal information are included. The evaluated base model identifier is openai/gpt-oss-20b, and execution requires a compatible local model deployment.
Acknowledgments
The authors gratefully acknowledge computational resources and research infrastructure provided by the Center for Advanced Research in Entity Resolution and Information Quality (ERIQ), University of Arkansas at Little Rock. This research was conducted under the ERIQ Center’s ongoing cooperative research program in census data quality.
Conflicts of Interest
The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.
Appendix A. Sample Evaluation Records
Table A1 presents representative records from the synthetic evaluation dataset. All data is entirely fictitious.
Table A1.
Representative evaluation records (one per category).
The implementation and reproducibility materials are publicly available at https://github.com/MohdMuzakkiruddinAhmed/census-pii-agents (accessed on 22 July 2026). The repository provides the fixed-seed generator and documents how to regenerate the complete 700-record synthetic benchmark used in the evaluation.
Appendix B. Illustrative Retry-Pathway Traces
Table A2 documents the three logged retry-pathway activations for transparency. The cases were not drawn from a pre-specified stress corpus, are not a statistical sample, and are not used to estimate a recovery rate or causal benefit.
Acceptable post-retry outputs were observed in these three logged cases. This statement describes the traces only; it does not establish pathway reliability, a three-out-of-three recovery rate, or a contribution to the pipeline’s overall benchmark score.
Table A2.
Three illustrative retry-pathway traces observed during the controlled pilot. The cases are not a statistical sample and are not used to estimate a recovery rate.
References
- U.S. Census Bureau. 2020 Census Results. 2021. Available online: https://www.census.gov/2020census (accessed on 15 January 2025).
- Christen, P. Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection; Springer: Berlin/Heidelberg, Germany, 2012. [Google Scholar] [CrossRef] [Scilit]
- Fellegi, I.P.; Sunter, A.B. A Theory for Record Linkage. J. Am. Stat. Assoc. 1969, 64, 1183–1210. [Google Scholar] [CrossRef]
- Goldberg, D.W.; Wilson, J.P.; Knoblock, C.A. From Text to Geographic Coordinates: The Current State of Geocoding. URISA J. 2007, 19, 33–46. [Google Scholar]
- Mohammed, O.K.; Syed, K.; Talburt, J.; Tarannum, A.; Kashif, A.K.K.; Khan, S.; Syed, N.; Mehdi, S.Y. A Pattern-Based Approach to Name and Address Parsing with Active Learning. In Proceedings of the 17th International Conference on Agents and Artificial Intelligence (ICAART 2025), Porto, Portugal, 23–25 February 2025; SCITEPRESS: Set’ubal, Portugal, 2025; Volume 3, pp. 70–77. [Google Scholar] [CrossRef] [Scilit]
- Borkar, V.; Deshmukh, K.; Sarawagi, S. Automatic Segmentation of Text into Structured Records. In Proceedings of the 2001 ACM SIGMOD International Conference on Management of Data, Santa Barbara, CA, USA, 21–24 May 2001; ACM: New York, NY, USA, 2001; pp. 175–186. [Google Scholar] [CrossRef] [Scilit]
- Lafferty, J.; McCallum, A.; Pereira, F.C.N. Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data. In Proceedings of the 18th International Conference on Machine Learning (ICML 2001), Williamstown, MA, USA, 28 June–1 July 2001; Morgan Kaufmann: San Francisco, CA, USA, 2001; pp. 282–289. [Google Scholar]
- Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2019), Minneapolis, MN, USA, 2–7 June 2019; ACL: Stroudsburg, PA, USA, 2019; Volume 1, pp. 4171–4186. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Sun, A.; Han, J.; Li, C. A Survey on Deep Learning for Named Entity Recognition. IEEE Trans. Knowl. Data Eng. 2020, 34, 50–70. [Google Scholar] [CrossRef] [Scilit]
- Brown, T.B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language Models Are Few-Shot Learners. In Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Virtual, 6–12 December 2020; Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M.F., Lin, H., Eds.; Curran Associates: Red Hook, NY, USA, 2020; Volume 33, pp. 1877–1901. [Google Scholar]
- Tarannum, A.; Mohammed, M.A.; Cakmak, M.C.; Al Mandalawi, S.; Talburt, J. A System for Name and Address Parsing with Large Language Models. arXiv 2026, arXiv:2601.18014. [Google Scholar]
- Liu, N.F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; Liang, P. Lost in the Middle: How Language Models Use Long Contexts. Trans. Assoc. Comput. Linguist. 2024, 12, 157–173. [Google Scholar] [CrossRef] [Scilit]
- OpenAI. Introducing GPT-OSS-20b: Open-Weight Mixture-of-Experts Model. 2025. Available online: https://openai.com/index/introducing-gpt-oss/ (accessed on 1 June 2025).
- Borgman, C.L.; Siegfried, S.L. Getty’s Synoname and Its Cousins: A Survey of Applications of Personal Name-Matching Algorithms. J. Am. Soc. Inf. Sci. 1992, 43, 459–476. [Google Scholar]
- Al-Rfou, R. libpostal: A Library for Parsing/Normalizing Street Addresses Around the World. 2016. Available online: https://github.com/openvenues/libpostal (accessed on 15 January 2025).
- Cunningham, H.; Maynard, D.; Bontcheva, K.; Tablan, V. GATE: A Framework and Graphical Development Environment for Robust NLP Tools and Applications. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL 2002), Philadelphia, PA, USA, 7–12 July 2002; ACL: Stroudsburg, PA, USA, 2002; pp. 168–175. [Google Scholar] [CrossRef] [Scilit]
- Cucerzan, S.; Yarowsky, D. Language Independent Named Entity Recognition Combining Morphological and Contextual Evidence. In Proceedings of the 1999 Joint SIGDAT Conference on Empirical Methods in Natural Language Processing and Very Large Corpora (EMNLP-VLC), College Park, MD, USA, 21–22 June 1999; ACL: Stroudsburg, PA, USA, 1999; pp. 90–99. [Google Scholar]
- Comber, S.; Zeng, R. International Address Parsing Using Conditional Random Fields. GeoInformatica 2019, 23, 431–452. [Google Scholar]
- Sharma, H.; Jadon, R.S.; Shukla, B.K. Address Parsing Using Deep Learning Approach. In Proceedings of the 2018 International Conference on Computing, Power and Communication Technologies (GUCON 2018), Greater Noida, India, 28–29 September 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 545–549. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. In Proceedings of the 31st Conference on Neural Information Processing Systems (NeurIPS 2017), Long Beach, CA, USA, 4–9 December 2017; Curran Associates: Red Hook, NY, USA, 2017; Volume 30, pp. 5998–6008. [Google Scholar]
- Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Proceedings of the 36th Conference on Neural Information Processing Systems (NeurIPS 2022), New Orleans, LA, USA, 28 November–9 December 2022; Curran Associates: Red Hook, NY, USA, 2022; Volume 35, pp. 24824–24837. [Google Scholar]
- Willard, B.T.; Louf, R. Efficient Guided Generation for Large Language Models. arXiv 2023, arXiv:2307.09702. [Google Scholar]
- Xu, D.; Chen, W.; Peng, W.; Zhang, C.; Xu, T.; Zhao, X.; Wu, X.; Zheng, Y.; Wang, Y.; Chen, E. Large Language Models for Generative Information Extraction: A Survey. arXiv 2023, arXiv:2312.17617. [Google Scholar]
- Peng, B.; Galley, M.; He, P.; Cheng, H.; Xie, Y.; Hu, Y.; Huang, Q.; Liden, L.; Yu, Z.; Chen, W.; et al. Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback. arXiv 2023, arXiv:2302.12813. [Google Scholar]
- Khattab, O.; Singhvi, A.; Maheshwari, P.; Zhang, Z.; Santhanam, K.; Vardhamanan, S.; Haq, S.; Sharma, A.; Joshi, T.T.; Moazam, H.; et al. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines. arXiv 2023, arXiv:2310.03714. [Google Scholar]
- Li, R.; Peng, H.; Li, J. CodeIE: Large Code Generation Models Are Better Few-Shot Information Extractors. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023), Toronto, ON, Canada, 9–14 July 2023; ACL: Stroudsburg, PA, USA, 2023; pp. 8396–8411. [Google Scholar] [CrossRef] [Scilit]
- Lu, Y.; Bartolo, M.; Moore, A.; Riedel, S.; Stenetorp, P. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL 2022), Dublin, Ireland, 22–27 May 2022; ACL: Stroudsburg, PA, USA, 2022; pp. 8086–8098. [Google Scholar] [CrossRef] [Scilit]
- Kadavath, S.; Conerly, T.; Askell, A.; Henighan, T.; Drain, D.; Perez, E.; Schiefer, N.; Hatfield-Dodds, Z.; DasSarma, N.; Tran-Johnson, E.; et al. Language Models (Mostly) Know What They Know. arXiv 2022, arXiv:2207.05221. [Google Scholar]
- Wooldridge, M.; Jennings, N.R. Intelligent Agents: Theory and Practice. Knowl. Eng. Rev. 1995, 10, 115–152. [Google Scholar] [CrossRef] [Scilit]
- Russell, S.J.; Norvig, P. Artificial Intelligence: A Modern Approach, 4th ed.; Pearson: Hoboken, NJ, USA, 2021. [Google Scholar]
- Ferber, J. Multi-Agent Systems: An Introduction to Distributed Artificial Intelligence; Addison-Wesley: Boston, MA, USA, 1999. [Google Scholar]
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. ReAct: Synergizing Reasoning and Acting in Language Models. In Proceedings of the 11th International Conference on Learning Representations (ICLR 2023), Kigali, Rwanda, 1–5 May 2023; OpenReview: Kigali, Rwanda, 2023. [Google Scholar]
- Wu, Q.; Bansal, G.; Zhang, J.; Wu, Y.; Li, B.; Zhu, E.; Jiang, L.; Zhang, X.; Zhang, S.; Liu, J.; et al. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv 2023, arXiv:2308.08155. [Google Scholar]
- Moura, J. CrewAI: Framework for Orchestrating Role-Playing AI Agents. 2024. Available online: https://github.com/joaomdmoura/crewAI (accessed on 15 January 2025).
- LangChain. LangGraph: Build Stateful, Multi-Actor Applications with LLMs. 2024. Available online: https://github.com/langchain-ai/langgraph (accessed on 15 January 2025).
- Althaf, A.M.; Mohammed, M.A.; Milanova, M.; Talburt, J.; Cakmak, M.C. Multi-Agent RAG Framework for Entity Resolution: Advancing Beyond Single-LLM Approaches with Specialized Agent Coordination. Computers 2025, 14, 525. [Google Scholar] [CrossRef] [Scilit]
- Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: Language Agents with Verbal Reinforcement Learning. In Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS 2023), New Orleans, LA, USA, 10–16 December 2023; Curran Associates: Red Hook, NY, USA, 2023; Volume 36. [Google Scholar]
- Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. Self-Refine: Iterative Refinement with Self-Feedback. In Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS 2023), New Orleans, LA, USA, 10–16 December 2023; Curran Associates: Red Hook, NY, USA, 2023; Volume 36. [Google Scholar]
- Mohammed, M.A.; Talburt, J.R.; Mohammed, A.; Syed, K. Entity Resolution with Household Movement Discovery Using Google Generative AI. In Proceedings of the International Conference on Information Technology—New Generations (ITNG 2025), Las Vegas, NV, USA, 14–16 April 2025; Springer: Cham, Switzerland, 2025; pp. 469–481. [Google Scholar]
- Mohammed, M.A.; Talburt, J.R.; Althaf, A.M.; Milanova, M. Multi-LLM Record Linkage: A Comparative Analysis Framework for Co-Residence Pattern Discovery. In Proceedings of the 12th Annual Conference on Computational Science and Computational Intelligence (CSCI 2025), Las Vegas, NV, USA, 4–7 November 2025; IEEE: Piscataway, NJ, USA, 2025. [Google Scholar]
- Mohammed, M.A.; Al Mandalawi, S.; Maclean, H.; Talburt, J.R. Multilingual Customer Record Linkage: A Novel Approach Using LLMs for Cross-Lingual Entity Resolution. In Proceedings of the 12th Annual Conference on Computational Science and Computational Intelligence (CSCI 2025), Las Vegas, NV, USA, 4–7 November 2025; IEEE: Piscataway, NJ, USA, 2025. [Google Scholar]
- Mohammed, M.A.; Talburt, J.R.; Claassens, L.; Marais, A. Retrieval-Augmented Multi-LLM Ensemble for Industrial Part Specification Extraction. In Proceedings of the 17th International Conference on Knowledge and System Engineering (KSE 2025), Ho Chi Minh City, Vietnam, 6–8 November 2025; IEEE: Piscataway, NJ, USA, 2025; pp. 1–6. [Google Scholar]
- Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. LLaMA: Open and Efficient Foundation Language Models. arXiv 2023, arXiv:2302.13971. [Google Scholar]
- Dwork, C.; Roth, A. The Algorithmic Foundations of Differential Privacy. Found. Trends Theor. Comput. Sci. 2014, 9, 211–407. [Google Scholar] [CrossRef] [Scilit]
- McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; y Arcas, B.A. Communication-Efficient Learning of Deep Networks from Decentralized Data. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS 2017), Fort Lauderdale, FL, USA, 20–22 April 2017; PMLR: Cambridge, MA, USA, 2017; Volume 54, pp. 1273–1282. [Google Scholar]
- Al Mandalawi, S.; Mohammed, M.A.; Maclean, H.; Cakmak, M.C.; Talburt, J.R. Policy-Aware Generative AI for Safe, Auditable Data Access Governance. In Proceedings of the 17th International Conference on Knowledge and System Engineering (KSE 2025), Ho Chi Minh City, Vietnam, 6–8 November 2025; IEEE: Piscataway, NJ, USA, 2025; pp. 1–6. [Google Scholar]
- Sweller, J. Cognitive Load during Problem Solving: Effects on Learning. Cogn. Sci. 1988, 12, 257–285. [Google Scholar] [CrossRef] [PubMed]
- Carlini, N.; Tramer, F.; Wallace, E.; Jagielski, M.; Herbert-Voss, A.; Lee, K.; Roberts, A.; Brown, T.; Song, D.; Erlingsson, D.; et al. Extracting Training Data from Large Language Models. In Proceedings of the 30th USENIX Security Symposium, Virtual, 11–13 August 2021; USENIX Association: Berkeley, CA, USA, 2021; pp. 2633–2650. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
