Next Article in Journal
A Multi-Agent Framework for SQL Injection Auditing with LLM-Driven Post-Exploitation and Privilege Escalation
Previous Article in Journal
Joint Feature Selection and Hyperparameter Optimization Using Evolutionary Algorithms for Diabetes Identification Across Multiple Datasets
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

An Evidence-Based LLM Framework for Expert-Validated Clean Architecture Conformance Assessment in Enterprise Software Systems

by
Eder Jahir Gonzalez Bravo
,
Gabriel Chavira Juarez
*,
Eduardo Alvarez Navarro
,
Adriana Montoto Gonzalez
and
Jose Luis Diaz Juarez
Faculty of Engineering Tampico, Autonomous University of Tamaulipas, Centro Universitario Tampico Madero S/N, Universidad Poniente, Tampico 89336, Tamaulipas, Mexico
*
Author to whom correspondence should be addressed.
Computers 2026, 15(10), 657; https://doi.org/10.3390/computers15100657
Submission received: 1 June 2026 / Revised: 21 June 2026 / Accepted: 29 June 2026 / Published: 28 September 2026

Abstract

Architectural erosion and technical debt make it difficult to preserve intended design principles in long-lived enterprise software systems. Clean Architecture offers principles for separating business rules from frameworks, persistence and interface concerns, but its assessment often remains informal and difficult to scale across heterogeneous repositories. This article presents a hybrid framework for evaluating Clean Architecture conformance using traceable static evidence, a criterion-based rubric and structured large language model (LLM) prompts. The framework was evaluated over an inventoried corpus of 50 open-source enterprise repositories, from which 12 repositories and 36 modules were selected for an expanded validation phase. Two LLM-based configurations produced 1136 auditable candidate findings, reviewed blindly by three expert reviewers and resolved through a 2-of-3 consensus rule, with ordinal-median resolution for eight severity ties. In the final expanded gold standard, all 1136 candidate findings were consensus-supported, no candidate finding reached majority false-positive consensus, 975 rows had sufficient evidence and 161 had partial evidence. The final mean absolute score error was 0.0118 and severity agreement was 0.9525. The results suggest that LLMs can support AI-enhanced software engineering when constrained by evidence, rubrics and expert validation, but should not be treated as autonomous architectural judges.

1. Introduction

Enterprise software systems evolve under sustained pressure from changing business processes, integration needs, regulatory constraints and delivery deadlines. Over time, this evolution can weaken the separation between domain rules, application services, infrastructure, persistence and interface concerns. The resulting erosion is not merely a local code-quality issue: architectural coupling, framework leakage and persistence dependencies can increase maintenance costs, reduce testability and make business rules harder to evolve [1,2,3,4].
Clean Architecture is widely used by practitioners to express a desired separation between inner business policies and outer technical details [5]. Its central dependency rule states that dependencies should point inward, so that business rules do not depend on frameworks, databases, user interfaces or external delivery mechanisms. Related architectural traditions, including ports and adapters, domain-driven design and enterprise application patterns, also emphasize isolating the domain from infrastructure and representing dependencies through stable abstractions [6,7,8]. However, evaluating whether a real system conforms to these principles is difficult. Directory names and high-level diagrams are insufficient; the assessment requires evidence from imports, modules, framework references, persistence artifacts, dependency-inversion mechanisms and test artifacts.
Architecture conformance research provides a useful formal basis for this problem. Reflexion models compare an intended architecture with relations extracted from the implementation [9,10]. Architecture reconstruction and compliance-checking approaches similarly use source-level evidence, dependency relations and rules to detect divergence between planned and implemented structures [11,12,13,14]. These methods are rigorous, but they often require explicit architectural models, structured rules or manual mappings that are expensive to maintain across heterogeneous enterprise systems.
Large language models (LLMs) introduce a different possibility for AI-enhanced software engineering. Recent work shows that LLMs can reason over code, repositories and software tasks, but also that they are sensitive to context, prompts, benchmark contamination, hallucinations and incomplete evidence [15,16,17,18]. Similarly, the LLM-as-a-judge literature shows that rubric-based evaluation can align with human judgment in some tasks, but remains vulnerable to bias and prompt sensitivity [19,20,21,22]. These risks are particularly relevant for software quality, maintainability and architectural conformance in enterprise- and microservice-oriented systems, where plausible architectural claims can be costly if they are not grounded in inspectable evidence. Therefore, an LLM-based architecture evaluator should not be used as an oracle. It should be constrained by traceable evidence, explicit criteria, auditable outputs and expert validation.
This article presents a hybrid framework for evaluating Clean Architecture conformance in enterprise software systems. The central claim is not that LLMs should replace architectural judgment, but that an evidence–rubric–expert workflow can turn LLM-generated observations into auditable software-engineering evidence. The framework combines static evidence extraction, a Clean Architecture adherence rubric, structured prompts, auditable LLM outputs and blind expert review. The scientific novelty lies in this controlled workflow: static evidence is packaged into fixed prompts, LLM outputs are constrained by an explicit architecture rubric and structured JSON contract, and candidate findings are resolved through blind expert consensus rather than accepted as model judgments. The study addresses the following research questions:
  • RQ1: Can a structured LLM-based framework produce traceable findings for Clean Architecture conformance assessment across heterogeneous enterprise modules?
  • RQ2: How closely do the framework outputs align with a blind expert-reviewed gold standard?
  • RQ3: Which rubric criteria and validation signals reveal limits or uncertainty in the framework?
The main contribution is an empirically evaluated method, not a single prompt. The framework was applied to an initial corpus of 50 open-source enterprise repositories, an analytically stratified selection of 12 repositories and an expanded validation phase of 36 modules. Three expert reviewers evaluated 1136 auditable candidate findings under a blind protocol, producing a final expanded gold standard. The results show that the framework can produce consensus-supported findings while also revealing partial evidence cases, score corrections and severity disagreements that motivate a cautious interpretation.

2. Related Work

2.1. Architecture Conformance and Reconstruction

Architecture conformance assessments examine whether the implemented system remains consistent with an intended architecture. Reflexion models introduced a systematic way to compare a high-level model with source-level relations, identifying convergence, divergence and absence between design and implementation [9,10]. Later work extended this direction through reconstruction guidelines, compliance checking and rule-based tools [11,12,13,14,23]. These approaches motivate the evidence-first stance of the present framework: architectural claims must be grounded in observable source-code relations rather than in generic textual judgments.
Architecture erosion, architectural technical debt and architecture smells further motivate conformance assessment. Erosion occurs when implemented structures drift from intended architectural decisions, often through local changes that accumulate over time [2,4]. Architectural technical debt captures design decisions or structural compromises that increase future maintenance cost [3,24,25]. Architecture smells provide a useful vocabulary for describing recurring symptoms of erosion and debt, including excessive coupling, concern mixing and unstable dependency structures. These symptoms align with the kinds of evidence targeted by the rubric, particularly dependency direction, framework leakage, persistence coupling, cohesion and testability. Enterprise systems are especially exposed to such problems because they are long-lived, integrated and organizationally embedded.

2.2. Clean Architecture as an Evaluable Rubric

Clean Architecture, ports and adapters, domain-driven design and enterprise application patterns share a common concern for separating stable business rules from volatile technical details [5,6,7,8]. In practice, however, evaluating Clean Architecture requires operational criteria. For example, dependency direction can be inspected through imports and package relations; framework independence can be inspected through framework types, annotations and base classes; database independence can be inspected through active record models, migrations and query artifacts; dependency inversion can be inspected through interfaces, ports and adapters.
The framework therefore translates Clean Architecture principles into eight auditable criteria. The criteria are not intended to impose a single folder structure on all systems. Rather, they provide a common lens for examining whether a module separates responsibilities, keeps domain logic independent of technical details and provides sufficient evidence for each judgment.

2.3. LLMs for Software Engineering and Rubric-Based Evaluation

LLMs have been applied to issue resolution, program repair, repository question answering and code review support [15,16,17]. These tasks require multi-file context, evidence retrieval and careful reasoning over code. Hybrid transformer-based architectures have also been explored in adjacent AI domains, including cybersecurity-oriented content classification [26]. Such work illustrates the broader interest in combining neural components for complex classification tasks. The present study differs by applying LLM-based reasoning to software-architecture evidence under rubric-based expert validation. At the same time, code hallucination studies show that LLM outputs may contain plausible but unsupported claims, incorrect dependencies or incomplete reasoning [18,27]. This risk is particularly important for architectural evaluation, where a plausible explanation can be misleading if it is not tied to concrete files, classes or dependencies.
The LLM-as-a-judge literature suggests that structured rubrics and explicit evaluation criteria can improve alignment with human judgment [19,21]. However, studies also report position bias, prompt sensitivity and vulnerability to superficial presentation features [20,22,28]. The present study responds to these risks by requiring evidence fields, using blind expert review and computing agreement against an expert-resolved gold standard.

2.4. Design Science Framing

The study follows a design science orientation: it builds and evaluates an artifact intended to improve a practical software-engineering task [29,30,31]. Evaluation is not limited to demonstrating that a tool can run; it examines whether the artifact produces auditable outputs that expert reviewers can assess. This framing is appropriate because the contribution is a method and evaluation workflow for architecture assessment, where the quality of the artifact depends on traceability, interpretability and expert-resolved validity rather than on automation alone. In contrast to studies that treat LLMs primarily as autonomous judges or code-generation assistants, this work treats the LLM component as one part of a design artifact whose validity depends on evidence packaging, rubric operationalization, structured outputs and expert-resolved validation.

3. Materials and Methods

3.1. Framework Overview

The framework follows an evidence-first pipeline designed to reduce unsupported architectural claims, as summarized in Figure 1. First, candidate enterprise repositories are inventoried and modules are selected as units of analysis. Second, static evidence is extracted from source structures, imports, framework references, persistence artifacts, tests and module organization. Third, structured prompts present the evidence and rubric criteria to LLM configurations. Fourth, each output is transformed into an auditable row containing the criterion, score, severity, evidence, finding, recommendation and review flags. Finally, blind expert review resolves the final gold standard.

3.2. Corpus and Sample

The initial corpus contained 50 open-source enterprise repositories. These repositories covered domains such as CRM, ERP, IT service management, accounting, inventory, documentation, low-code internal tools and project management. The primary languages included PHP, Python, JavaScript, Java, TypeScript, Ruby, Perl, D, Dart and Go. No builds, dependency installations, tests, containers or services were executed during the corpus inventory.
Repositories were eligible when they represented enterprise-oriented systems, were publicly accessible, contained analyzable source code and exposed module-level structures suitable for static architectural inspection. Repositories were excluded when they were primarily libraries, templates, toy projects, documentation-only repositories, archived projects without sufficient source structure or systems whose modules could not be meaningfully separated for architectural analysis.
From this corpus, 12 repositories were selected through an analytically stratified selection across business domains and implementation technologies. Table 1 summarizes the resulting empirical units. The stratification goal was analytical coverage of heterogeneous enterprise systems rather than proportional population inference.
The selection did not apply fixed quantitative popularity thresholds, such as minimum stars, forks or release counts. Instead, repositories were retained when they provided sufficient architectural surface for inspection, sufficiently developed module boundaries, traceable source-code evidence and diversity across enterprise domains and technology stacks. The selected repositories were dolibarr, odoo, ofbiz-framework, tryton, SuiteCRM, twenty, glpi, LedgerSMB, InvenTree, openproject, appsmith and BookStack. The expanded validation phase evaluated 36 modules derived from these repositories. The module was used as the unit of analysis because Clean Architecture conformance is expressed through relationships among responsibilities, layers, dependencies and technical details. Evaluating entire repositories would dilute module-level findings, while isolated files would remove architectural context.
The selected sample included widely used open-source enterprise systems across multiple domains, such as Odoo for ERP [32], SuiteCRM for customer relationship management [33], and OpenProject for project management [34]. Additional repositories were selected to preserve domain and technology diversity.

3.3. Rubric Operationalization

The rubric used an ordinal score from 0 to 4, where 0 indicates strong violation or absence of conformance evidence, and 4 indicates strong and traceable conformance. Table 2 summarizes the eight criteria.
The rubric operationalizes Clean Architecture principles into auditable criteria that can be evaluated from static repository evidence. Table 3 summarizes the construct rationale used to connect each criterion with observable indicators. The purpose of the rubric is not to enforce a single folder convention, but to make architectural judgments traceable to dependency direction, boundary separation, framework and persistence coupling, abstraction mechanisms, cohesion, testability and evidence quality.

3.4. LLM Evaluation Configurations

The expanded phase used two hosted LLM configurations, DeepSeek V4 Flash and Qwen 3.6 Plus. The hosted LLM configurations were identified at execution time as DeepSeek V4 Flash and Qwen 3.6 Plus. DeepSeek was accessed as a hosted model service provided by DeepSeek (Hangzhou, China), whereas Qwen was accessed through Alibaba Cloud/Tongyi Qianwen (Alibaba Cloud, Hangzhou, China). The configurations were applied across two prompt families, resulting in 144 valid LLM runs over 72 fixed input packages. Table 4 summarizes the execution controls used for reproducibility. Each package contained the structured prompt, the CA01–CA08 rubric and the static evidence extracted for the corresponding module. Each output was required to follow a structured JSON contract. In the expanded phase, all 144 outputs were valid JSON and there were no parse errors.
The execution protocol is interface-agnostic: equivalent APIs, local tools or command-line harnesses can reproduce the procedure if they preserve the same prompt packages, model identifiers, output schema, execution dates and logging artifacts.
The framework did not execute the target systems. This was a deliberate methodological decision: the goal was to evaluate static architectural evidence across heterogeneous enterprise repositories without introducing build-system failures, dependency installation differences or runtime configuration barriers. This decision improves comparability but limits the claims to static evidence and expert assessment of that evidence.
This boundary is especially relevant for CA05, dependency inversion. In this study, CA05 was interpreted as an assessment of observable static indicators, including interfaces, ports, adapters, dependency-injection declarations, factories and service boundaries. These indicators can reveal whether a module appears to mediate dependencies through abstractions, but they do not fully verify runtime wiring, reflection-based dependencies, dynamic imports, service-container behavior or deployment-time configuration. CA05 should therefore be read as static evidence of dependency-inversion support rather than as a complete runtime verification of dependency wiring.

3.5. Blind Expert Review and Gold-Standard Construction

The expanded review package contained 1136 auditable candidate findings. Three expert reviewers independently evaluated all 1136 rows, producing 3408 expert review records. Table 5 summarizes the anonymized reviewer profiles. The expert reviewers were not manuscript authors, did not participate in prompt construction and did not develop the framework; their role in the blind expert-review stage was limited to independent assessment of the prepared artifacts. The review package anonymized the origin of the artifacts so that reviewers did not see the producing model or prompt family during evaluation.
The expert profiles were used to cover complementary perspectives on code quality, architecture, maintainability, testing and software validation. The procedure, summarized in Figure 2, was designed as blind expert validation of the artifacts. Reviewers completed their initial assessments independently using the same rubric, codebook and response definitions. The origin of each artifact, including model and prompt family, was hidden during review. No group discussion was used to alter initial reviewer judgments; consensus was computed after individual reviews using the predefined rules.
For finding acceptance, reviewers used three values: yes, partial and no. A yes value indicates that the finding was supported by the available static evidence. A partial value indicates that the finding was plausible or conditionally supported, but that its evidence, scope or formulation required caution. A no value indicates that the finding was not supported. For consensus acceptance, yes and partial were treated as favorable signals, because both indicate that the finding retained some expert-supported validity; this rule is reported explicitly to avoid interpreting consensus support as unanimous or perfect acceptance.
Finding acceptance, false-positive status and evidence sufficiency were resolved using a 2-of-3 majority rule. The final score was resolved using the median of expert final scores. Severity was resolved by majority vote; when no majority existed, an ordinal-median rule was applied using the order none < low < medium < high < critical < unknown. Eight severity ties required this documented rule. No fourth vote was added and no expert responses were altered.

3.6. Metric Definitions

The mean absolute score error compares the original LLM-assigned ordinal score, s c o r e i L L M , with the final expert-consensus score, denoted s c o r e i c o n s e n s u s :
M A E s c o r e = 1 N ∑ i = 1 N s c o r e i L L M − s c o r e i c o n s e n s u s .
Severity agreement is the proportion of rows for which the original LLM-assigned severity matched the final consensus severity after expert review and documented tie resolution. Expert-variable agreement is reported as exact agreement among the three reviewers for each reviewed row. To complement exact agreement, pairwise Cohen’s weighted kappa and Gwet’s AC1/AC2 were computed as reliability and prevalence-sensitivity checks, respectively. Krippendorff’s alpha was computed as a global multi-expert reliability coefficient for the central ordinal variables of final score and final severity. McNemar exact tests were used only to identify systematic differences in binary expert decisions; they were not treated as agreement coefficients. These metrics are used to evaluate alignment with the final gold standard, not superiority over non-LLM baselines.

3.7. Lightweight Non-LLM Baseline

To contextualize the proposed workflow, a lightweight deterministic baseline was included using only the static-evidence features extracted before LLM evaluation. The baseline did not use LLM-generated findings, recommendations, scores or severities. Instead, it operated at the module-criterion level, producing one score and one severity prediction for each of the 36 modules and eight rubric criteria, for a total of 288 baseline units.
The baseline mapped pre-existing static signals to rule-based risk scores. The signals included architectural role, source-file count, import count, framework indicators, persistence indicators, test indicators, coupling indicators and snippet availability. For example, CA01 penalized domain modules with framework or persistence signals; CA03 penalized framework dependence; CA04 penalized persistence coupling; CA05 used indirect dependency-inversion signals such as direct detail dependencies without visible abstraction indicators; CA07 used test-signal absence and coupling indicators; and CA08 used the availability of traceable snippets. Risk scores were converted deterministically into ordinal scores and severities. The baseline is therefore a transparent non-LLM reference point, not a full architecture-conformance tool.
Because the baseline produces one prediction per module and criterion, the expert-resolved gold standard was aggregated to the same module-criterion level. Consensus scores were aggregated by rounded median, and consensus severities by ordinal median. This design makes the baseline comparable at the same empirical unit while preserving the original row-level gold standard for the main analysis.

4. Results

4.1. Expanded Gold Standard

The expanded gold standard contains 1136 resolved rows. Table 6 summarizes the global results, Table 7 reports the expert-review friction behind the consensus and Table 8 reports stricter sensitivity scenarios. The main result is that all 1136 candidate findings were resolved as consensus-supported findings under the reported rule that treats both yes and partial as favorable acceptance signals. This result should be interpreted as consensus support rather than as perfect agreement or complete evidence: 161 rows had only partial evidence sufficiency, and expert reviewers corrected scores and severities in multiple cases. Thus, the result supports the usefulness of structured, evidence-based findings, not the infallibility of the LLM configurations.
Table 8 separates the primary consensus-support metric from stricter interpretations of expert agreement. The main 1136/1136 result therefore represents rule-based consensus support, not unanimous full acceptance. Under a stricter reading, 847 rows reached unanimous full finding acceptance, 1134 rows received at least two full yes finding-acceptance votes and 161 rows retained only partial evidence sufficiency.

4.2. Agreement Across Expert Review Variables

Table 9 reports agreement across expert-review variables. Agreement was strongest for false-positive assessment and final severity. Finding acceptance agreement was lower, reflecting differences in how reviewers treated fully acceptable versus partially acceptable findings. This pattern is important for interpretation: the final consensus was favorable, but it was not reached through uniform reviewer judgments.
Table 10 reports the additional reliability checks for the expanded expert review. Pairwise Cohen’s weighted kappa and Gwet AC2 were high for the two central ordinal variables, and Krippendorff’s alpha showed strong global multi-expert reliability for final score and final severity. Evidence sufficiency is reported as a diagnostic variable: although exact agreement was 0.8178, Krippendorff’s alpha was lower when yes, partial and no were treated as an ordinal scale. This indicates that reviewers differed in how they used the partial evidence category, which supports the cautious interpretation of consensus-supported findings.

4.3. Lightweight Non-LLM Baseline Results

Table 11 reports the lightweight rule-based baseline results. The baseline covered all 288 module-criterion units. Its mean absolute score error was 0.8819 against the aggregated expert gold standard, with exact score agreement of 0.3915 and severity agreement of 0.4097. At the same module-criterion aggregation level, the LLM-assisted workflow had a mean absolute score error of 0.0142 and severity agreement of 0.9479. McNemar exact tests over paired correctness decisions indicated systematic differences favoring the LLM-assisted workflow for both exact score match and exact severity match ( p < 0.001 ). These results should be interpreted as contextual evidence against a simple deterministic reference baseline, not as a claim of superiority over all possible static-analysis or architecture-conformance tools.

4.4. Results by Model

Table 12 shows the final metrics by model. Both models produced candidate findings that were resolved as consensus-supported, and neither produced parse errors in the expanded phase. Qwen had a lower mean absolute score error and slightly higher severity agreement, while DeepSeek had a lower proportion of rows with evidence sufficiency marked as yes.

4.5. Results by Rubric Criterion

Table 13 reports results by criterion. CA05, dependency inversion, had the lowest evidence-sufficiency rate and the lowest severity agreement. This is consistent with the nature of the criterion: dependency inversion is often harder to assess from static evidence alone because interfaces, injection mechanisms and adapter boundaries can be implicit, configured externally or framework-dependent. The CA05 result should therefore not be interpreted as an isolated model failure; rather, it identifies the rubric dimension where expert reviewers most often required caution because the available evidence did not fully expose dependency contracts or runtime wiring.

4.6. Pilot and Expanded Phase

The pilot phase contained 274 auditable rows and was used to stabilize the review protocol, output format and expert-review process. The expanded phase increased the evaluation to 36 modules and 1136 rows. This progression is important because it separates feasibility testing from the main empirical validation. The article therefore uses the expanded gold standard as its main result, while the pilot is treated as an earlier methodological phase.

5. Discussion

5.1. Interpretation of the Main Findings

The results support the feasibility of a constrained LLM-based framework for Clean Architecture conformance assessment. All 1136 candidate findings were resolved as consensus-supported, and the final score error was low. However, this result must be interpreted through the expert-friction and consensus-sensitivity signals reported above. Consensus support is not equivalent to unanimous acceptance, and acceptance is not equivalent to perfect evidence. In 161 cases, evidence was only partially sufficient. These cases are methodologically useful because they identify where the LLM output was plausible and relevant but not fully supported by the available static evidence.
The criterion-level results also show that some Clean Architecture dimensions are easier to evaluate than others. CA08, evidence traceability, achieved full evidence sufficiency and severity agreement because it is directly tied to whether the output cites verifiable paths, symbols and justifications. By contrast, CA05, dependency inversion, was more difficult. Static indicators can reveal visible abstractions, direct dependencies and adapter structures, but they may not capture how components are actually connected through dependency-injection containers, reflection, dynamic imports or framework-specific configuration. This suggests that future versions of the framework should improve extraction of explicit interface contracts, dependency-injection configuration, adapter registration and runtime wiring evidence.

5.2. Why Expert Validation Matters

The study does not claim that LLMs independently determine architectural truth. The final gold standard was produced through blind expert review and consensus. This distinction is central: the LLM configurations generated auditable candidate findings, while expert review determined whether those findings were supported, whether the evidence was sufficient and whether scores or severities required correction. The framework is therefore best understood as an architecture-review assistance mechanism. The methodological contribution is the controlled transformation of LLM-generated architectural observations into auditable, evidence-linked and expert-resolved findings.

5.3. Implications for Architecture Governance

The framework can help architecture teams and researchers structure reviews that would otherwise remain informal. It provides a common rubric, requires each claim to cite evidence and produces rows that can be audited or sampled by experts. This is especially relevant in enterprise systems, where modules mix business rules, persistence, user interfaces, framework conventions and legacy structures. A structured approach can make architectural erosion more visible without requiring full execution of every target system.
From an architecture-governance perspective, the candidate findings can be treated as auditable review items rather than as automatic decisions. Repeated CA01, CA03, CA04 or CA05 findings may indicate dependency-rule erosion, framework leakage, persistence coupling or weak abstraction boundaries; CA07 and CA08 findings can expose testability and traceability gaps. These signals can support periodic architecture reviews, technical-debt triage, refactoring prioritization and governance dashboards by giving teams a shared evidence base for discussion.
This application remains decision-support oriented. Enterprise architecture decisions still require human judgment about business criticality, ownership, release constraints, operational risk and planned evolution. The framework therefore fits best as a structured input to architecture governance and software-quality processes, where expert reviewers or architecture boards decide which findings become backlog items, refactoring tasks or follow-up investigations.
The results also clarify the boundary of the contribution. The lightweight non-LLM baseline provides a transparent deterministic reference point and shows that simple static rules were less aligned with the expert-resolved gold standard than the LLM-assisted workflow. However, this does not demonstrate superiority over manual review, reflexion-model tools, architecture-reconstruction systems or mature industrial conformance-checking tools. Stronger external baselines would be valuable in future work.

5.4. Threats to Validity

Construct validity is limited by the operationalization of Clean Architecture into eight criteria. Although these criteria are grounded in dependency direction, framework independence, persistence independence, dependency inversion, cohesion, testability and traceability, they do not capture every possible architectural decision.
Internal validity is affected by the static-evidence design. The framework did not execute builds, tests, services or containers. Runtime dependency injection, dynamic imports, reflection and deployment-time configuration may not be fully visible in static evidence. This limitation is particularly relevant for CA05 because dependency inversion can depend on framework-specific wiring mechanisms that are only partially observable from repository structure and source-code indicators. A similar caution applies to CA06, where cohesion and module responsibility were assessed from observable source-code structure, naming, dependency patterns and responsibility distribution, rather than from runtime behavior or long-term maintenance evolution.
Conclusion validity is affected by the consensus process. Three experts reviewed all rows, and eight severity ties were resolved through a documented ordinal-median rule. This rule closes the gold standard consistently, but it does not replace an additional expert vote. In addition, the consensus-acceptance rule treats partial as a favorable signal, so consensus support should be interpreted as expert-supported validity rather than unqualified correctness. The consensus-sensitivity analysis mitigates this threat by reporting stricter full-acceptance and evidence-sufficiency scenarios separately from the primary consensus-support metric.
External validity is limited to the selected open-source enterprise modules. The selected repositories provide analytical diversity across domains and technologies, but they do not constitute a statistically representative sample of all enterprise systems. The initial corpus contained 50 repositories, but the closed empirical validation applies to the selected 36 modules, not to the entire corpus or to all enterprise systems.
Reliability depends on the stability of prompts, models, evidence extraction and review instructions. The study mitigates this risk through structured outputs, blind review, fixed criteria and documented metrics, but future replications should version prompts, model parameters, repository snapshots and extraction scripts. Exact output-level replication may also be affected by provider-side model updates and execution-interface defaults; this risk was mitigated by preserving fixed prompt packages, model identifiers, execution dates, raw outputs, parsed outputs, exit codes and validation reports. This study did not experimentally vary model capacity, context-window length, inference costs or prompt formulations; these factors remain relevant for scalability, cost-effectiveness and robustness under controlled deployments. The expert panel was intentionally composed to cover complementary software-engineering perspectives; nevertheless, broader replications with additional reviewers would further strengthen external confidence in the review protocol.

6. Conclusions

This article presented a hybrid LLM-based framework for evaluating Clean Architecture conformance in enterprise software systems. The framework combines static evidence, a rubric, structured prompts, auditable outputs and blind expert review. In the expanded validation phase, 36 modules produced 1136 auditable candidate findings reviewed by three experts. The final gold standard resolved all rows as consensus-supported, with no majority false-positive consensus, 975 rows with sufficient evidence, 161 rows with partial evidence, a mean absolute score error of 0.0118 and severity agreement of 0.9525.
The findings indicate that LLMs can support architecture conformance assessment when their outputs are constrained by evidence and evaluated by experts. The strongest contribution is not automation alone, but rather the combination of evidence extraction, rubric-based judgment and consensus validation. Future work should extend the framework with stronger external static-analysis and architecture-conformance baselines, stronger extraction of dependency-inversion evidence, execution-aware checks where feasible, controlled analysis of model capacity, context length, inference costs and prompt sensitivity, and replication across additional systems and industrial contexts.

Author Contributions

Conceptualization, E.J.G.B.; methodology, E.J.G.B.; software, E.J.G.B.; validation, E.J.G.B., G.C.J., E.A.N., A.M.G. and J.L.D.J.; formal analysis, E.J.G.B.; investigation, E.J.G.B. and G.C.J.; resources, G.C.J.; data curation, E.J.G.B., A.M.G. and J.L.D.J.; writing—original draft preparation, E.J.G.B.; writing—review and editing, E.J.G.B., G.C.J., E.A.N., A.M.G. and J.L.D.J.; visualization, E.J.G.B.; supervision, E.A.N.; project administration, A.M.G. and J.L.D.J. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable. The study analyzed open-source software artifacts and used anonymized expert review records for methodological validation.

Informed Consent Statement

Not applicable.

Data Availability Statement

The list of repositories analyzed, aggregated metrics, rubric definitions and review protocol descriptions can be made available upon reasonable request, subject to anonymization and repository licensing constraints.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Perry, D.E.; Wolf, A.L. Foundations for the Study of Software Architecture. ACM Sigsoft Softw. Eng. Notes 1992, 17, 40–52. [Google Scholar] [CrossRef] [Scilit]
  2. Rosik, J.; Le Gear, A.; Buckley, J.; Babar, M.A.; Connolly, D. Assessing Architectural Drift in Commercial Software Development: A Case Study. Softw. Pract. Exp. 2011, 41, 63–86. [Google Scholar] [CrossRef] [Scilit]
  3. Besker, T.; Martini, A.; Bosch, J. Managing Architectural Technical Debt: A Unified Model and Systematic Literature Review. J. Syst. Softw. 2018, 135, 1–16. [Google Scholar] [CrossRef] [Scilit]
  4. Li, R.; Liang, P.; Soliman, M.; Avgeriou, P. Understanding Software Architecture Erosion: A Systematic Mapping Study. J. Softw. Evol. Process 2022, 34, e2423. [Google Scholar] [CrossRef] [Scilit]
  5. Martin, R.C. Clean Architecture: A Craftsman’s Guide to Software Structure and Design; Pearson: London, UK, 2017. [Google Scholar]
  6. Cockburn, A. Hexagonal Architecture; Humans and Technology: Salt Lake City, UT, USA, 2005; Available online: https://alistair.cockburn.us/hexagonal-architecture/ (accessed on 19 May 2026).
  7. Evans, E. Domain-Driven Design: Tackling Complexity in the Heart of Software; Addison-Wesley: Boston, MA, USA, 2003. [Google Scholar]
  8. Fowler, M. Patterns of Enterprise Application Architecture; Addison-Wesley: Boston, MA, USA, 2002. [Google Scholar]
  9. Murphy, G.C.; Notkin, D.; Sullivan, K.J. Software Reflexion Models: Bridging the Gap between Source and High-Level Models. In Proceedings of the 3rd ACM SIGSOFT Symposium on Foundations of Software Engineering; Association for Computing Machinery: New York, NY, USA, 1995; pp. 18–28. [Google Scholar] [CrossRef] [Scilit]
  10. Murphy, G.C.; Notkin, D.; Sullivan, K.J. Software Reflexion Models: Bridging the Gap between Design and Implementation. IEEE Trans. Softw. Eng. 2001, 27, 364–380. [Google Scholar] [CrossRef] [Scilit]
  11. Kazman, R.; Carrière, S.J. Playing Detective: Reconstructing Software Architecture from Available Evidence; Technical Report CMU/SEI-97-TR-010; Software Engineering Institute, Carnegie Mellon University: Pittsburgh, PA, USA, 1997. [Google Scholar]
  12. Ducasse, S.; Pollet, D. Software Architecture Reconstruction: A Process-Oriented Taxonomy. IEEE Trans. Softw. Eng. 2009, 35, 573–591. [Google Scholar] [CrossRef] [Scilit]
  13. Knodel, J.; Popescu, D. A Comparison of Static Architecture Compliance Checking Approaches. In Proceedings of the Sixth Working IEEE/IFIP Conference on Software Architecture; IEEE: New York, NY, USA, 2007. [Google Scholar] [CrossRef] [Scilit]
  14. Pruijt, L.J.; Köppe, C.; Brinkkemper, S. HUSACCT: Architecture Compliance Checking with Rich Sets of Module and Rule Types. In Proceedings of the 29th ACM/IEEE International Conference on Automated Software Engineering; Association for Computing Machinery: New York, NY, USA, 2014; pp. 851–854. [Google Scholar] [CrossRef] [Scilit]
  15. Jimenez, C.E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; Narasimhan, K. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? In Proceedings of the International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024. [Google Scholar]
  16. Zhang, Y.; Ruan, H.; Fan, Z.; Roychoudhury, A. AutoCodeRover: Autonomous Program Improvement. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis; Association for Computing Machinery: New York, NY, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
  17. Liu, J.; Tian, J.L.; Daita, V.; Wei, Y.; Ding, Y.; Wang, Y.K.; Yang, J.; Zhang, L. RepoQA: Evaluating Long Context Code Understanding. arXiv 2024, arXiv:2406.06025. [Google Scholar] [CrossRef] [Scilit]
  18. Tian, Y.; Yan, W.; Yang, Q.; Zhao, X.; Chen, Q.; Wang, W.; Luo, Z.; Ma, L.; Song, D. CodeHalu: Investigating Code Hallucinations in LLMs via Execution-Based Verification. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Washington, DC, USA, 2025. [Google Scholar]
  19. Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; Zhu, C. G-Eval: NLG Evaluation Using GPT-4 with Better Human Alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 2511–2522. [Google Scholar] [CrossRef] [Scilit]
  20. Zheng, L.; Chiang, W.L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2023. [Google Scholar]
  21. Kim, S.; Shin, J.; Cho, Y.; Jang, J.; Longpre, S.; Lee, H.; Yun, S.; Shin, R.; Kim, S.; Thorne, J.; et al. Prometheus: Inducing Fine-Grained Evaluation Capability in Language Models. In Proceedings of the International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024. [Google Scholar]
  22. Zeng, Z.; Yu, J.; Gao, T.; Meng, Y.; Goyal, T.; Chen, D. Evaluating Large Language Models at Evaluating Instruction Following. In Proceedings of the International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024. [Google Scholar]
  23. Kazman, R.; O’Brien, L.; Verhoef, C. Architecture Reconstruction Guidelines; Technical Report CMU/SEI-2001-TR-026; Software Engineering Institute, Carnegie Mellon University: Pittsburgh, PA, USA, 2001. [Google Scholar] [CrossRef]
  24. Martini, A.; Bosch, J.; Chaudron, M. Investigating Architectural Technical Debt Accumulation and Refactoring over Time: A Multiple-Case Study. Inf. Softw. Technol. 2015, 67, 237–253. [Google Scholar] [CrossRef] [Scilit]
  25. Verdecchia, R.; Kruchten, P.; Lago, P.; Malavolta, I. Building and Evaluating a Theory of Architectural Technical Debt in Software-Intensive Systems. J. Syst. Softw. 2021, 176, 110925. [Google Scholar] [CrossRef] [Scilit]
  26. Chechkin, A.; Pleshakova, E.; Gataullin, S. A Hybrid Neural Network Transformer for Detecting and Classifying Destructive Content in Digital Space. Algorithms 2025, 18, 735. [Google Scholar] [CrossRef] [Scilit]
  27. Dou, S.; Jia, H.; Wu, S.; Zheng, H.; Zhou, W.; Wu, M.; Chai, M.; Chai, M.; Fan, J.; Xi, Z.; et al. What Is Wrong with Your Code Generated by Large Language Models? Sci. China Inf. Sci. 2025, 69, 112107. [Google Scholar] [CrossRef] [Scilit]
  28. Shi, L.; Ma, C.; Liang, W.; Diao, X.; Ma, W.; Vosoughi, S. Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge. In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia–Pacific Chapter of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 293–303. [Google Scholar]
  29. Hevner, A.R.; March, S.T.; Park, J.; Ram, S. Design Science in Information Systems Research. MIS Q. 2004, 28, 75–105. [Google Scholar] [CrossRef] [Scilit]
  30. Peffers, K.; Tuunanen, T.; Rothenberger, M.A.; Chatterjee, S. A Design Science Research Methodology for Information Systems Research. J. Manag. Inf. Syst. 2007, 24, 45–77. [Google Scholar] [CrossRef] [Scilit]
  31. Wieringa, R.J. Design Science Methodology for Information Systems and Software Engineering; Springer: Berlin/Heidelberg, Germany, 2014. [Google Scholar] [CrossRef] [Scilit]
  32. Odoo, S.A. Odoo [Source Code]; GitHub: San Francisco, CA, USA; Available online: https://github.com/odoo/odoo (accessed on 19 May 2026).
  33. SalesAgility Ltd. SuiteCRM [Source Code]; GitHub: San Francisco, CA, USA; Available online: https://github.com/SuiteCRM/SuiteCRM (accessed on 19 May 2026).
  34. OpenProject GmbH. OpenProject [Source Code]; GitHub: San Francisco, CA, USA; Available online: https://github.com/opf/openproject (accessed on 19 May 2026).
Figure 1. Evidence-first framework for Clean Architecture conformance assessment. The LLM component operates on structured evidence and a fixed rubric, while expert review resolves the final gold standard.
Figure 1. Evidence-first framework for Clean Architecture conformance assessment. The LLM component operates on structured evidence and a fixed rubric, while expert review resolves the final gold standard.
Computers 15 00657 g001
Figure 2. Blind expert review and gold-standard construction. Artifact origins were hidden during review; the private mapping was used only after consensus resolution to compute metrics by model.
Figure 2. Blind expert review and gold-standard construction. Artifact origins were hidden during review; the private mapping was used only after consensus resolution to compute metrics by model.
Computers 15 00657 g002
Table 1. Sampling stages and empirical units.
Table 1. Sampling stages and empirical units.
StageUnitCountRationale
Initial inventoryRepository50Enterprise-oriented open-source systems.
Analytically stratified selectionRepository12Analyzable module boundaries, architectural evidence and diversity across domains and technologies.
Expanded validationModule36Three analyzable modules per selected repository.
Evaluation matrixCandidate finding1136Auditable rows reviewed by experts.
Table 2. Clean Architecture adherence rubric.
Table 2. Clean Architecture adherence rubric.
IDCriterionExpected Static Evidence
CA01Dependency rule between layersImports, namespaces, packages, folders, manifests and module references.
CA02Separation of domain, application, infrastructure and interface concernsDirectory structure, module names, configuration files, classes and conventions.
CA03Framework independenceFramework imports in domain modules, framework-driven inheritance, annotations or decorators.
CA04Database independenceActive models, queries, migrations, repositories, DAOs and domain-persistence coupling.
CA05Dependency inversionInterfaces, services, dependency injection, adapters, ports and factories.
CA06Single responsibility and modular cohesionClass/file focus, responsibility mixing and location of cross-cutting logic.
CA07TestabilityTests, mocks/fakes, interfaces, fixtures and decoupling from external services.
CA08Evidence traceabilityFile paths, snippets, symbols, cross-references and verifiable justification.
Table 3. Construct rationale for the Clean Architecture adherence rubric.
Table 3. Construct rationale for the Clean Architecture adherence rubric.
IDClean Architecture RationaleObservable EvidenceAssessment Focus
CA01Dependencies should point inward toward stable business rules.Imports, package references and cross-layer dependencies.Detects outward dependencies from domain or application code toward technical details.
CA02Business rules, use cases, interfaces and infrastructure should remain separated.Module roles, directory structure, controllers, services, models and adapters.Evaluates separation of responsibilities across architectural boundaries.
CA03Frameworks should be treated as external details rather than core policy.Framework-specific imports, base classes, annotations and configuration references.Identifies framework coupling in modules expected to contain business or application logic.
CA04Persistence mechanisms should not dominate domain or use-case logic.ORM models, SQL, migrations, repositories, database APIs and persistence configuration.Evaluates coupling between architectural logic and database or storage details.
CA05High-level policy should depend on abstractions rather than concrete details.Interfaces, ports, adapters, dependency injection, factories and service boundaries.Assesses whether static evidence suggests abstraction-mediated dependencies.
CA06Modules should have coherent responsibilities and limited concern mixing.File distribution, module size, role consistency and mixed framework/persistence/interface signals.Evaluates cohesion and single responsibility at module level.
CA07Decoupled architecture should support testing of policies and use cases.Test files, test directories, fixtures, mocks and isolation indicators.Assesses whether the module exposes evidence that supports testability.
CA08Architectural claims should be auditable and tied to concrete evidence.File paths, snippets, imports, symbols and explicit justifications.Evaluates the traceability and inspectability of each judgment.
Table 4. LLM execution and reproducibility controls.
Table 4. LLM execution and reproducibility controls.
ElementControl Used in the Expanded Phase
Model configurationsDeepSeek V4 Flash and Qwen 3.6 Plus.
Prompt familiesscoring and recommendation_synthesis.
Input packages and runs72 fixed prompt packages and 144 LLM runs.
Execution harnessLocal command-line evaluation harness used to submit fixed packages and preserve execution artifacts.
Execution constraintsNo tool use, no file editing, no command execution and no inference from builds, tests or runtime behavior.
Output contractStructured JSON schema with criterion, score, severity, finding, evidence, recommendation, confidence and review flags.
Preserved artifactsRaw outputs, parsed JSON files, standard-error logs, exit codes and validation reports.
Validation result144/144 valid JSON outputs, 0 parse errors and 0 failed executions.
Runtime activity0 builds, 0 dependency installations, 0 tests and 0 containers or services executed.
Table 5. Anonymized expert-reviewer profiles.
Table 5. Anonymized expert-reviewer profiles.
ReviewerExperienceEvaluation Focus
Expert 18 yearsProgramming languages, data structures, databases, version control and programming best practices, with emphasis on code quality, clarity, maintainability, modularity and correct implementation of system logic.
Expert 210 yearsScalability, security, separation of concerns, design patterns and project organization, with emphasis on whether the codebase is robust, orderly and sustainable for future evolution.
Expert 36 yearsDevelopment methodologies, technical documentation, testing, usability and software validation, with emphasis on functional objectives, project logic and established evaluation criteria.
Table 6. Global results against the final expanded gold standard.
Table 6. Global results against the final expanded gold standard.
MetricValue
Auditable rows in final expanded gold standard1136
Expert reviewers3
Combined expert review records3408
Consensus-supported candidate findings1136
Candidate findings reaching majority false-positive consensus0
Evidence sufficient: yes975
Evidence sufficient: partial161
Evidence sufficient: no0
Mean absolute score error0.0118
Severity agreement0.9525
Pending ties after documented resolution0
Table 7. Expert-review friction behind the final consensus.
Table 7. Expert-review friction behind the final consensus.
Review SignalValue
Rows with unanimous full finding acceptance (yes, yes, yes)847
Rows with two yes votes and one partial vote287
Rows with one yes vote and two partial votes2
Individual partial finding-acceptance evaluations291
Individual score evaluations marked no295
Individual severity evaluations marked no257
Individual evidence evaluations marked partial339
Individual evidence evaluations marked no31
Individual false-positive votes2
Table 8. Consensus-sensitivity analysis for finding acceptance and evidence sufficiency.
Table 8. Consensus-sensitivity analysis for finding acceptance and evidence sufficiency.
Sensitivity ScenarioRowsRate
Primary consensus support under yes/partial rule11361.0000
Unanimous full finding acceptance (yes, yes, yes)8470.7456
At least two full yes finding-acceptance votes11340.9982
One yes and two partial finding-acceptance votes20.0018
Unanimous sufficient evidence (yes, yes, yes)9290.8178
Majority sufficient evidence yes9750.8583
Consensus partial evidence1610.1417
Table 9. Agreement by expert-review variable in the expanded phase.
Table 9. Agreement by expert-review variable in the expanded phase.
VariableAgreement
Finding acceptance0.7456
Evidence sufficiency0.8178
False-positive assessment0.9982
Final score0.8917
Final severity0.9516
Table 10. Reliability checks for expert-review variables in the expanded phase.
Table 10. Reliability checks for expert-review variables in the expanded phase.
VariableRoleExact AgreementKrippendorff’s α Interpretation
Final scorePrimary ordinal0.89250.9703Strong global reliability
Final severityPrimary ordinal0.95240.9813Strong global reliability
Evidence sufficiencyDiagnostic ordinal0.81780.2932Partial-evidence friction
Table 11. Lightweight non-LLM rule-based baseline against the aggregated expert gold standard.
Table 11. Lightweight non-LLM rule-based baseline against the aggregated expert gold standard.
MetricRule-Based BaselineLLM-Assisted Workflow
Module-criterion units288288
Mean absolute score error0.88190.0142
Exact score agreement0.39150.9858
Score within one point0.79861.0000
Severity agreement0.40970.9479
Table 12. Model-level results against the final expanded gold standard.
Table 12. Model-level results against the final expanded gold standard.
ModelRowsConsensus SupportFP ConsensusEvidence YesScore ErrorSeverity Agreement
DeepSeek V4 Flash5601.00000.00000.82140.01480.9500
Qwen 3.6 Plus5761.00000.00000.89410.00900.9549
Table 13. Criterion-level results against the final expanded gold standard.
Table 13. Criterion-level results against the final expanded gold standard.
CriterionRowsConsensus SupportEvidence YesScore ErrorSeverity Agreement
CA011421.00000.90140.00000.9718
CA021421.00000.91550.01450.9648
CA031421.00000.85920.03650.9437
CA041421.00000.80280.02920.9437
CA051421.00000.66200.00000.8873
CA061421.00000.90850.01450.9507
CA071421.00000.81690.00000.9577
CA081421.00001.00000.00001.0000
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Gonzalez Bravo, E.J.; Chavira Juarez, G.; Alvarez Navarro, E.; Montoto Gonzalez, A.; Diaz Juarez, J.L. An Evidence-Based LLM Framework for Expert-Validated Clean Architecture Conformance Assessment in Enterprise Software Systems. Computers 2026, 15, 657. https://doi.org/10.3390/computers15100657

AMA Style

Gonzalez Bravo EJ, Chavira Juarez G, Alvarez Navarro E, Montoto Gonzalez A, Diaz Juarez JL. An Evidence-Based LLM Framework for Expert-Validated Clean Architecture Conformance Assessment in Enterprise Software Systems. Computers. 2026; 15(10):657. https://doi.org/10.3390/computers15100657

Chicago/Turabian Style

Gonzalez Bravo, Eder Jahir, Gabriel Chavira Juarez, Eduardo Alvarez Navarro, Adriana Montoto Gonzalez, and Jose Luis Diaz Juarez. 2026. "An Evidence-Based LLM Framework for Expert-Validated Clean Architecture Conformance Assessment in Enterprise Software Systems" Computers 15, no. 10: 657. https://doi.org/10.3390/computers15100657

APA Style

Gonzalez Bravo, E. J., Chavira Juarez, G., Alvarez Navarro, E., Montoto Gonzalez, A., & Diaz Juarez, J. L. (2026). An Evidence-Based LLM Framework for Expert-Validated Clean Architecture Conformance Assessment in Enterprise Software Systems. Computers, 15(10), 657. https://doi.org/10.3390/computers15100657

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop