1. Introduction
Enterprise software systems evolve under sustained pressure from changing business processes, integration needs, regulatory constraints and delivery deadlines. Over time, this evolution can weaken the separation between domain rules, application services, infrastructure, persistence and interface concerns. The resulting erosion is not merely a local code-quality issue: architectural coupling, framework leakage and persistence dependencies can increase maintenance costs, reduce testability and make business rules harder to evolve [
1,
2,
3,
4].
Clean Architecture is widely used by practitioners to express a desired separation between inner business policies and outer technical details [
5]. Its central dependency rule states that dependencies should point inward, so that business rules do not depend on frameworks, databases, user interfaces or external delivery mechanisms. Related architectural traditions, including ports and adapters, domain-driven design and enterprise application patterns, also emphasize isolating the domain from infrastructure and representing dependencies through stable abstractions [
6,
7,
8]. However, evaluating whether a real system conforms to these principles is difficult. Directory names and high-level diagrams are insufficient; the assessment requires evidence from imports, modules, framework references, persistence artifacts, dependency-inversion mechanisms and test artifacts.
Architecture conformance research provides a useful formal basis for this problem. Reflexion models compare an intended architecture with relations extracted from the implementation [
9,
10]. Architecture reconstruction and compliance-checking approaches similarly use source-level evidence, dependency relations and rules to detect divergence between planned and implemented structures [
11,
12,
13,
14]. These methods are rigorous, but they often require explicit architectural models, structured rules or manual mappings that are expensive to maintain across heterogeneous enterprise systems.
Large language models (LLMs) introduce a different possibility for AI-enhanced software engineering. Recent work shows that LLMs can reason over code, repositories and software tasks, but also that they are sensitive to context, prompts, benchmark contamination, hallucinations and incomplete evidence [
15,
16,
17,
18]. Similarly, the LLM-as-a-judge literature shows that rubric-based evaluation can align with human judgment in some tasks, but remains vulnerable to bias and prompt sensitivity [
19,
20,
21,
22]. These risks are particularly relevant for software quality, maintainability and architectural conformance in enterprise- and microservice-oriented systems, where plausible architectural claims can be costly if they are not grounded in inspectable evidence. Therefore, an LLM-based architecture evaluator should not be used as an oracle. It should be constrained by traceable evidence, explicit criteria, auditable outputs and expert validation.
This article presents a hybrid framework for evaluating Clean Architecture conformance in enterprise software systems. The central claim is not that LLMs should replace architectural judgment, but that an evidence–rubric–expert workflow can turn LLM-generated observations into auditable software-engineering evidence. The framework combines static evidence extraction, a Clean Architecture adherence rubric, structured prompts, auditable LLM outputs and blind expert review. The scientific novelty lies in this controlled workflow: static evidence is packaged into fixed prompts, LLM outputs are constrained by an explicit architecture rubric and structured JSON contract, and candidate findings are resolved through blind expert consensus rather than accepted as model judgments. The study addresses the following research questions:
RQ1: Can a structured LLM-based framework produce traceable findings for Clean Architecture conformance assessment across heterogeneous enterprise modules?
RQ2: How closely do the framework outputs align with a blind expert-reviewed gold standard?
RQ3: Which rubric criteria and validation signals reveal limits or uncertainty in the framework?
The main contribution is an empirically evaluated method, not a single prompt. The framework was applied to an initial corpus of 50 open-source enterprise repositories, an analytically stratified selection of 12 repositories and an expanded validation phase of 36 modules. Three expert reviewers evaluated 1136 auditable candidate findings under a blind protocol, producing a final expanded gold standard. The results show that the framework can produce consensus-supported findings while also revealing partial evidence cases, score corrections and severity disagreements that motivate a cautious interpretation.
3. Materials and Methods
3.1. Framework Overview
The framework follows an evidence-first pipeline designed to reduce unsupported architectural claims, as summarized in
Figure 1. First, candidate enterprise repositories are inventoried and modules are selected as units of analysis. Second, static evidence is extracted from source structures, imports, framework references, persistence artifacts, tests and module organization. Third, structured prompts present the evidence and rubric criteria to LLM configurations. Fourth, each output is transformed into an auditable row containing the criterion, score, severity, evidence, finding, recommendation and review flags. Finally, blind expert review resolves the final gold standard.
3.2. Corpus and Sample
The initial corpus contained 50 open-source enterprise repositories. These repositories covered domains such as CRM, ERP, IT service management, accounting, inventory, documentation, low-code internal tools and project management. The primary languages included PHP, Python, JavaScript, Java, TypeScript, Ruby, Perl, D, Dart and Go. No builds, dependency installations, tests, containers or services were executed during the corpus inventory.
Repositories were eligible when they represented enterprise-oriented systems, were publicly accessible, contained analyzable source code and exposed module-level structures suitable for static architectural inspection. Repositories were excluded when they were primarily libraries, templates, toy projects, documentation-only repositories, archived projects without sufficient source structure or systems whose modules could not be meaningfully separated for architectural analysis.
From this corpus, 12 repositories were selected through an analytically stratified selection across business domains and implementation technologies.
Table 1 summarizes the resulting empirical units. The stratification goal was analytical coverage of heterogeneous enterprise systems rather than proportional population inference.
The selection did not apply fixed quantitative popularity thresholds, such as minimum stars, forks or release counts. Instead, repositories were retained when they provided sufficient architectural surface for inspection, sufficiently developed module boundaries, traceable source-code evidence and diversity across enterprise domains and technology stacks. The selected repositories were dolibarr, odoo, ofbiz-framework, tryton, SuiteCRM, twenty, glpi, LedgerSMB, InvenTree, openproject, appsmith and BookStack. The expanded validation phase evaluated 36 modules derived from these repositories. The module was used as the unit of analysis because Clean Architecture conformance is expressed through relationships among responsibilities, layers, dependencies and technical details. Evaluating entire repositories would dilute module-level findings, while isolated files would remove architectural context.
The selected sample included widely used open-source enterprise systems across multiple domains, such as Odoo for ERP [
32], SuiteCRM for customer relationship management [
33], and OpenProject for project management [
34]. Additional repositories were selected to preserve domain and technology diversity.
3.3. Rubric Operationalization
The rubric used an ordinal score from 0 to 4, where 0 indicates strong violation or absence of conformance evidence, and 4 indicates strong and traceable conformance.
Table 2 summarizes the eight criteria.
The rubric operationalizes Clean Architecture principles into auditable criteria that can be evaluated from static repository evidence.
Table 3 summarizes the construct rationale used to connect each criterion with observable indicators. The purpose of the rubric is not to enforce a single folder convention, but to make architectural judgments traceable to dependency direction, boundary separation, framework and persistence coupling, abstraction mechanisms, cohesion, testability and evidence quality.
3.4. LLM Evaluation Configurations
The expanded phase used two hosted LLM configurations, DeepSeek V4 Flash and Qwen 3.6 Plus. The hosted LLM configurations were identified at execution time as DeepSeek V4 Flash and Qwen 3.6 Plus. DeepSeek was accessed as a hosted model service provided by DeepSeek (Hangzhou, China), whereas Qwen was accessed through Alibaba Cloud/Tongyi Qianwen (Alibaba Cloud, Hangzhou, China). The configurations were applied across two prompt families, resulting in 144 valid LLM runs over 72 fixed input packages.
Table 4 summarizes the execution controls used for reproducibility. Each package contained the structured prompt, the CA01–CA08 rubric and the static evidence extracted for the corresponding module. Each output was required to follow a structured JSON contract. In the expanded phase, all 144 outputs were valid JSON and there were no parse errors.
The execution protocol is interface-agnostic: equivalent APIs, local tools or command-line harnesses can reproduce the procedure if they preserve the same prompt packages, model identifiers, output schema, execution dates and logging artifacts.
The framework did not execute the target systems. This was a deliberate methodological decision: the goal was to evaluate static architectural evidence across heterogeneous enterprise repositories without introducing build-system failures, dependency installation differences or runtime configuration barriers. This decision improves comparability but limits the claims to static evidence and expert assessment of that evidence.
This boundary is especially relevant for CA05, dependency inversion. In this study, CA05 was interpreted as an assessment of observable static indicators, including interfaces, ports, adapters, dependency-injection declarations, factories and service boundaries. These indicators can reveal whether a module appears to mediate dependencies through abstractions, but they do not fully verify runtime wiring, reflection-based dependencies, dynamic imports, service-container behavior or deployment-time configuration. CA05 should therefore be read as static evidence of dependency-inversion support rather than as a complete runtime verification of dependency wiring.
3.5. Blind Expert Review and Gold-Standard Construction
The expanded review package contained 1136 auditable candidate findings. Three expert reviewers independently evaluated all 1136 rows, producing 3408 expert review records.
Table 5 summarizes the anonymized reviewer profiles. The expert reviewers were not manuscript authors, did not participate in prompt construction and did not develop the framework; their role in the blind expert-review stage was limited to independent assessment of the prepared artifacts. The review package anonymized the origin of the artifacts so that reviewers did not see the producing model or prompt family during evaluation.
The expert profiles were used to cover complementary perspectives on code quality, architecture, maintainability, testing and software validation. The procedure, summarized in
Figure 2, was designed as blind expert validation of the artifacts. Reviewers completed their initial assessments independently using the same rubric, codebook and response definitions. The origin of each artifact, including model and prompt family, was hidden during review. No group discussion was used to alter initial reviewer judgments; consensus was computed after individual reviews using the predefined rules.
For finding acceptance, reviewers used three values: yes, partial and no. A yes value indicates that the finding was supported by the available static evidence. A partial value indicates that the finding was plausible or conditionally supported, but that its evidence, scope or formulation required caution. A no value indicates that the finding was not supported. For consensus acceptance, yes and partial were treated as favorable signals, because both indicate that the finding retained some expert-supported validity; this rule is reported explicitly to avoid interpreting consensus support as unanimous or perfect acceptance.
Finding acceptance, false-positive status and evidence sufficiency were resolved using a 2-of-3 majority rule. The final score was resolved using the median of expert final scores. Severity was resolved by majority vote; when no majority existed, an ordinal-median rule was applied using the order none < low < medium < high < critical < unknown. Eight severity ties required this documented rule. No fourth vote was added and no expert responses were altered.
3.6. Metric Definitions
The mean absolute score error compares the original LLM-assigned ordinal score,
, with the final expert-consensus score, denoted
:
Severity agreement is the proportion of rows for which the original LLM-assigned severity matched the final consensus severity after expert review and documented tie resolution. Expert-variable agreement is reported as exact agreement among the three reviewers for each reviewed row. To complement exact agreement, pairwise Cohen’s weighted kappa and Gwet’s AC1/AC2 were computed as reliability and prevalence-sensitivity checks, respectively. Krippendorff’s alpha was computed as a global multi-expert reliability coefficient for the central ordinal variables of final score and final severity. McNemar exact tests were used only to identify systematic differences in binary expert decisions; they were not treated as agreement coefficients. These metrics are used to evaluate alignment with the final gold standard, not superiority over non-LLM baselines.
3.7. Lightweight Non-LLM Baseline
To contextualize the proposed workflow, a lightweight deterministic baseline was included using only the static-evidence features extracted before LLM evaluation. The baseline did not use LLM-generated findings, recommendations, scores or severities. Instead, it operated at the module-criterion level, producing one score and one severity prediction for each of the 36 modules and eight rubric criteria, for a total of 288 baseline units.
The baseline mapped pre-existing static signals to rule-based risk scores. The signals included architectural role, source-file count, import count, framework indicators, persistence indicators, test indicators, coupling indicators and snippet availability. For example, CA01 penalized domain modules with framework or persistence signals; CA03 penalized framework dependence; CA04 penalized persistence coupling; CA05 used indirect dependency-inversion signals such as direct detail dependencies without visible abstraction indicators; CA07 used test-signal absence and coupling indicators; and CA08 used the availability of traceable snippets. Risk scores were converted deterministically into ordinal scores and severities. The baseline is therefore a transparent non-LLM reference point, not a full architecture-conformance tool.
Because the baseline produces one prediction per module and criterion, the expert-resolved gold standard was aggregated to the same module-criterion level. Consensus scores were aggregated by rounded median, and consensus severities by ordinal median. This design makes the baseline comparable at the same empirical unit while preserving the original row-level gold standard for the main analysis.
5. Discussion
5.1. Interpretation of the Main Findings
The results support the feasibility of a constrained LLM-based framework for Clean Architecture conformance assessment. All 1136 candidate findings were resolved as consensus-supported, and the final score error was low. However, this result must be interpreted through the expert-friction and consensus-sensitivity signals reported above. Consensus support is not equivalent to unanimous acceptance, and acceptance is not equivalent to perfect evidence. In 161 cases, evidence was only partially sufficient. These cases are methodologically useful because they identify where the LLM output was plausible and relevant but not fully supported by the available static evidence.
The criterion-level results also show that some Clean Architecture dimensions are easier to evaluate than others. CA08, evidence traceability, achieved full evidence sufficiency and severity agreement because it is directly tied to whether the output cites verifiable paths, symbols and justifications. By contrast, CA05, dependency inversion, was more difficult. Static indicators can reveal visible abstractions, direct dependencies and adapter structures, but they may not capture how components are actually connected through dependency-injection containers, reflection, dynamic imports or framework-specific configuration. This suggests that future versions of the framework should improve extraction of explicit interface contracts, dependency-injection configuration, adapter registration and runtime wiring evidence.
5.2. Why Expert Validation Matters
The study does not claim that LLMs independently determine architectural truth. The final gold standard was produced through blind expert review and consensus. This distinction is central: the LLM configurations generated auditable candidate findings, while expert review determined whether those findings were supported, whether the evidence was sufficient and whether scores or severities required correction. The framework is therefore best understood as an architecture-review assistance mechanism. The methodological contribution is the controlled transformation of LLM-generated architectural observations into auditable, evidence-linked and expert-resolved findings.
5.3. Implications for Architecture Governance
The framework can help architecture teams and researchers structure reviews that would otherwise remain informal. It provides a common rubric, requires each claim to cite evidence and produces rows that can be audited or sampled by experts. This is especially relevant in enterprise systems, where modules mix business rules, persistence, user interfaces, framework conventions and legacy structures. A structured approach can make architectural erosion more visible without requiring full execution of every target system.
From an architecture-governance perspective, the candidate findings can be treated as auditable review items rather than as automatic decisions. Repeated CA01, CA03, CA04 or CA05 findings may indicate dependency-rule erosion, framework leakage, persistence coupling or weak abstraction boundaries; CA07 and CA08 findings can expose testability and traceability gaps. These signals can support periodic architecture reviews, technical-debt triage, refactoring prioritization and governance dashboards by giving teams a shared evidence base for discussion.
This application remains decision-support oriented. Enterprise architecture decisions still require human judgment about business criticality, ownership, release constraints, operational risk and planned evolution. The framework therefore fits best as a structured input to architecture governance and software-quality processes, where expert reviewers or architecture boards decide which findings become backlog items, refactoring tasks or follow-up investigations.
The results also clarify the boundary of the contribution. The lightweight non-LLM baseline provides a transparent deterministic reference point and shows that simple static rules were less aligned with the expert-resolved gold standard than the LLM-assisted workflow. However, this does not demonstrate superiority over manual review, reflexion-model tools, architecture-reconstruction systems or mature industrial conformance-checking tools. Stronger external baselines would be valuable in future work.
5.4. Threats to Validity
Construct validity is limited by the operationalization of Clean Architecture into eight criteria. Although these criteria are grounded in dependency direction, framework independence, persistence independence, dependency inversion, cohesion, testability and traceability, they do not capture every possible architectural decision.
Internal validity is affected by the static-evidence design. The framework did not execute builds, tests, services or containers. Runtime dependency injection, dynamic imports, reflection and deployment-time configuration may not be fully visible in static evidence. This limitation is particularly relevant for CA05 because dependency inversion can depend on framework-specific wiring mechanisms that are only partially observable from repository structure and source-code indicators. A similar caution applies to CA06, where cohesion and module responsibility were assessed from observable source-code structure, naming, dependency patterns and responsibility distribution, rather than from runtime behavior or long-term maintenance evolution.
Conclusion validity is affected by the consensus process. Three experts reviewed all rows, and eight severity ties were resolved through a documented ordinal-median rule. This rule closes the gold standard consistently, but it does not replace an additional expert vote. In addition, the consensus-acceptance rule treats partial as a favorable signal, so consensus support should be interpreted as expert-supported validity rather than unqualified correctness. The consensus-sensitivity analysis mitigates this threat by reporting stricter full-acceptance and evidence-sufficiency scenarios separately from the primary consensus-support metric.
External validity is limited to the selected open-source enterprise modules. The selected repositories provide analytical diversity across domains and technologies, but they do not constitute a statistically representative sample of all enterprise systems. The initial corpus contained 50 repositories, but the closed empirical validation applies to the selected 36 modules, not to the entire corpus or to all enterprise systems.
Reliability depends on the stability of prompts, models, evidence extraction and review instructions. The study mitigates this risk through structured outputs, blind review, fixed criteria and documented metrics, but future replications should version prompts, model parameters, repository snapshots and extraction scripts. Exact output-level replication may also be affected by provider-side model updates and execution-interface defaults; this risk was mitigated by preserving fixed prompt packages, model identifiers, execution dates, raw outputs, parsed outputs, exit codes and validation reports. This study did not experimentally vary model capacity, context-window length, inference costs or prompt formulations; these factors remain relevant for scalability, cost-effectiveness and robustness under controlled deployments. The expert panel was intentionally composed to cover complementary software-engineering perspectives; nevertheless, broader replications with additional reviewers would further strengthen external confidence in the review protocol.
6. Conclusions
This article presented a hybrid LLM-based framework for evaluating Clean Architecture conformance in enterprise software systems. The framework combines static evidence, a rubric, structured prompts, auditable outputs and blind expert review. In the expanded validation phase, 36 modules produced 1136 auditable candidate findings reviewed by three experts. The final gold standard resolved all rows as consensus-supported, with no majority false-positive consensus, 975 rows with sufficient evidence, 161 rows with partial evidence, a mean absolute score error of 0.0118 and severity agreement of 0.9525.
The findings indicate that LLMs can support architecture conformance assessment when their outputs are constrained by evidence and evaluated by experts. The strongest contribution is not automation alone, but rather the combination of evidence extraction, rubric-based judgment and consensus validation. Future work should extend the framework with stronger external static-analysis and architecture-conformance baselines, stronger extraction of dependency-inversion evidence, execution-aware checks where feasible, controlled analysis of model capacity, context length, inference costs and prompt sensitivity, and replication across additional systems and industrial contexts.