Highlights
Please indicate how your work links to systems science via your contributions to systems practice, theory, and/or methodology.
- A task-oriented functional semantics (TOFS) representation is introduced to systematically bridge SysML functional models and component-level physical architectures within the MBSE process.
- A retrieval-augmented architecture generation methodology integrates component retrieval with LLM-based synthesis, providing a knowledge-grounded approach to automated system architecture design.
What are the main findings and/or the implications of the main findings?
- Hybrid component retrieval using TOFS improves the identification of engineering components relevant to system functions and design constraints.
- Grounding architecture generation in retrieved component knowledge and TOFS improves the functional completeness and engineering validity of generated architectures.
Abstract
Automated system architecture generation is a key task in Model-Based Systems Engineering (MBSE), where requirement and functional models need to be transformed into a physical architecture composed of selected components and port-level connections. However, this task remains challenging because functional information is often semi-structured and distributed, and is represented at a different level of abstraction from physical architectures. Existing knowledge-driven methods provide systematic and traceable reasoning but require labor-intensive formal modeling and rule definition, whereas direct Large Language Model (LLM)-based generation improves automation but may suffer from component hallucination and low interpretability. To address these limitations, this study proposes a retrieval-augmented system architecture generation approach using Task-Oriented Functional Semantics (TOFS). TOFS is constructed from SysML requirements and functional models through an LLM-based process, and represents system goals, functions, logical entities, ports, and engineering constraints in a unified representation. Based on TOFS, a hybrid component retrieval method is proposed, which integrates interface matching, parameter constraint matching, and semantic matching to retrieve candidate components from an engineering component library. TOFS and the retrieved component sets are then jointly organized as the generation context to guide LLM-based architecture generation. A hydronic heating control system is used as a case study, and scenario-based comparative experiments demonstrate improved component retrieval performance, functional completeness, and engineering validity.
1. Introduction
Model-Based Systems Engineering (MBSE) provides a model-centered paradigm for managing the complexity of modern engineering systems. By using formal and semi-formal models as the primary artifacts throughout the system lifecycle, MBSE supports requirement analysis, functional modeling, architecture design, verification, and traceability [1]. In many MBSE processes, system development follows a Requirement–Function–Logical–Physical (RFLP) logic [2], in which stakeholder and system requirements are progressively transformed into functional descriptions, logical architectures, and physical architectures [3]. As system scale and complexity continue to increase, traditional architecture design approaches, which heavily rely on individual experience and manual effort, have become a major bottleneck limiting system design quality [4].
However, automatically generating physical architectures from upstream MBSE artifacts remains challenging. Requirements and functions are often semi-structured, distributed across different SysML model elements, and represented at a different level of abstraction from physical architectures. As a result, the information required for physical architecture generation is not directly available as a unified and machine-processable representation. This creates a semantic gap between SysML artifacts and the component-level decisions required for physical architecture generation.
Existing studies have explored knowledge-driven methods to support automated or semi-automated architecture generation. These methods usually rely on formal functional representations, component knowledge bases, and reasoning and matching techniques to support function–component mapping and architecture synthesis. Their main advantage lies in their systematic process and traceable intermediate results. By explicitly representing functions, components, and related interface and parameter constraints, knowledge-driven methods can provide interpretable reasoning paths from functional requirements to architecture solutions. Nevertheless, these methods often assume that required information has already been formalized in advance. In practical MBSE scenarios, constructing and maintaining such formal knowledge representations from semi-structured SysML artifacts is labor-intensive and difficult to scale [5]. Therefore, although knowledge-driven methods are rigorous in principle, their practical application is limited by the high cost of knowledge preparation and rule maintenance.
Recent advances in Large Language Models (LLMs) provide new opportunities for automating MBSE-related modeling and architecture generation tasks. LLMs can interpret natural language requirements, extract structured information, and generate model-like descriptions with limited manual rule definition. These capabilities make it possible to automatically process semi-structured MBSE artifacts and reduce the dependence on manually defined formal representations. However, direct LLM-based architecture generation is usually weakly grounded in system-specific engineering knowledge. Without explicit constraints from component libraries, interfaces, parameters, and functional flows, LLMs may generate components that do not exist in the target engineering context, establish unsupported connections between components, or produce architectures with limited traceability to the original requirement and functional models [6]. Therefore, the key challenge is not merely to generate architecture descriptions from text, but to construct an intermediate system-functional representation that can connect SysML artifacts, engineering component knowledge, and component-level physical architecture generation.
The analysis of existing automated architecture generation methods shows that, when applied to practical MBSE architecture design tasks, current methods still struggle to achieve a balance between reducing manual modeling effort and maintaining the engineering validity of generated architectures.
- (1)
- Difficulty in automatically constructing formal functional semantics. Existing knowledge-driven architecture generation methods usually rely on mapping between functions and logical/physical architectures based on functional semantics. However, its effective operation depends on a strong prerequisite: the semantics on both the requirement side and the component side must have been manually organized into formal representations. In practical MBSE architecture design scenarios, however, this prerequisite is difficult to satisfy. The fundamental reason lies in the semi-structured and multi-view nature of SysML models. On the one hand, SysML model artifacts often use semi-structured natural language descriptions and usually lack explicit and strictly formalized functional descriptions. On the other hand, component retrieval needs to satisfy not only functional semantics, but also other engineering constraints such as interfaces and performance. These constraints are distributed across different model artifacts rather than derived from a single functional model.
- (2)
- Lack of engineering constraints and interpretability in direct LLM-based generation. Compared with knowledge-driven methods, LLM-based methods depend less on strictly formalized knowledge and can handle natural language requirements and semi-structured model information more flexibly. However, on the one hand, due to the lack of engineering knowledge input specific to the target system, they may generate components that do not exist in the real component library, leading to component hallucination. On the other hand, direct LLM-based generation usually lacks explicit intermediate semantic representations and inspectable generation evidence, such as functions, logical components, and physical components. As a result, when the generated architecture is unsatisfactory, it is difficult to locate defects and improve the generation process.
To address this challenge, this study proposes a retrieval-augmented system architecture generation approach for MBSE using Task-Oriented Functional Semantics (TOFS). TOFS is introduced as an intermediate system-functional representation that transforms distributed and semi-structured information in SysML requirements and functional models into structured functional semantics and engineering constraints. An LLM-assisted process is used to construct TOFS, based on which candidate physical components are retrieved from an engineering component library through interface, parameter-constraint, and semantic matching. TOFS and the retrieved candidate component sets are then jointly used as the generation context to guide LLM-based component selection and port-level connection generation.
The main contributions of this study are threefold. First, TOFS is proposed, which bridges semi-structured SysML artifacts and component-level physical architecture generation without requiring fully formalized functional representations to be manually constructed in advance. Second, a TOFS-based hybrid component retrieval method is developed, which provides an engineering-constrained candidate component space. Third, a retrieval-augmented LLM generation process is proposed to generate component-level physical architectures using both TOFS and the retrieved candidate component sets as the generation context. A hydronic heating control system (HHCS) is used as a case study to demonstrate the feasibility of the proposed approach. The experiments show that, by combining task-oriented functional semantics with retrieved component knowledge, the proposed method reduces unsupported component generation and improves the functional completeness and engineering validity of generated architectures.
The remainder of this paper is organized as follows. Section 2 reviews related work. Section 3 presents the materials and methods, including the method overview, the definition and construction of TOFS, the retrieval-augmented architecture generation approach, and the case description and experimental setup. Section 4 presents the case study and quantitative evaluation results. Section 5 discusses the main findings, advantages, limitations, and broader applicability of the proposed method. Section 6 concludes the paper and outlines directions for future work.
2. Literature Review
2.1. System Architecture Design in MBSE
In the MBSE field, there is still no unified definition of system architecture. ISO/IEC/IEEE 42010 defines architecture as “the fundamental concepts or properties of the entity considered in its environment, which may include constituent elements, interactions among the elements and with the environment, behavior and structure, and principles governing its design, use, operation and evolution” [7]. NASA defines system architecture as “the high-level unifying structure containing a set of rules, guidelines, and constraints that defines a cohesive and coherent structure consisting of constituent parts, relationships and connections that establish how those parts fit and work together” [8]. In various general MBSE methodologies [9], such as OOSEM [10], Harmony SE [11], and MBSAP [12], system architecture is regarded as a hierarchical description of a system from different perspectives, including operational architecture, functional architecture, logical architecture, and physical architecture. Although these definitions differ in wording and scope, they commonly view system architecture as a high-level organization of system elements, relationships among elements, and associated constraints, and emphasize that architecture should be described from multiple perspectives. According to this common understanding, this paper does not attempt to cover the complete holistic architecture, but instead focuses on the evolution from functional architecture to physical architecture.
In terms of architecture design methods, Mhenn et al. [13] proposed a SysML-based architecture design method for mechatronic systems, in which the design process is divided into two stages: black-box analysis and white-box analysis, respectively, supporting requirement elicitation and the elaboration of internal structure and behavior. Lemazurier et al. [14] focused on the transformation from requirements to functional architecture, and proposed an MBSE method composed of requirement, context, and behavior views to support requirement representation, architecture organization, and preliminary verification. Granrath et al. [15] proposed a model transformation method based on the Compositional Unified System-Based Engineering (CUBE) methodology. After the allocation from functions to logical components is manually completed, SysML activity diagrams are automatically transformed into logical architectures in the form of internal block diagrams. Vazquez-Santacruz et al. [16] proposed the V-cube methodology, which uses a centralized SysML model as its core and integrates conceptual design, specification definition, mission analysis, logical architecture, detailed design, and physical implementation through six parallel and interacting V-models, while supporting multidisciplinary tool collaboration and lifecycle traceability. Ma et al. [17] proposed a multi-architecture modeling method that integrates multiple languages, including UPDM, SysML, BPMN, and EAST-ADL, using a unified metamodel. Following the MOFLP process, this method supports cross-stage architecture modeling and shifts static verification and dynamic simulation earlier into the architecture design stage.
2.2. Automated System Architecture Generation
As discussed above, existing studies on automated system architecture generation can be broadly classified into two categories: knowledge-driven methods and LLM-based methods. The former rely on explicit engineering knowledge and map functions to architecture solutions through component retrieval, rule-based reasoning, constraint solving, and related techniques. The latter leverage the implicit prior knowledge and generative reasoning capabilities of LLMs to automatically generate candidate architecture solutions. The related work on these two categories of methods is reviewed below.
2.2.1. Knowledge-Driven Architecture Generation
Knowledge-driven methods are founded on the explicit representation of knowledge. In terms of functional knowledge representation, flow-based functional representation [18] is one of the most classic theories of function representation, which regards a function as a transformation applied to flows [19]. Based on this theory, a variety of more formalized and computer-understandable functional semantic modeling methods have been developed for specific automation tasks. For example, for automated function decomposition, Yuan et al. [20] proposed the concept of functional effects, and Chen et al. [21] introduced flow attribute constraints to formally model the semantics of functional verbs. For software–physical co-design, Cao et al. [22] proposed hybrid functional semantics that integrate functional effort and Computational Tree Logic (CTL). To address the difficulty of unifying functional models across multiple use cases, Yildirim et al. [23] proposed a flow-heuristic-based MBSE functional modeling method.
Based on formal functional representations, many automated architecture generation methods are proposed using functional semantic matching and logical reasoning. Chen et al. [24,25] proposed an automated physical architecture generation method that covers the complete process of function-component matching, component combination compatibility checking, and architecture evaluation, and further developed an automated architecture verification method based on this framework. Cao et al. [26,27] proposed a software-physical integrated system design method based on hybrid functional semantics, as well as a corresponding automated system behavior verification method. Wang et al. [28] introduced reinforcement learning into the architecture design space exploration, thereby generating a large number of reliable system architectures efficiently. Della Bella et al. [29] proposed a systematic architecture generation method based on design structure matrices, introducing parallel and branched architectures during the mapping from functions to physical architectures to achieve safety redundancy.
In addition to the approaches based on engineering design theory discussed above, importing multi-source heterogeneous design knowledge into knowledge graphs and realizing automated architecture generation through reasoning over such knowledge has also become an important research direction. Zhang et al. [30] proposed a knowledge-graph-based method for the automated generation of system model diagrams. By integrating text mining and model retrieval techniques, their method enables the automated transformation from multi-source heterogeneous data to standardized system model diagrams. Huang et al. [31] constructed a multi-source knowledge graph integrating product, feature, and decision data. Initial architecture solutions are generated through graph matching, and a closed loop of architecture generation and verification is formed by combining multi-objective optimization and simulation-based validation. Wang et al. [32] proposed a retrieval-augmented generation framework that integrates patent knowledge graphs with small language models. Through multi-hop retrieval over the knowledge graph, the framework expands the semantic space and constrains model generation, thereby producing high-quality answers for intelligent question answering in police equipment design.
2.2.2. LLM-Based Architecture Generation
In recent years, LLMs have demonstrated strong capabilities in natural language understanding and structured text generation. They can semantically align natural language requirements with structured representations and, to some extent, learn general patterns from data, thereby reducing the dependence on manually predefined rules and knowledge bases. As a result, they provide a new technical pathway for MBSE automation. Existing studies on LLM-based MBSE model generation can be grouped into the following categories.
One group of studies mainly focuses on the automated generation of SysML models from natural language requirements. Apvrille and Sultan [33] proposed using LLMs to automatically construct structural and behavioral views from system specifications, while also pointing out that the quality of generated models depends on well-defined knowledge bases and automated feedback mechanisms. Schleifer et al. [34] designed a sequential prompting process consisting of intra-use-case and inter-use-case prompts to generate use case diagrams from natural language requirements. Rafique et al. [35] compared the performance of different LLMs in SysML v2 model generation, and indicated that LLMs can generate initial model skeletons, but relying entirely on LLMs to generate complete models still tends to produce syntactic, semantic, and structural errors. Johnson and Williams [36] addressed the task of transforming legacy documents into SysML models by converting unstructured documents into SysML BDDs, and improved element identification and relationship extraction using RAG and graph-theoretic invariants. Cibrián et al. [37] proposed an agent-based SysML v2 model generation framework for industrial scenarios, which integrates LLMs, RAG, and a syntax validation engine to improve the syntactic correctness and semantic consistency of generated models through structured retrieval and iterative validation. An et al. [38] proposed a fully offline lightweight LLM framework for generating SysML v1 PlantUML models from unstructured natural language requirements, combining QLoRA fine-tuning, same-diagram Jaccard retrieval, and a multi-round syntax-correction loop to improve syntactic validity. Bouamra et al. [39] proposed SysTemp, a multi-agent and template-based framework for generating SysML v2 models from natural language specifications, showing the usefulness of agent collaboration and parser feedback for improving syntactic correctness in SysML model generation.
Another group of studies focuses on applying LLMs to model verification and refinement. To address the problem of multi-view consistency in SysML models, Sultan and Apvrille [40] proposed a consistency detection and repair framework that combines formal rules with LLMs to identify and correct inconsistencies between use case diagrams and block definition diagrams. Sugawara et al. [41] transformed system models into Neo4j (Neo4j, Inc., San Mateo, CA, USA) graph structures and used LLMs to convert natural language questions into Cypher queries, thereby enabling natural language question answering for non-SysML experts. Bonner et al. [42] proposed an LLM-based semi-automated traceability link recommendation method for establishing traceability between requirements and MBSE models.
In addition, LLMs have also been applied to the architecture design process. Obieke et al. [43] proposed the AICED framework, which uses multi-agent LLMs to support problem definition, product design specification generation, and conceptual solution image generation. Von Heissen et al. [44] developed an AI4Cameo plugin for LLM-assisted system architecture generation, where system descriptions and selected requirements are used to generate functional and logical architecture views in Cameo Systems Modeler. Through a comprehensive assessment, Timperley et al. [45] demonstrated the successful automated generation of spacecraft system architectures using LLMs, while highlighting the importance of human input in maintaining output quality. Yuan et al. [46] proposed a conceptual design framework integrating an MBSE knowledge graph with hierarchical LLM reasoning, which constructs a multigranular knowledge index offline and performs top-down parallel search and pruning online to enable end-to-end generation of structured aerial bomb design solutions from natural-language tactical requirements.
To clarify the positioning of this study, Table 1 compares the major research fields related to several dimensions, including the use of formal functional semantics, the automation level of semantic or model construction, the degree of engineering grounding, and the availability of intermediate artifacts. Here, engineering grounding refers to the extent to which a method explicitly uses external, curated engineering knowledge sources, such as project- or enterprise-specific component libraries, interface information, parameter constraints, or historical design assets, while intermediate artifacts refer to inspectable representations or process outputs that can support traceability and error localization.
Table 1.
Comparison of related studies on automated architecture generation.
3. Materials and Methods
3.1. Method Overview
To address the above-mentioned gaps, this study integrates traditional engineering design theory with the generative and reasoning capabilities of LLMs, and proposes a retrieval-augmented architecture generation approach based on TOFS. First, to address the first issue, i.e., the difficulty of automatically constructing formal functional semantics, this study introduces the concept of TOFS. As shown in Figure 1, TOFS consists of two parts: Core Functional Semantics (CFS) and Task-oriented Semantic Extension (TSE).
Figure 1.
Basic idea of TOFS.
CFS is used to represent task-neutral system-level functional semantics, including high-level goals, functions, and the flows among them. Its structure is similar to that of traditional functional semantics. However, unlike traditional function representations that require strictly formalized descriptions of transformations, CFS describes transformations and flow semantics in natural language. Therefore, it is more suitable for automated extraction from SysML models using LLMs.
TSE is a semantic extension built upon CFS for specific downstream tasks and is used to compensate for the lack of engineering constraints in functional semantics. In this study, the downstream tasks are component retrieval and architecture synthesis. Therefore, TSE mainly includes engineering semantic constraints such as entities, required ports, normalized parameter constraints, entity name candidates, and port name candidates.
To address the second issue, i.e., the lack of system-specific engineering constraints in direct LLM-based generation, this study further proposes a TOFS-based retrieval-augmented architecture generation approach, as shown in Figure 2. The method first uses the information contained in TOFS to retrieve components from the component library from three perspectives: interface matching, parameter constraint matching, and semantic matching. In this way, a candidate component set is obtained for each entity. Then, based on these results, the candidate component sets and TOFS are jointly organized into the RAG context, which guides the LLM to identify missing entities, select components, and generate port-level connections within a constrained candidate component space. As a result, a feasible system architecture that satisfies functional requirements and engineering constraints can be generated.
Figure 2.
Retrieval-augmented architecture generation approach.
3.2. Definition and Construction of TOFS
In modeling-language research, semantics is commonly distinguished from syntax and refers to the interpretation or behavior denoted by model expressions [47]. For example, for functions modeled as SysML Activities, the semantics mainly concerns the behavior denoted by the activity. However, only limited functional information, such as the activity name and its input/output relationships, may be explicitly represented. The corresponding functional semantics therefore needs to be further extracted and made explicit to support subsequent tasks. In the engineering design domain, functional semantics is usually described as a transformation applied to material, energy, or signal flows [18]. Combining these two views, this study defines Task-Oriented Functional Semantics (TOFS) as a structured intermediate representation that makes the semantics embedded in SysML requirements and functional models explicit for subsequent tasks.
3.2.1. Formal Definition of TOFS
As discussed above, the basic idea of TOFS is to first represent system-level core functional semantics in a semi-formal way, and then extend corresponding engineering semantic constraints according to downstream engineering tasks. Therefore, TOFS is formally defined as the following tuple:
- Core functional semantics (CFS)
is used to represent task-neutral system-level functional semantics. Its purpose is to describe the goals that the system needs to achieve, the functions it needs to perform, and the interactions among these functions. Its structure is similar to traditional flow-based functional representation [19,48], in which product functions are represented through functional actions and the flows of material, energy, or signal being transformed, but it uses natural language to describe the transformations applied to flows. In addition, high-level system goals are introduced to describe global functional and performance information. Formally, is defined as
- (1)
- denotes the set of high-level system goals. Each describes, in natural language, a system objective [49] extracted from the requirement model, providing the purpose context for the functional behavior.
- (2)
- denotes the set of system functions. Each describes the semantics of a function in natural language, and is parsed from an action in the functional model. It is defined as:
- (3)
- denotes the set of input and output flows among functions. Each flow is defined as:
Taking the HHCS as an example, one of the high-level goals is “to maintain the indoor temperature around a user-defined setpoint”. A function may be “Generate Thermal Energy”, and a corresponding flow may be the thermal-fluid flow transferred from the heat generation function to the downstream circulation process.
As can be seen from the above definition, unlike traditional functional representations, CFS does not require functions to be described in a strictly formalized manner. Instead, while retaining basic flow type information, it uses natural language to describe functional transformations and flow semantics. This makes the representation more flexible for semantic matching and can be aligned with the textual information that is commonly available in engineering component libraries without introducing additional formal semantic modeling effort for every component.
Compared with directly using raw requirements and functional descriptions as retrieval text in RAG, CFS has two main advantages. First, it distinguishes different types of information through explicit schema fields such as high-level goals, functions, and flows, which can play different roles in the retrieval process, such as text-based similarity matching and constraint-based filtering. Second, CFS serves as an intermediate functional modeling result that can be reviewed and refined by human engineers in the LLM-assisted human-in-the-loop TOFS construction process. This makes the process more transparent and trustworthy than an end-to-end RAG process.
It should be noted that CFS mainly captures core functional semantics, while detailed non-functional requirements, such as quality, performance, safety, and physical constraints, are not encoded as long natural-language CFS descriptions, but are represented separately in TSE as normalized parameter constraints.
- Task-oriented semantic extension (TSE)
TSE is used to supplement CFS with engineering semantics related to downstream tasks. Since the downstream tasks considered in this study are component retrieval and system architecture synthesis, the structure of TSE is designed for the following reasons: First, according to the RFLP views [2], a logical layer is needed to bridge functional descriptions and physical realization, so TSE introduces entities and function–entity allocation relations. Second, since component retrieval depends on explicit engineering conditions [49], TSE normalizes interface and parameter constraints embedded in requirements and functional descriptions into retrieval-oriented constraints. Third, because the terminology used in SysML models may differ from that used in component libraries, TSE introduces candidate name sets to expand retrieval queries. Therefore, TSE mainly contains entities that may realize system functions, together with their interface and parameter constraints, and further extends candidate name sets to improve retrieval performance. Formally, TSE is defined as:
- (1)
- denotes logical components in the RFLP process. An Entity represents a logical-level realization element of one or more functions, such as sensor, controller, processor, and actuator. The semantic level of entities lies between functions and physical components. Each entity is defined as:
- (2)
- denotes the set of semantic alignment candidates. It is used to describe synonymous expressions, candidate names, or domain-equivalent names that may correspond to entities. Each semantic alignment candidate can be represented as:
- (3)
- denotes the set of port name candidates. It is used to describe possible naming forms corresponding to the required ports. Each port name candidate can be represented as:
- (4)
- A = {} denotes the set of function–entity allocation relations. It is used to describe the allocation and traceability relationships between functions in CFS and entities in TSE. Each allocation relation can be represented as:
For instance, the abstract function “Generate Thermal Energy” can be associated with the entity Heater. This entity requires ports such as command input, energy input, return-water input, and heated-water output, and includes parameter constraints such as minimum heating power and compatible water-flow conditions.
The descriptions and usages of the elements in TOFS are summarized in Table 2.
Table 2.
Definition of TOFS Elements.
3.2.2. LLM-Assisted Construction of TOFS
The construction of TOFS is formulated as an LLM-assisted human-in-the-loop process, which is divided into three subtasks: extraction, inference, and generation. The process is shown in Figure 3.
Figure 3.
LLM-assisted construction of TOFS process.
(1) The extraction task refers to identifying explicitly existing information from SysML requirement models and functional models, and converting it into structured fields. Since all information in CFS is derived from the requirement and functional models, with only some fields described in natural language, the LLM is therefore used to extract candidate CFS elements according to the TOFS schema and extraction rules, while human engineers verify their completeness and correctness. The extraction task can be defined as:
where denotes the SysML requirement model, denotes the SysML functional model, denotes the schema of the CFS output, denotes the set of rules that should be followed in the extraction task, and denotes the structured CFS obtained through extraction.
(2) The inference task refers to deriving engineering semantic information that is not explicitly given in the input models but can be reasonably supported by the functional semantics under the constraints of CFS. Since Entities represent abstract logical elements required to realize system functions, they are not explicitly provided in the requirements or functional models, but need to be inferred based on functional requirements, functional flows, and their contextual relationships. Because this task involves logical-level architecture decisions, such as n-m relations in function–entity allocation and preferred entity selections, the inferred results are not directly accepted as final outputs. The LLM is used to propose candidate Entities and function–entity relations, while human engineers review, refine, and confirm them. The inference task can be defined as:
where is obtained from the preceding extraction task; and are the same as defined above; denotes the schema of the inferred Entities output; denotes the set of rules that should be followed in the inference task; and denotes the structured set of Entities obtained through inference.
The identification of entities is not a trivial task, because it involves logical-level architecture decisions. A function may be realized by multiple entities, and multiple functions may also be assigned to an integrated entity. Therefore, human engineers are involved in reviewing and confirming candidate entities and function–entity relations proposed by the LLM. For example, for the function Circulate hot water, the LLM may propose several candidate entities, such as Pump, Thermosiphon Loop, and Integrated Boiler-Pump Unit. These candidates represent different technical realization strategies and corresponding function–entity mappings. Based on the specified flow-rate and head constraints in the requirements, the engineer finally confirms Pump as the entity used in this system because it directly satisfies the water-transport requirements with clear traceability and limited architectural coupling.
(3) The generation task refers to generating candidate expressions around existing entities and port semantics. The off-the-shelf LLM is used only to propose possible terminology variants and domain-equivalent expressions using its general knowledge. These candidates are reviewed and adjusted by human engineers before being used for component retrieval. The generation task can be defined as:
where is obtained from the preceding extraction task, is obtained from the preceding inference task, denotes the schema of the generated entity and port name candidates, denotes the set of rules that should be followed in the generation task, and and denote the generated entity name candidates and port name candidates, respectively. Sample rules for the three tasks are shown in Table 3.
Table 3.
Sample rules.
The rationale for adopting the above element-aware task decomposition process is that different TOFS elements have different semantic sources and construction mechanisms. If all elements are generated as a whole in a single step, explicit information extraction, entity inference, and candidate name generation may become mixed with each other, thereby reducing the interpretability and verifiability of the construction results. In contrast, by decomposing the construction process into three subtasks, more explicit inputs, output formats, and constraint rules can be specified for different types of semantic fields. When deviations occur in the construction results, it also becomes easier to identify the source of the problem.
Although the task decomposition and construction rules help constrain the LLM-assisted TOFS construction process, several risks may still arise in different subtasks. In the extraction task, the LLM may miss or misinterpret information explicitly contained in the SysML requirement and functional models. In the inference task, the LLM may introduce over-inferred entities, inappropriate entity granularity, or unreasonable mappings between functions and entities. In the generation task, the LLM may introduce terminology bias when generating entity and port name candidates. Therefore, the proposed method adopts a human-in-the-loop construction process. However, human review can only reduce these risks rather than completely eliminate them. More systematic mitigation mechanisms will be investigated in future work.
3.3. Retrieval-Augmented Architecture Generation
Based on TOFS, this study follows the systematic process of conceptual design in engineering design theory to generate architectures. In this study, the hybrid retrieval strategy is adopted, which integrates interface matching, parameter constraint matching, and functional semantic matching. In this way, a candidate component set can be obtained that is not only functionally relevant but also satisfies the engineering constraints. In addition, in the architecture synthesis stage, a RAG mechanism is introduced. The engineering-feasible component space provided by retrieval and the engineering constraints provided by TOFS are organized as the context, enabling the LLM to generate high-quality feasible architectures.
3.3.1. Hybrid Component Retrieval
Let the component library be denoted as . For any component , it is represented as:
where denotes the semantic description of the component, including textual information such as the component name, functional description, and capability description; denotes the set of component interfaces; and denotes the set of component parameters. Furthermore, the interface set can be represented as:
where and denote the input interface set and output interface set of the component, respectively. The parameter set is represented as:
where denotes the parameter name, denotes the value of component with respect to this parameter, and denotes the parameter unit. The following sections provide detailed descriptions of the three matching tasks.
- Interface matching
Interface matching is used to determine whether candidate components in the component library can satisfy the port requirements of entities in TOFS. The criteria for interface matching can be formalized into three parts: direction consistency, type compatibility, and name matching. For a required port of an entity , if there exists a port in component that satisfies these three conditions simultaneously, the required port is considered to be successfully matched.
For any required port of entity ,
let the port set of component be:
and each component port be represented as:
(1) Direction matching. Direction matching is satisfied when the required port and the component port have the same direction. If either port is bidirectional, the two ports can also be regarded as direction-compatible. The matching rule is as follows:
(2) Type matching. Type matching is satisfied when the required port and the component port have the same or compatible types. Here, denotes a predefined set of interface type compatibility relations. For example, digital signal and data signal may be regarded as compatible according to domain rules. The matching rule is as follows:
(3) Name matching. Name matching is satisfied when the component port name belongs to the set consisting of the required port name and its candidate port names, or when its similarity to this set exceeds a given threshold. The matching rule is as follows:
where denotes the semantic similarity function for port names, and denotes the name matching threshold.
- Parameter constraint matching
Parameter constraint matching is used to determine whether candidate components in the component library satisfy the normalized constraints specified by entities in TOFS in terms of performance, physical properties, or engineering attributes. Parameter constraint matching can be formalized into three sub-judgments: parameter name alignment, unit consistency or convertibility, and constraint expression satisfaction.
Let the normalized parameter constraint set of entity in TOFS be:
where each parameter constraint is defined as:
Here, denotes the parameter name, denotes the constraint operator, denotes the constraint value or value range, and denotes the constraint unit. For any component in the component library, its parameter set is represented as:
where denotes the component parameter name, denotes the value of component with respect to this parameter, and denotes the component parameter unit.
(1) Parameter name alignment. For constraint , a corresponding component parameter first needs to be identified from the component parameter set . The matching rule is as follows:
where denotes the semantic similarity function for parameter names, denotes the name similarity threshold, and denotes a predefined set of synonymous or equivalent parameter relations.
(2) Unit consistency or convertibility. After parameter name alignment, it is necessary to determine whether the constraint unit and the component parameter unit are identical or convertible. The matching rule is as follows:
where indicates whether the component parameter unit can be converted into the constraint unit.
(3) Numerical constraint satisfaction. After parameter name alignment and unit unification are completed, the component parameter value is checked against the constraint expression in TOFS. The matching rule is as follows:
For example, for , if the component parameter is , parameter name alignment gives , and unit conversion gives . Then, since holds, this parameter constraint is satisfied.
- Semantic matching
Semantic matching uses the entity-specific natural-language semantic descriptions in TOFS to perform text-based similarity matching against the natural-language semantic descriptions of components in the component library, thereby determining whether a component is functionally suitable for undertaking a certain entity.
For an entity in TOFS, semantic matching mainly uses the following information to construct the query semantics:
where denotes the transformation fields of the set of functions allocated to according to the function–entity allocation relation A; denotes the content fields of the set of input and output flows associated with these functions.
The semantic description of each component in the component library is denoted as , which includes textual information such as the component name, functional description, and capability description. A text embedding model or LLM-based semantic scoring method is then used to calculate the similarity between the query semantics and the component semantic description, resulting in the semantic matching score . A higher score indicates that component is more likely to be functionally suitable for undertaking entity .
3.3.2. Architecture Synthesis Based on LLM
After obtaining candidate component sets through hybrid retrieval, the subsequent task is to synthesize the candidate components into a complete system architecture. The workflow of the proposed retrieval-augmented automated architecture generation method is shown in Figure 4.
Figure 4.
Workflow of LLM-based Architecture synthesis.
- Context construction
The goal of context construction is to transform the hybrid retrieval results into a structured generation context that can be used by the LLM. This context includes TOFS, the candidate component sets, and architecture generation rules. The context for architecture generation can be represented as:
where denotes the candidate component sets corresponding to all entities. For each component , is also recorded, indicating the retrieval matching score between component and entity . denotes the architecture generation rules, such as that components should preferably be selected from the candidate sets, connections must be established between real ports, and port directions and types should be consistent or compatible.
- Missing entity identification
Although hybrid retrieval returns candidate components for each entity, the candidate space may still be incomplete for many reasons. For example, some functional flows may lack intermediate transformation components, interface adaptation components may be missing between two candidate components, or supporting entities such as control, processing, communication, and power supply may not have been explicitly identified in the previous stage. Therefore, this step uses the semantic reasoning capability of the LLM to identify potential missing entities under the constraints of TOFS and the candidate component context. This step is analogous to identifying supporting functions in traditional functional decomposition.
The input of this step is the constructed in the previous stage. The output is a set of missing entities:
where each missing entity can be represented as:
Here, denotes the name of the missing entity, denotes the function it should undertake, denotes its required input and output ports, denotes the engineering constraints it should satisfy, and denotes the reason why this entity is identified.
The execution principle of this step is to let the LLM examine potential gaps between TOFS and the current candidate component context. In this study, missing entities are mainly identified from two types of gaps: boundary-flow gaps and physical interface gaps.
A boundary-flow gap occurs when a functional flow in TOFS interacts with the system boundary but lacks an explicit source, target, or contextual entity in the current architecture context. For example, a temperature setpoint flow requires a source that provides the user-defined setpoint to the controller, but without an explicit function in the original functional model to represent it; hence, a user interface entity is identified in this step.
A physical interface gap occurs after candidate physical components have been retrieved but have incompatible physical ports. For example, a controller candidate may provide a 5 V digital control output, whereas a heater candidate may require a 24 V control input. Because of this voltage mismatch, an additional interface-adaptation entity, such as a signal-level converter, may need to be supplemented.
- Component selection and connection
This step determines the final selected components within the candidate component space and generates the connection relationships among components based on their real ports. Its input is the updated generation context , namely the candidate space formed by the original candidate component sets and the supplementary retrieval results for missing entities. The output is a system architecture model:
where denotes the set of finally selected components, and denotes the set of connection relationships among component ports.
The execution principle of this step consists of two stages. First, the LLM selects suitable components from the candidate set corresponding to each entity according to the retrieval matching scores, port satisfaction, and parameter constraint satisfaction of the candidate components. Second, based on the flows in TOFS, the LLM maps system-level functional interactions into component-level port connections. For any functional flow , the LLM needs to identify port connections among the selected components that can realize this functional flow, while satisfying conditions such as consistent port directions, compatible port types, and alignable port names or candidate names.
Through the above process, the proposed method finally generates a structured system architecture consisting of a component set and a port connection set. Since the architecture is derived from the candidate component library and constrained by the functional semantics, interface constraints, and parameter constraints in TOFS, it is expected to improve engineering validity and functional completeness compared with direct LLM-based architecture generation from raw textual inputs.
Before the generated architecture is accepted as the output of the workflow, a preliminary check is conducted as an early screening step. This step checks whether the required entities are covered by selected components, whether the main functional flows can be mapped to component-level port connections, whether the selected components are available in the candidate component library, whether the generated connections satisfy port direction and type compatibility, and whether the selected components satisfy the parameter constraints. If all the checks are passed, the generated architecture is accepted as the output architecture. If one or more checks fail, the failed check items are used by engineers to trace the source of inconsistency and adjust the corresponding inputs or constraints before rerunning the relevant generation process when necessary.
3.4. Case Description and Experimental Setup
3.4.1. Case System and Input SysML Models
A case study is used to show how the proposed approach is applied. This case study considers a closed-loop hydronic heating control system (HHCS) for indoor thermal regulation, which is illustrated in Figure 5. The system is a representative cyber-physical system that integrates temperature sensing, control decision-making, heat generation, thermal-fluid circulation, heat buffering, and heat release. It interacts with the indoor environment, the user-defined temperature setpoint, and the external energy supply. Therefore, it provides a suitable case for evaluating the transformation from SysML requirements and functional models to component-level physical architecture.
Figure 5.
Illustrative view of the HHCS system.
Two SysML models are constructed as the input MBSE artifacts for the proposed architecture generation process. The requirement model, shown in Figure 6, provides system-level information, including the system objective, functional requirements, performance requirements, and engineering constraints. The functional model, shown in Figure 7, provides process-level information about how the system objective is functionally realized.
Figure 6.
Requirement model of the HHCS system.
Figure 7.
Functional model of the HHCS system.
3.4.2. Experiment Implementation Setup
The experiments use a self-constructed heating-domain component library containing 100 candidate components. The library covers the main component categories required by the hydronic heating control system, covering sensors, controllers, heaters, buffer tanks, circulation pumps, radiators, user interfaces, and power supply units. Each component is represented by its name, functional description, input and output ports, and key engineering parameters. These fields correspond to the three retrieval dimensions used in this study: functional descriptions support semantic matching, ports support interface matching, and parameters support constraint satisfaction checking.
To make the retrieval task non-trivial, the library includes both feasible and infeasible alternatives for each entity. Some components are functionally similar but do not satisfy required interface or parameter constraints, while others satisfy partial engineering constraints but are semantically less suitable. Therefore, the component library provides a controlled basis for evaluating whether the proposed hybrid retrieval method can retrieve components that are not only functionally relevant but also engineering-feasible. These components are imported into a Neo4j knowledge graph to support component retrieval. GPT-5.5 (OpenAI, San Francisco, CA, USA) is used as the LLM.
In this study, Magic Systems of Systems Architect (Dassault Systèmes, Vélizy-Villacoublay, France) was used as the SysML modeling environment to create the input requirements and functional models and to view the generated architecture models. The interface between the modeling environment and the LLM is implemented through an XMI–JSON–XMI workflow. REQ and ACT are first exported as XMI files. A preprocessing script extracts key elements from the XMI files and converts them into structured JSON. After architecture generation, the LLM outputs a structured architecture JSON containing selected components, ports, and port-level connections. A postprocessing script then converts this JSON into XMI, including model elements, their attributes, and explicit relationships, rather than diagram layout or other graphical presentation information, which can be opened in the SysML modeling environment. The imported model data constitute the generated architecture model, while the BDD and IBD diagrams shown in this study are manually created and arranged from these imported model elements for visualization and inspection.
3.4.3. TOFS Scenarios Design
To evaluate the performance of different retrieval methods under diverse system requirements, 24 TOFS scenarios (https://github.com/fuchenghui001-cloud/sysml-senarios accessed on 23 September 2026) covering various typical control requirements in heating systems are designed in this study, including rapid heating, steady-state temperature control, low-noise operation, energy-saving-oriented control, disturbance-resistant control, and multi-scenario switching. To further analyze the influence of semantic complexity and constraint complexity on retrieval performance, the TOFS scenarios are divided into four categories: basic control scenarios, parameter-constrained scenarios, semantically abstract scenarios, and complex combined scenarios. The scenario categories and samples are shown in Table 4. Among them, Category A scenarios mainly describe basic control tasks in heating systems, where the requirement objectives are relatively explicit. Category B scenarios emphasize constraints such as parameter ranges, interface compatibility, or operating boundaries. Category C scenarios place more emphasis on high-level semantic objectives such as comfort, robustness, and adaptability. Category D scenarios simultaneously contain multiple control objectives and constraint conditions, representing complex combined tasks that are closer to real system design requirements.
Table 4.
Categories and samples of TOFS scenarios.
For each TOFS scenario, a corresponding ground-truth component set is manually annotated to represent the components that are expected to be retrieved under the given requirement conditions. This ground-truth set is then used as the reference for computing the evaluation metrics.
3.4.4. Baselines and Metrics of Component Retrieval
To evaluate the performance of different component retrieval strategies, the following four methods are implemented and compared.
- (1)
- S-only retrieves components based solely on semantic similarity. It ranks components according to the embedding-based similarity between the TOFS description and the natural-language descriptions of components, and returns the top- retrieval results.
- (2)
- R-only performs retrieval based solely on interfaces and constraints, i.e., rule-based matching. It filters the component library according to interface and parameter constraints matching rules, without using semantic similarity information.
- (3)
- S→R first retrieves components using semantic similarity, and then filters or re-ranks the recalled candidates according to interface matching and parameter constraints.
- (4)
- R→S first filters the component library using interface matching and parameter constraints to narrow the candidate space, and then ranks the remaining components based on semantic similarity. This strategy is also adopted in the case study.
Precision, Recall, and F1 are adopted as the main evaluation metrics. To assess the overall performance of different methods, macro-averaged metrics are further calculated over the 24 TOFS scenarios. Specifically, Precision, Recall, and F1 are first computed for each scenario, and then averaged across all scenarios. This metric provides a balanced characterization of the average performance of each method across diverse scenarios.
3.4.5. Ablation Settings and Metrics of Architecture Generation
To evaluate the effects of different generation contexts on architecture generation, this study conducts an ablation study with four settings as shown in Table 5. The ablation focuses on two types of information: structured functional semantics provided by TOFS and retrieved component knowledge provided by the candidate component set (CCS). These settings are designed to analyze how TOFS and CCS separately and jointly affect the engineering validity and functional completeness of the generated architectures.
Table 5.
Ablation settings.
The four ablation settings are defined as follows.
Five metrics are used in this evaluation. The first three, i.e., Component Validity (CV), Connection Validity (CNV), and Parameter Satisfaction Rate (PSR), measure whether the generated architecture satisfies real component and constraint requirements in engineering practice, thereby reflecting Engineering Validity. The last two metrics, i.e., Entity Coverage (EC) and Flow Coverage (FC), measure whether the entities and flows defined in TOFS are covered, thereby reflecting Functional Completeness.
(1) Component Validity measures whether the generated components are real components in the component library, and thus reflects the degree of component hallucination. It is defined as:
where denotes the number of generated components that can be matched to components in the component library, and denotes the total number of generated components.
(2) Connection Validity measures whether the generated connections are established between real component ports and satisfy port direction and type compatibility constraints. It is defined as:
where denotes the number of connections that satisfy conditions such as component existence, port existence, direction compatibility, and type compatibility, and denotes the total number of generated connections.
(3) Parameter Satisfaction Rate measures whether the finally selected components satisfy the normalized parameter constraints defined in TOFS. It is defined as:
where denotes the number of satisfied parameter constraints, and denotes the total number of parameter constraints in TOFS that need to be checked.
(4) Entity Coverage measures whether the entities in TOFS are covered by components in the generated architecture. It is defined as:
where denotes the number of entities undertaken by at least one valid component, and denotes the total number of entities in TOFS.
(5) Flow Coverage measures whether the functional flows in TOFS are covered by the generated port-level connections. It is defined as:
where denotes the number of functional flows implemented through valid component connections, and denotes the total number of functional flows in TOFS.
In addition to the five individual metrics, a Weighted Comparative Performance Index (WCPI), which is commonly used to summarize multidimensional performance characteristics into a single comparative indicator [52], is introduced to provide a concise aggregated comparison. It is defined as:
In this study, the weights are assigned according to a group-based equal-weighting strategy. CV, CNV, and PSR jointly characterize engineering validity, and this metric group is assigned a total weight of 0.5. EC and FC jointly characterize functional completeness, and this metric group is also assigned a total weight of 0.5. Within each group, equal weights are used because no additional domain-specific priority among the metrics is assumed in this study.
4. Results
4.1. Case Study Results
4.1.1. TOFS Construction Results
Taking the requirement and functional models as inputs, the three subtasks of LLM-based extraction, inference, and generation are executed sequentially. It should be emphasized that the entities in TSE are not predefined in the input SysML models. Instead, they are inferred by the LLM and evaluated by engineers based on the extracted CFS information. In this example, 6 entities, i.e., TemperatureSensor, Controller, Heater, Tank, Pump, Radiator, are inferred, which correspond to functions f1 to f6, respectively. As a result, the structured TOFS shown in Table 6 can be automatically generated. This structure is consistent with the TOFS element definitions presented in Table 2.
Table 6.
TOFS of the HHCS system.
The LLM-assisted construction of TOFS may introduce several potential biases. First, prompt-induced bias may occur because the wording of prompts and task rules can influence the inferred TOFS elements. For example, if the prompt emphasizes “closed-loop hydronic circulation,” the LLM may be biased toward inferring a conventional Pump entity, while alternative realizations such as a Thermosiphon Loop or an Integrated Boiler-Pump Unit may be under-considered. Second, prior-knowledge bias may occur because the LLM tends to prefer common engineering terms or typical architecture patterns. For example, for indoor heating control, the LLM may naturally prefer common entities such as TemperatureSensor, Controller, Heater, Tank, Pump, and Radiator, although other domain-specific decompositions may also be feasible. Third, terminology bias may affect the generated entity and port name candidates, thereby influencing semantic matching and interface matching in the retrieval stage. For example, the input port of the TemperatureSensor may be expressed as room_air_in, ambient_air_in, indoor_air_in, or air_in; if the candidate set does not cover the terminology used in the component library, relevant sensor components may not be retrieved. Finally, granularity bias may occur when the LLM decomposes entities too coarsely or too finely. For example, the heat-generation and circulation functions may be represented as separate Heater and Pump entities, or may be merged into an Integrated Boiler-Pump Unit, leading to different function–entity allocation relations and retrieval results.
To mitigate these biases, this study does not treat the LLM output as the final TOFS directly. Instead, the LLM is used to propose candidate TOFS elements, while human engineers review and confirm them before they are used for component retrieval and architecture generation. For example, alternative entities for the hot-water circulation function are reviewed against the specified flow-rate and head constraints, and Pump is finally confirmed because it directly satisfies the water-transport requirement with clear traceability and limited architectural coupling.
However, this human-in-the-loop strategy also introduces limitations. First, manual review increases the effort required for TOFS construction and may limit the scalability of the method when larger systems or larger scenario sets are considered. Second, human review can reduce but cannot completely eliminate bias. More systematic bias-mitigation mechanisms, such as multi-prompt consistency checking and multi-model cross-checking, can be introduced in the future. Finally, the effectiveness of the review also depends on the completeness of the input SysML models and the coverage of the component library. Therefore, the proposed method should be understood as an LLM-assisted and engineer-reviewed construction process rather than a fully automated or bias-free TOFS generation method.
4.1.2. Components Retrieval Results
For each entity in TOFS, the proposed hybrid retrieval method can be used to retrieve, from the existing engineering component library, a set of components that can realize the corresponding entity. The Heater entity is used as an example to illustrate the retrieval process. The Heater entity requires the component to have a Heater_Command (HC) input port, an EnergyPort (EP) input port, and input/output ports of the ThermalFluidPort (TF) type. It also needs to satisfy parameter constraints such as rated power rated_power ∈ [800, 2500] W (P), outlet temperature outlet_temperature ≥ 70 °C (T), and flow rate flow_rate ≥ 4 L/min (Q). Meanwhile, the entity name and its candidate names, such as ImmersionHeater and ResistiveWaterHeater, are used to construct the semantic query.
In the interface matching stage, the system filters the component library according to port direction consistency, port type compatibility, and port name matching. For the Heater entity, Electric boiler components satisfying the above interface constraints are retained, whereas tanks, radiators, pumps, and other components lacking a control command port or an energy input port are filtered out. This stage narrows the candidate space from the entire component library to a set of electrically connected boiler-type candidates.
The parameter constraint matching stage is then performed. For the candidates that pass interface filtering, the system sequentially performs parameter name mapping, unit normalization, and numerical range checking. As shown in Table 7, for the Heater entity, ElectricBoiler_Mini_LowPower_0.5kW fails in this stage because p < 800 W. Only components that satisfy all three parameter constraints are passed to semantic ranking.
Table 7.
Retrieval results of Heater.
In the semantic matching stage, the system calculates the cosine similarity between the query semantics and the description vectors of the retained candidate components. The candidate set that has passed interface and parameter filtering is then semantically ranked, so that the component most consistent with the functional semantics can be further identified among engineering-feasible components. Since semantic matching is applied only to candidates that already satisfy the interface and parameter constraints, this process avoids introducing invalid components that are functionally relevant but fail to satisfy interface or parameter requirements. As shown in Table 7, for the Heater entity, ElectricBoiler_Inline_2kW_ClosedLoop_V1 satisfies all interface and parameter constraints and ranks first with a semantic score of 0.91. The stepwise filtering and ranking results are shown below.
The above retrieval process is repeated for the 6 core entities in TOFS, forming the constrained component space for subsequent RAG-based architecture generation. Table 8 lists the top-1 component retrieved for each of the 6 entities, together with its constraint verification parameters.
Table 8.
Top-1 components of the 6 entities.
4.1.3. Architecture Generation Results
Based on the retrieved candidate component sets, the architecture model of the HHCS system can be generated using the proposed retrieval-augmented architecture generation approach. First, TOFS and the candidate component sets are integrated into the RAG context. A fragment of the context is shown in Figure 8, which is related to the Heater entity.
Figure 8.
Context fragment.
Subsequently, the LLM examines the candidate context and identifies the missing entities. In this process, three supplementary entities are identified: UserInterface, PowerSupply, and RoomAir. Based on these three supplementary entities, corresponding components can be further retrieved, and the architecture generation context is updated accordingly.
Finally, the updated context is input into the LLM. The LLM selects components and generates port-level connections and produces the BDD and IBD shown in Figure 9 and Figure 10, respectively.
Figure 9.
Generated BDD of the HHCS physical architecture.
Figure 10.
Generated IBD of the HHCS physical architecture.
4.2. Component Retrieval Evaluation Results
4.2.1. Overall Performance Comparison
Table 9 reports the evaluation results of different retrieval strategies. The R→S strategy achieves the best performance across Precision, Recall, and F1. This result indicates that, under the current experimental setting, first narrowing the candidate space through rule-based matching and then applying semantic ranking can improve retrieval effectiveness. In contrast, S-only performs relatively poorly, suggesting that semantic similarity alone is insufficient to satisfy interface and parameter constraints. R-only already shows strong competitiveness, demonstrating the effectiveness of rule-based information. The performance of S→R is lower than that of R→S, indicating that introducing semantic recall at an early stage may bring in candidates that do not satisfy engineering constraints, thereby degrading the overall retrieval performance.
Table 9.
Performance comparison of different retrieval strategies.
4.2.2. Performance Analysis Across TOFS Scenario Types
To further examine how the methods perform under different types of requirements, this study computes macro-averaged F1 results for the four categories of TOFS scenarios. As shown in Table 10, clear performance differences can be observed across scenario types.
Table 10.
Macro-averaged F1 results of the TOFS scenario categories.
In Category A, which consists of basic control scenarios, R-only achieves the best performance, followed by R→S. This suggests that rule-based matching is sufficient to characterize requirements when the control objectives are explicit. In Category B, which focuses on parameter-constrained scenarios, R-only and R→S achieve comparable performance, indicating that parameter and interface constraints remain the dominant discriminative factors.
In Category C, which contains semantically abstract scenarios, R→S achieves the best performance, outperforming R-only and S-only significantly. This demonstrates the importance of semantic information in retrieving components for high-level requirements. In Category D, which consists of complex combined scenarios, the hybrid methods generally outperform single-source methods, with S→R and R→S. This indicates that integrating multiple types of information is more suitable for complex requirements.
Overall, rule-based retrieval is effective for simple scenarios with explicit constraints, whereas hybrid retrieval shows more stable performance in semantically abstract and complex combined tasks.
4.3. Ablation Results on Architecture Generation
Table 11 reports the macro-averaged results of different ablation settings over the 24 TOFS scenarios. Since the relative performance of the ablation settings across different types of TOFS scenarios is largely consistent with the overall results, and no clear scenario-type-specific differences are observed, the architecture generation metrics for each scenario category are not separately analyzed.
Table 11.
Performance comparison of different ablation settings.
First, the comparison between LLM-only and LLM+CCS, as well as between LLM+TOFS and LLM+TOFS+CCS, shows the contribution of CCS. It can be observed that introducing CCS substantially improves the engineering validity of the generated architectures. Compared with LLM-only, LLM+CCS increases Component Validity from 0.29 to 0.97, Connection Validity from 0.25 to 0.76, and Parameter Satisfaction Rate from 0.14 to 0.54. Similarly, compared with LLM+TOFS, LLM+TOFS+CCS increases Component Validity from 0.33 to 0.99, Connection Validity from 0.24 to 0.83, and Parameter Satisfaction Rate from 0.15 to 0.66. This indicates that the real component, port, and parameter information contained in the retrieved candidate component sets can effectively constrain the generation space of the LLM, reducing component hallucination and unsupported connections, and thereby improving the engineering usability of the generated architectures.
Second, the comparison between LLM-only and LLM+TOFS, as well as between LLM+CCS and LLM+TOFS+CCS, shows the contribution of TOFS. The introduction of TOFS significantly improves the functional completeness of the generated architectures. Compared with LLM-only, LLM+TOFS increases Entity Coverage from 0.93 to 0.99 and Flow Coverage from 0.64 to 0.91. Compared with LLM+CCS, LLM+TOFS+CCS increases Entity Coverage from 0.83 to 0.98 and Flow Coverage from 0.57 to 0.92. These results show that the explicitly defined functions, flows, and entities in TOFS can effectively guide the LLM to cover the required system functions and generate a more complete architecture structure. In contrast, providing only the candidate component set can improve component validity, but cannot ensure that the LLM fully understands and covers the functional relationships of the system.
Overall, LLM+TOFS+CCS substantially outperforms the other three settings. This demonstrates the complementary entities of TOFS and CCS in architecture generation: TOFS provides functional semantics and entity-level constraints, while CCS provides real component, port, and parameter information. Combining the two can improve both functional completeness and engineering validity. However, the scores of the proposed method still do not reach 1, indicating that certain limitations remain. This is mainly because the current method adopts a one-shot generation strategy, and the LLM may overlook some key information in TOFS or in the candidate component sets during generation, such as failing to cover certain entities, missing some functional flows, or not strictly using the ports and parameters provided by the candidate components. In addition, the current method does not yet incorporate explicit component combination and optimization mechanisms. As a result, locally unreasonable connections or partially unsatisfied parameter constraints may still occur.
5. Discussion
The proposed method can be regarded as an integration of traditional engineering design theory and LLM-based generation, and is further embedded into the RFLP process of MBSE to support the transformation from requirement and functional models to component-level physical architectures. It builds upon the formal representation and systematic design process of traditional engineering design theory, while leveraging the reasoning and generation capabilities of LLMs. In this way, it partially addresses two limitations of existing approaches: the high modeling cost of traditional functional representation methods, and the tendency of LLM-based methods to suffer from component hallucination and limited interpretability. Based on the case study and evaluation results, the proposed approach has the following advantages compared with existing studies.
- (1)
- Enabling LLM-assisted construction of functional semantics for engineering tasks. Compared with existing formalized functional representations reported in previous studies [20,21,22], TOFS preserves a similar function-flow structure, but represents detailed functional semantics in natural language and augments them with task-oriented engineering information. This design facilitates LLM-assisted construction from semi-structured SysML models using LLMs, while still providing sufficient constraints to support effective component retrieval and architecture generation.
- (2)
- Constraining open-ended LLM generation with TOFS and retrieved component knowledge. Compared with direct LLM-driven methods such as [33,44,45], TOFS and the retrieved candidate component set are jointly used as the generation context in the proposed method. TOFS provides functional semantics and engineering constraints, while the candidate component set provides a component space grounded in an engineering component library. Their combination guides the LLM toward functionally complete and engineering-feasible architecture solutions, thereby reducing component hallucination and incomplete functional coverage.
- (3)
- Providing a systematic and inspectable generation process. Compared with approaches such as [37,38,39] that mainly focus on end-to-end generation from requirements to SysML model outputs, the proposed method decomposes it into systematic steps with explicit intermediate results. This makes the generation process more traceable and inspectable.
Although the proposed method achieves promising results, it still has the following limitations.
- (1)
- Lack of multi-turn validation and iterative refinement. In the current architecture generation process, all the LLM tasks are performed in a one-shot manner. The method has not yet incorporated interactive automated feedback, multi-turn refinement, or validation-and-repair mechanisms. As a result, errors in component selection, connection generation, or local constraint satisfaction cannot be progressively detected and corrected, which may limit the quality and robustness of the generated architectures. In practical applications, engineers therefore need to use the failed checking results to diagnose the source of inconsistency and manually refine the corresponding inputs before rerunning the generation process.
- (2)
- Lack of architecture-level compatibility checking and design-space evaluation. In the current implementation, candidate components retrieved for different entities are directly provided to the LLM for architecture generation. However, component compatibility and architecture-level evaluation criteria are not explicitly modeled. As a result, the LLM tends to select the top-ranked component for each entity, rather than systematically comparing alternative component combinations. Therefore, the generated results should be regarded as feasible architectures satisfying basic functional and engineering constraints, rather than optimized solutions. Before being used for downstream detailed design or engineering implementation, the generated architecture still requires architecture-level evaluation, simulation-based analysis, trade-off comparison, and expert review.
In addition, practical deployment of the proposed method still faces challenges in scalability, stability, and reproducibility. First, although human review helps improve reliability, it also introduces manual effort. As the number of requirements, functions, and constraints increases, manual review and correction may become time-consuming and limit scalability. Second, the quality of the generated architecture depends on the quality of the input SysML models and component library. Incompleteness and ambiguity may lead to incomplete or inaccurate results. Third, because LLM outputs can be sensitive to factors such as model versions, prompt settings and generation parameters, the generated results may vary across runs. Therefore, more standardized modeling inputs, curated component libraries, and controlled LLM configurations are needed to improve robustness and engineering assurance.
Emerging developments in systems engineering and digital engineering, especially the transition toward SysML v2 and integrated digital engineering environments, may help address some of the above limitations in two ways. First, at the model-quality level, SysML v2 provides a more formal, precise, and machine-processable modeling foundation. It can make input models and engineering libraries more explicit and consistent, which may reduce ambiguity in TOFS construction and improve component retrieval quality. Second, at the implementation level, integrated digital engineering environments and SysML v2-based model services may improve the current XMI–JSON–XMI workflow. Instead, the proposed approach could be implemented as a service-oriented architecture, in which standardized model APIs and service interfaces could support direct access to SysML models, component repositories, and LLM-based generation modules. This would lead to a more controlled and traceable workflow.
The proposed method also has potential implications for broader socio-technical systems. Its core logic, including semantic structuring in TOFS, retrieval-based grounding, and constrained LLM-based synthesis, can be adapted to other complex systems. However, applying the method to socio-technical systems requires extensions in at least three aspects. First, the notion of components should be broadened from physical components and software modules to heterogeneous architecture elements, including human roles, organizations, regulations, policies, and service capabilities. Second, TOFS should be extended to represent not only technical functions and constraints, but also broader socio-technical semantics such as stakeholder objectives, human–organizational activities, and institutional and environmental constraints. Third, the validation criteria should be expanded beyond functional completeness and engineering validity to include other socio-technical evaluation criteria such as stakeholder satisfaction, organizational feasibility, and regulatory compliance.
6. Conclusions
In the RFLP architecture design process, one key challenge is how to transform upstream requirements and functional models into component-level physical architectures while reducing manual functional-semantic modeling effort and maintaining engineering validity. This study proposed a retrieval-augmented architecture generation approach based on TOFS. TOFS is proposed to bridge semi-structured SysML artifacts and physical architecture generation. Based on it, engineering-feasible candidate components can be retrieved. TOFS and the retrieved candidate component sets are jointly employed to guide the LLM-based architecture generation process.
The case study and scenario-based experiments demonstrate the feasibility and effectiveness of the proposed approach. The component retrieval results show that combining rule-based filtering with semantic ranking provides more robust retrieval performance than using either semantic similarity or rule-based matching alone. The architecture generation results further show that TOFS and the retrieved candidate component sets play complementary roles: TOFS improves functional completeness, while the candidate component sets improve engineering validity by grounding the generation process in real components. These findings indicate that the proposed method can reduce the manual effort required for formal functional-semantic modeling, while still maintaining generation quality by grounding LLM-based architecture generation in explicit functional semantics and retrieved engineering component knowledge.
Future work will focus on three directions. First, multi-turn validation and iterative refinement will be introduced to progressively check and revise generated architectures. Second, architecture-level compatibility checking and design-space evaluation will be incorporated to support the optimization of alternative component combinations. Third, the proposed method will be extended to more complex engineering systems with larger-scale component libraries and richer domain constraints.
Author Contributions
Conceptualization, C.F.; methodology, C.F., Y.C. and Y.L.; software, C.F.; validation, Y.C. and Y.L.; data curation, C.F.; writing—original draft preparation, C.F.; writing—review and editing, Y.C. and Y.L.; supervision, Y.C.; project administration, Y.L.; funding acquisition, Y.C. All authors have read and agreed to the published version of the manuscript.
Funding
This research was funded by the Natural Science Foundation of Zhejiang Province, grant number LMS26F020022.
Data Availability Statement
Data available on request due to restrictions. The data presented in this study are available on request from the corresponding author.
Conflicts of Interest
Author Yu Liu was employed by the company China State Shipbuilding Corporation. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
References
- Madni, A.M.; Sievers, M. Model-based systems engineering: Motivation, current status, and research opportunities. Syst. Eng. 2018, 21, 172–190. [Google Scholar] [CrossRef] [Scilit]
- Kleiner, S.; Kramer, C. Model Based Design with Systems Engineering Based on RFLP Using V6. In Smart Product Engineering; Lecture Notes in Production Engineering; Springer: Berlin/Heidelberg, Germany, 2013; pp. 93–102. [Google Scholar]
- Weilkiens, T.; Lamm, J.G.; Roth, S.; Walker, M. Model-Based System Architecture; John Wiley & Sons: Hoboken, NJ, USA, 2022. [Google Scholar]
- Jankovic, M.; Eckert, C. Architecture decisions in different types of complex systems. Artif. Intell. Eng. Des. Anal. Manuf. 2016, 30, 217–234. [Google Scholar]
- Sievers, M. (Ed.) Semantics, Metamodels, and Ontologies; Spinger: Cham, Switzerland, 2023. [Google Scholar]
- Zhang, W.; Cockburn, C.; Henshaw, M.; Douglas, P.; Palmer, P.; Olivier-Myall, J.; Ji, S. MBSE Co-Pilot: A Research Roadmap. Syst. Eng. 2025, 29, 20–33. [Google Scholar] [CrossRef] [Scilit]
- ISO IEC IEEE 42010-2022; Software, Systems and Enterprise—Architecture Description. ISO: Geneva, Switzerland; IEC: Geneva, Switzerland; IEEE: Piscataway, NJ, USA, 2022.
- NASA (Ed.) NASA Systems Engineering Handbook; NASA: Washington, DC, USA, 2007.
- Estefan, J.A. Survey of Model-Based Systems Engineering (MBSE) Methodologies; INCOSE MBSE Focus Group: Seattle, WA, USA, 2008. [Google Scholar]
- Lykins, H.; Friedenthal, S.; Meilich, A. Adapting UML for an object oriented systems engineering method (OOSEM). In Proceedings of the 10th International INCOSE Symposium, Minneapolis, MN, USA, 16–20 July 2000. [Google Scholar]
- Douglass, B.P. (Ed.) Harmony aMBSE Deskbook; IBM Corporation: Armonk, NY, USA, 2017. [Google Scholar]
- Borky, J.M.; Bradley, T.H. (Eds.) Effective Model-Based Systems Engineering; Springer: Cham, Switzerland, 2018. [Google Scholar]
- Mhenni, F.; Choley, J.-Y.; Penas, O.; Plateaux, R.; Hammadi, M. A SysML-based methodology for mechatronic systems architectural design. Adv. Eng. Inform. 2014, 28, 218–231. [Google Scholar] [CrossRef] [Scilit]
- Lemazurier, L.; Chapurlat, V.; Grossetête, A. An MBSE Approach to Pass from Requirements to Functional Architecture. In Proceedings of the IFAC-PapersOnLine, Toulouse, France, 9–14 July 2017; Elsevier: Amsterdam, The Netherlands, 2017; pp. 7260–7265. [Google Scholar]
- Granrath, C.; Kugler, C.; Michael, J.; Rumpe, B.; Wachtmeister, L. Generating Logical Architectures from SysML Behavior Models. Syst. Eng. 2025, 28, 762–778. [Google Scholar] [CrossRef] [Scilit]
- Vazquez-Santacruz, J.A.; Portillo-Velez, R.; Torres-Figueroa, J.; Marin-Urias, L.F.; Portilla-Flores, E. Towards an integrated design methodology for mechatronic systems. Res. Eng. Des. 2023, 34, 497–512. [Google Scholar] [CrossRef] [Scilit]
- Ma, J.; Wang, G.; Lu, J.; Zheng, X.; Yan, Y. Application of multi architecture modelling method in intelligent electric–vehicle design. Int. J. Prod. Res. 2025, 63, 5493–5511. [Google Scholar] [CrossRef] [Scilit]
- Beitz, W.; Pahl, G.; Grote, K. (Eds.) Engineering Design: A Systematic Approach, 3rd ed.; Springer: London, UK, 2007. [Google Scholar]
- Hirtz, J.; Stone, R.B.; McAdams, D.A.; Szykman, S.; Wood, K.L. A functional basis for engineering design: Reconciling and evolving previous efforts. Res. Eng. Des. 2002, 13, 65–82. [Google Scholar] [CrossRef] [Scilit]
- Yuan, L.; Liu, Y.; Sun, Z.; Cao, Y.; Qamar, A. A hybrid approach for the automation of functional decomposition in conceptual design. J. Eng. Des. 2016, 27, 333–360. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.; Liu, Z.-L.; Xie, Y.-B. A knowledge-based framework for creative conceptual design of multi-disciplinary systems. Comput.-Aided Des. 2012, 44, 146–153. [Google Scholar] [CrossRef] [Scilit]
- Cao, Y.; Liu, Y.; Ye, X.; Zhao, J.; Gao, S. Software-physical synergetic design methodology of mechatronic systems based on formal functional models. Res. Eng. Des. 2020, 31, 235–255. [Google Scholar] [CrossRef] [Scilit]
- Yildirim, U.; Campean, F.; Uddin, A. Function modeling in model-based systems engineering using flow heuristics. Artif. Intell. Eng. Des. Anal. Manuf. 2025, 39, e26. [Google Scholar] [CrossRef] [Scilit]
- Chen, R.; Chen, C.-H.; Liu, Y.; Ye, X. Ontology-based requirement verification for complex systems. Adv. Eng. Inform. 2020, 46, 101148. [Google Scholar] [CrossRef] [Scilit]
- Chen, R.; Liu, Y.; Fan, H.; Zhao, J.; Ye, X. An integrated approach for automated physical architecture generation and multi-criteria evaluation for complex product design. J. Eng. Des. 2019, 30, 63–101. [Google Scholar] [CrossRef] [Scilit]
- Cao, Y.; Liu, Y.; Qin, X. A hybrid approach to system verification in early design for complex mechatronic systems based on formal functional semantics. Adv. Eng. Inform. 2023, 58, 102201. [Google Scholar] [CrossRef] [Scilit]
- Cao, Y.; Liu, Y.; Ye, X.; Zhao, J. An Automated Approach for Execution Sequence-Driven Software and Physical Co-Design of Mechatronic Systems Based on Hybrid Functional Ontology. Comput.-Aided Des. 2021, 131, 102942. [Google Scholar] [CrossRef] [Scilit]
- Wang, R.; Wei, Z.; Li, H.; Wang, Z.; Huang, Y.; Wang, G. Intelligent design exploration method for complex engineered system architecture generation. J. Eng. Des. 2024, 36, 797–835. [Google Scholar] [CrossRef] [Scilit]
- Della Bella, E.; Jankovic, M.; Pommier-Budinger, V.; Delbecq, S.; Jezegou, J.; Lefebvre, A. Systematic Architecture Generation for the Design of Aircraft Propulsive Systems. J. Mech. Des. 2026, 148, 112002. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Q.; Liu, J.; Li, L.; Chen, X.; Wang, R. Automatic generation of system model diagrams driven by multi-source heterogeneous data. J. Eng. Des. 2024, 35, 1442–1486. [Google Scholar] [CrossRef] [Scilit]
- Huang, Y.; Wang, R.; Li, Y.; Wang, G.; Yan, Y. Knowledge graph-driven methodology for complex product architecture solution generation and simulation verification. Adv. Eng. Inform. 2025, 68, 103590. [Google Scholar] [CrossRef] [Scilit]
- Wang, R.; Liu, J.; Xie, Y.; Yu, K.; Song, Z. Facilitating police equipment design via retrieval-augmented generation with knowledge graphs and small language models. J. Eng. Des. 2026, 1–31. [Google Scholar] [CrossRef] [Scilit]
- Apvrille, L.; Sultan, B. System Architects are not alone Anymore-Automatic System Modeling with AI. In Proceedings of the MODELSWARD 2024: 12th International Conference on Model-Based Software and Systems Engineering, Rome, Italy, 21–23 February 2024; INSTICC: Lisboa, Portugal, 2024; pp. 27–38. [Google Scholar]
- Schleifer, S.; Lungu, A.; Kruse, B.; Goetz, S.; Wartzack, S. Large Language Model-based Generation of Use Case Diagrams from Requirement Specifications. INCOSE Int. Symp. 2025, 35, 261–275. [Google Scholar] [CrossRef] [Scilit]
- Rafique, K.A.; Shah, S.; Dalecke, Š.; Grimm, C. Enhancing Model-Based Systems Engineering with Large Language Models. INCOSE Int. Symp. 2025, 35, 1523–1543. [Google Scholar] [CrossRef] [Scilit]
- Johnson, T.; Williams, A. Automated Legacy Documentation to SysML Conversion. INCOSE Int. Symp. 2025, 35, 1544–1563. [Google Scholar] [CrossRef] [Scilit]
- Cibrián, E.; Olivert-Iserte, J.; Llorens, J.; Álvarez-Rodríguez, J.M. An agent-based approach for the automatic generation of valid SysMLv2 Models in industrial contexts. Comput. Ind. 2025, 172, 104350. [Google Scholar] [CrossRef] [Scilit]
- An, B.; Lei, T.; Tian, G. An LLM-Based Framework for the Automatic Generation of SysML Models. Sensors 2026, 26, 5133. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bouamra, Y.; Yun, B.; Poisson, A.; Armetta, F. SysTemp: A Multi-agent System for Template-Based Generation of SysML V2. In Advances in Practical Applications of Agents, Multi-Agent Systems, and Computational Social Science: The PAAMS Collection; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2026; pp. 41–52. [Google Scholar]
- Sultan, B.; Apvrille, L. AI-Driven Consistency of SysML Diagrams. In Proceedings of the ACM/IEEE 27th International Conference on Model Driven Engineering Languages and Systems, Linz, Austria, 22–27 September 2024; Association for Computing Machinery: New York, NY, USA, 2024; pp. 149–159. [Google Scholar]
- Sugawara, K.; Komatsu, Y.; Wada, A. Extracting Information from System Model as Graph Structure by Large Language Model in MBSE. INCOSE Int. Symp. 2025, 35, 496–517. [Google Scholar] [CrossRef] [Scilit]
- Bonner, M.; Zeller, M.; Schulz, G.; Savu, A. LLM-based Approach to Automatically Establish Traceability between Requirements and MBSE. INCOSE Int. Symp. 2024, 34, 2542–2560. [Google Scholar] [CrossRef] [Scilit]
- Obieke, C.C.; Bridgeman, J.; Han, J. A framework of AI collaboration in engineering design (AICED). Proc. Des. Soc. 2025, 5, 91–100. [Google Scholar] [CrossRef] [Scilit]
- Von Heissen, O.; Hanke, F.; Mpidi Bita, I.; Hovemann, A.; Dumitrescu, R. Toward intelligent generation of system architectures. In Proceedings of the NordDesign 2024, Reykjavik, Iceland, 12–14 August 2024; The Design Society: Glasgow, UK, 2024; pp. 504–513. [Google Scholar]
- Timperley, L.R.; Berthoud, L.; Snider, C.; Tryfonas, T. Assessment of large language models for use in generative design of model based spacecraft system architectures. J. Eng. Des. 2025, 36, 550–570. [Google Scholar] [CrossRef] [Scilit]
- Yuan, W.; Wang, K.; Lu, J.; Li, W.; Luo, W.; Liu, Y.; Liang, Z.; Niu, B. A Model-Based Systems Engineering- and Large Language Model-Driven Method for Intelligent Generation of Aerial Bomb Design Solution. J. Comput. Inf. Sci. Eng. 2026, 26, 071002. [Google Scholar] [CrossRef] [Scilit]
- Harel, D.; Rumpe, B. Meaningful modeling: What’s the semantics of “semantics”? Computer 2004, 37, 64–72. [Google Scholar] [CrossRef] [Scilit]
- Stone, R.B.; Wood, K.L. Development of a Functional Basis for Design. J. Mech. Des. 2000, 122, 359–370. [Google Scholar] [CrossRef] [Scilit]
- Madni, A.M.; Augustine, N.; Sievers, M. (Eds.) Handbook of Model-Based Systems Engineering; Springer Nature: Cham, Switzerland, 2023. [Google Scholar]
- Carpineto, C.; Romano, G. A Survey of Automatic Query Expansion in Information Retrieval. ACM Comput. Surv. 2012, 44, 1–50. [Google Scholar] [CrossRef] [Scilit]
- Friedenthal, S.; Moore, A.; Steiner, R. (Eds.) A Practical Guide to SysML: The Systems Modeling Language; Morgan Kaufmann: San Francisco, CA, USA, 2014. [Google Scholar]
- Sepulveda-Fontaine, S.A.; Amigo, J.M. Applications of Entropy in Data Analysis and Machine Learning: A Review. Entropy 2024, 26, 1126. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.









