Skip to Content
ElectronicsElectronics
  • Article
  • Open Access

25 September 2026

25 Pages

An Evidence-Gated KCMVP Pre-Certification Decision-Support Prototype: Rule-Based Analysis, Evidence Retrieval, and Constrained LLM Review

,
,
,
,
,
and
Department of Convergence Security, Hansung University, Seoul 02876, Republic of Korea
*
Author to whom correspondence should be addressed.

Abstract

The Korean Cryptographic Module Validation Program (KCMVP) requires the joint review of source code, submission documents, and their traceability. This study presents an evidence-gated decision-support prototype combining rule-based candidate generation (L1), rule-bound evidence retrieval (L2), program-fact verification, and constrained LLM-assisted review (L3). Its 166 YAML rule assets, comprising 97 code rules, 65 document rules, and four traceability rules, were developed with reference to official KCMVP procedures and National Intelligence Service guidance on submission preparation and cryptographic algorithm implementation. On an LEA-centered development corpus, L1 matched 115 of 128 author-annotated positive file–rule pairs (89.84%). However, because these annotations were created during system development, independent expert validation is required before interpreting this value as detection accuracy. In nine source-derived mutation groups, the baseline and improved-evidence conditions each produced six binary dispositions and three holds (exact McNemar test, p = 1.0 ). This result defines the current role of L2 as supplying traceable evidence, validating citations, and withholding unsupported decisions rather than independently determining compliance. L3 reviews only residual candidates supported by validated evidence and program facts under a structured output contract; failure to satisfy input, evidence, citation, or output-validation conditions results in a hold. Overall, the results demonstrate a feasible and auditable architecture that preserves evidence provenance and explicitly holds unresolved cases instead of forcing unsupported decisions.

1. Introduction

The Korean Cryptographic Module Validation Program (KCMVP [1,2]) evaluates cryptographic modules used to protect important information that requires security protection but is not formally classified in Korean government and public sector information systems. The program operates within the national cybersecurity and electronic government framework. Accordingly, KCMVP validation examines not only whether a cryptographic module satisfies the applicable security requirements, but also whether its implementation of approved cryptographic algorithms and the submitted supporting evidence conform to those requirements. The validation process covers application, submission review, technical testing, and a final validation decision [1,3].
During technical testing, the testing body examines the implementation of the approved cryptographic algorithms, applicable security requirements, and the consistency of the submitted evidence. Source code and supporting documents are therefore related review objects rather than independent submissions. Requests for supplementary material may require correction, renewed testing, and resubmission. Repeated cycles can prolong validation and delay the subsequent deployment of a cryptographic module. However, KCMVP-specific tools for identifying potential compliance issues before official submission remain limited. This limitation highlights the need for an integrated pre-certification tool that supports early review and correction.
This study addresses this engineering gap with an evidence-gated decision-support prototype for KCMVP pre-certification review. It combines rule-based candidate generation, rule-bound evidence retrieval, program-fact verification, and constrained LLM-assisted review. The prototype supports pre-certification review and does not replace official validation. The present evaluation is limited to rule behavior, evidence traceability, and conservative hold behavior.

Main Contributions

The main contributions of this study are summarized as follows.
Evidence-gated architecture for multi-artifact review.
The primary contribution is an evidence-gated architecture that integrates source code, submission documents, and traceability artifacts within one KCMVP pre-certification workflow. It connects each L1 candidate to rule-bound evidence, rule-scoped program facts, and constrained L3 review while preserving source provenance across stages. This architecture provides an integrated basis for reviewing implementation and submission evidence before official validation.
Accountable use of RAG and LLMs through evidence control.
This study contributes a method for incorporating RAG and LLMs into security compliance review while preserving evidence traceability and decision accountability. It converts model assessments into review records that can be traced to rules, accepted evidence sources, and program facts. This approach provides a practical basis for using generative models under human oversight in security-sensitive pre-certification review.
KCMVP-specific rule registry and implementation foundation.
The prototype provides a KCMVP-specific registry of 166 YAML rule assets, comprising 97 code rules, 65 document rules, and four traceability rules. The registry organizes code, document, and traceability checks through explicit inspection targets and execution paths. It provides a concrete foundation for extending pre-certification rules and conducting controlled evaluations.

2. Background

2.1. Overview of the KCMVP System and Validation Procedure

KCMVP operates under Article 9 of the Cyber Security Work Regulations and Article 69 of the Enforcement Decree of the Electronic Government Act [2,4,5]. It evaluates cryptographic modules used to protect important information that is not formally classified in government and public sector information systems. Eligible modules may be implemented in software, hardware, firmware, or a combination of these forms.
KCMVP applies KS X ISO/IEC 19790 for cryptographic-module security requirements and KS X ISO/IEC 24759 for the corresponding test requirements [6,7]. The standards define four security levels, with requirements becoming progressively more stringent from Level 1 to Level 4. Level 1 is the baseline level and applies the basic requirements in each relevant security area. This study limits its implementation analysis to Level 1 software modules.
KCMVP covers multiple approved cryptographic algorithms. This study selects LEA as the initial evaluation target [8,9]. Future work will extend the same evaluation protocol to rule sets for other KCMVP-approved algorithms.
The public description of the validation process can be summarized in four stages [1,3]:
  • Application: The applicant submits the module and required evidence to an authorized testing body.
  • Submission Review: The testing body checks the completeness and applicability of the submitted materials.
  • Technical Testing: The testing body evaluates the module against the applicable security and test requirements.
  • Validation Decision: The validation authority reviews the test results and determines whether to grant validation.
Supplement requests can lead to correction, renewed testing, and resubmission. Repeated cycles can extend the validation period and delay subsequent module deployment. This risk motivates an auxiliary pre-certification review tool designed to identify potential nonconformities before official submission.

2.2. Related Work

Table 1 compares the prototype with representative cryptographic-misuse analyzers and the algorithm-testing service in the United States CMVP ecosystem. The design also draws on machine-readable compliance, LLM-assisted code analysis, RAG evaluation, selective classification, and indirect prompt-injection research.
Table 1. Comparative analysis with related prior work. ✔ = supported, × = not supported.
CryptoGuard uses interprocedural slicing to detect 16 classes of cryptographic API misuse in Java [10]. CogniCrypt provides IDE-integrated checks for correct use of the Java Cryptography Architecture [11]. These systems show that cryptographic-code inspection can be automated for defined misuse classes. They do not target KCMVP-specific requirements or LEA.
NIST’s Cryptographic Algorithm Validation Program focuses on algorithm testing [12]. NIST OSCAL shows how control information can be represented in machine-readable XML, JSON, and YAML [13]. Empirical research on LLM code analysis reports both useful capabilities and material limitations [14]. Selective classification formalizes a reject option that exchanges automated coverage for lower decision risk [15]. Calibration research also shows that a model’s raw confidence need not match its probability of correctness [16]. Indirect prompt injection demonstrates that retrieved or embedded data can behave as adversarial instructions in an LLM-integrated application [17].
Against this background, the contribution claimed here is deliberately narrow. The prototype links KCMVP-oriented code, document, and traceability candidates to rule-scoped evidence, program facts, conservative hold behavior, and an auditable decision record. It does not claim that related compliance or LLM-analysis tools do not exist.

3. System Design

The prototype separates rule-based candidate generation (L1), rule-bound evidence retrieval (L2), rule-scoped program-fact verification, and constrained LLM-assisted review (L3). This separation supports broad discovery of potential review targets while preventing uncertain cases from being automatically finalized and retaining them for human review.
The system has three design goals.
Goal 1. Broad discovery of review targets. Cryptographic modules may implement semantically equivalent operations using different function names, macros, wrapper structures, and file organizations. Checks that rely on only a narrow set of strings or a single code form may therefore miss relevant implementations. L1 searches broadly for multiple signals expressible as rules. The missing path checks for absent required patterns, while the regex path records matching textual patterns. The semantic and ast paths additionally inspect surrounding context and code structure, respectively. Collecting review targets broadly at this initial stage reduces the likelihood that potentially relevant locations are omitted before downstream analysis and establishes a recall-oriented preliminary review process.
Goal 2. Conservative disposition of review targets. Rule matching can capture similar code patterns, public constants, test code, or code analyzed under an incomplete build context. The system therefore routes each review target through domain anchors, program facts, evidence checks, and constrained LLM review. Suppression requires explicit decision conditions. If those conditions are not satisfied, the target is retained or assigned hold.
Goal 3. Evidence-linked inspection of code and documents. The framework processes source archives and submitted PDF documents within one review pipeline. Rule-linked guidance is attached to a review target only when the evidence contract is satisfied. Each evidence unit is classified as an official requirement, official explanatory material, author-interpreted guidance, or a derived implementation rule. A review target generated by the framework does not independently establish an actual violation. Instead, the framework records the applicable rule, the location of the original requirement, the relevant code or document location, and the applicability conditions of the evidence so that a reviewer can trace how the result was produced and which evidence supported it.

3.1. Overall System Architecture

Figure 1 shows the funnel-shaped pipeline. Source archives and submission documents first pass through an input-admission and preprocessing boundary. For source code, the system builds function summaries and a symbol graph that links calls across files. For documents, it extracts body text and tables and organizes them into sections. The resulting artifacts record the file, line, page, or section coordinates used by later rules.
Figure 1. Overall pipeline of the proposed prototype. The revised figure shows the decision-authorization gate, the API-free program-fact disposition path, the hold and reviewer-required path, and the restricted L3 path entered only by evidence-authorized residual targets.
L1 applies YAML rules using four pattern types. The missing type checks for absent required signals, and regex records textual pattern matches. The semantic type uses keywords and surrounding context, whereas ast uses syntactic and structural information. Before the results are consolidated as review targets, domain anchors refine clear mismatches between the detected location and the intended rule scope.
Retrieval is separated from decision authorization. The atomic-claim registry records each evidence unit’s identifier, digest, source location, rule linkage, and applicability. Retrieval proposes potentially relevant evidence but does not authorize a decision by itself. A proposed evidence bundle can be used for downstream assessment only when its units are correctly identified, bound to the original source, applicable to the target rule, and consistent with the citation contract. Incomplete, conflicting, unmapped, or differently authoritative evidence results in hold.
L1 produces locations requiring additional review rather than confirmed violations. After provenance, evidence, and program-fact checks, a review target can receive a deterministic disposition, remain in hold, or enter the restricted L3 path. A run without provider API access does not fabricate an L3 result. Every reported L3 result must pass output-schema, identifier-allowlist, and citation-relation validation.
The evidence module is used both to prepare the evidence supplied to a permitted LLM assessment and to attach accepted evidence identifiers and source locations to the final review record. This allows the final report to present the applicable rule, verified program facts, and supporting evidence together with each review target.
Table 2 summarizes the role of each stage.
Table 2. Per-stage summary of the proposed KCMVP pre-certification pipeline.

3.2. Per-Stage Design of the Validation Pipeline

3.2.1. Preprocessing Stage

The rule engine and constrained LLM path require stable file, line, page, and section locations. A review-target description and the context supplied to later stages must refer to the same input location. Structural and call information also provides program facts that cannot be reliably recovered through string matching alone. Preprocessing therefore converts source code into line-oriented text, function summaries, and call summaries, and converts documents into body-text, table, and section structures. Because these structured artifacts retain both their original locations and derived analysis information, a reviewer can trace where a target originated and which code or document context was used in a later disposition.
The implementation produces an analysis-provenance record that links the original input, preprocessing results, and analysis-tool information. It contains the original source-byte identity, preprocessing artifacts, and the compiler or parser path used. When rule applicability depends on an actual shipping target, the record also includes the declared target and archive-membership information. This separates facts observed directly from the submitted material, facts derived by analysis tools, and facts requiring additional reviewer confirmation.
A content-based cache retains file digests and preprocessing provenance across reruns. For analysis paths governed by the build-binding contract, an HMAC-protected record is used to check whether an analyzed file belongs to the declared shipping target. If membership cannot be resolved, the result is not converted into a binary disposition and remains in hold.
AST and Symbol Graph
An abstract syntax tree (AST) represents the syntactic structure of source code and is a standard output of compiler parsing [18]. It allows operations, conditions, loops, array accesses, and calls within a function to be examined according to syntax rather than textual appearance alone. The symbol graph links function definitions, calls, type information, and selected constant arrays across source files. It is used to inspect cross-file program facts such as wrapper delegation and calls to functions intended to erase sensitive information.
C and C++ preprocessing can change token structure before parsing. Directives such as #define, #include, and #ifdef may cause the post-preprocessing syntax to differ from the original source layout. The parser and rule engine also depend on type definitions, constant values, and macro expansions. Missing headers or build flags can therefore leave identifiers unresolved, shift source locations, or produce incomplete syntax trees.
The implementation combines compiler preprocessing, project-header tracking, libclang parsing, and pycparser-based fallback paths. Each review record identifies the parser path used. These paths preserve analysis availability under different levels of build information; they do not assert equivalence between parser results.
  • Stage 1. libclang-first parsing. The system first invokes LLVM libclang with the available compilation arguments [19]. When parsing succeeds, it records function definitions, cross-file calls, type information, selected array-initialization metadata, and an AST summary. Compiler-provided symbol identifiers are used to connect calls to definitions where available, reducing ambiguities caused by repeated names or separated declarations and definitions. Name-based matching is used only as a fallback. The completeness of this result still depends on whether the supplied headers, macros, and compilation flags adequately represent the target build; structural array information alone does not establish whether a stored value is secret.
  • Stage 2. Compiler preprocessing with pycparser. If libclang is unavailable or parsing fails, the system expands the C source with an available compiler preprocessor and analyzes the result with pycparser [20]. This path is intended to recover the function definitions, call locations, and basic declarations required for preliminary review, not to reproduce the complete build environment. Newline counts are preserved so that reported locations can be mapped back to the submitted source, and the review record identifies the fallback path and unresolved symbols. When required, a minimal declaration preamble and approximated declarations derived from project headers and macros are supplied; these approximations are explicitly recorded as analysis provenance rather than treated as the actual compiled program.
  • Stage 3. Declaration-synthesized preprocessing. When a compiler preprocessor is unavailable, the system follows project-local quoted includes, removes comments while preserving line counts, and synthesizes declarations for selected function-like macros. The approximated result is then analyzed with pycparser. This stage attempts to recover only the minimum information required for parsing and does not assume equivalence with the actual build output.
  • Stage 4. Minimal preprocessing. If project paths or headers cannot be resolved, the system retries using a fixed declaration preamble and declarations derived directly from the source file. As a final fallback, it replaces preprocessor-directive lines with blank lines and parses the comment-free body. Because this path has the most limited context, unresolved types and calls are explicitly recorded.
When one stage fails, the system records the failure and proceeds to the next fallback. If an inserted declaration preamble changes line positions, an offset is applied to map the analysis location back to the original source line. Listing 1 shows an example of the AST summary recorded during preprocessing.
Listing 1. Example JSON structure of an AST preprocessing result that records functions and call locations.
Electronics 15 04428 i001
Document Section Extraction
Document preprocessing separates running text from tables. pdfplumber detects table candidates and reconstructs their cells as two-dimensional arrays [21]. The system records each table’s bounding box. When PyMuPDF extracts body text, text blocks that substantially overlap a table are excluded from the body stream [22]. This procedure reduces duplicate extraction of the same content while preserving the row-and-column structure of tables.
After extracting body text and tables, the system divides the document into sections. It first attempts a table-of-contents-based path. If a usable table of contents is unavailable, it uses a heading-pattern fallback. The recovered section structure allows document review targets to be reported at both page and section levels.
  • Stage 1. Table-of-contents path. The system identifies a table of contents in the combined text and extracts entry identifiers, titles, and printed page numbers. It locates the first title in the body to estimate the offset between printed and physical PDF pages. Body text and tables within the corrected page range are assigned to the corresponding section. Repeated headers, footers, and isolated page numbers are removed when detected.
  • Stage 2. Heading-pattern path. If a usable table of contents cannot be recovered, regular expressions identify hierarchical heading candidates in the body text. The locations of the detected headings are then used to estimate each section’s physical page range.
Each section object records the document type, logical file identifier, section title, hierarchy depth, physical page range, body text, and tables assigned to the range.
If both paths fail, the complete document is represented as a single section. Even in this case, L3 receives a bounded region containing the review target and adjacent context rather than the complete document. The failure to recover section structure is retained in preprocessing metadata so that the reviewer can inspect the resulting context limitation.
With explicit deployment approval, a scanned PDF without a usable text layer can be sent to an OCR provider. OCR output is treated as untrusted external input, and text extracted by OCR cannot authorize a rule disposition by itself.

3.2.2. L1: Rule-Based Static Analysis

L1 applies conditions from the rule repository to the preprocessed artifacts and generates rule-matched review targets. Figure 2 shows this flow.
Figure 2. L1 flow of the proposed prototype. L1 output is interpreted as a location-level review target rather than a confirmed violation.
For each source file, L1 receives the original text, comment-stripped text, AST summary, and project-level symbol graph. The rule engine evaluates YAML conditions at file or project scope and records each match with its source location and parser provenance.
Each review target includes a file or document identifier, line or page range, rule identifier, detection type, generation reason, and parser path. Because this location information is preserved through the final review record, the reviewer can navigate from a result to the corresponding source or document context. The number of review targets associated with a file supports navigation and does not represent that file’s severity or violation status.
Table 3 presents the rule-asset domains and their execution paths. The current repository contains 97 code rules executed in L1, 65 document rules processed by the document-validation service, and four traceability rules processed using a separate schema and service. The four detection types described below apply to the code and document rule sets, but not to the four traceability rules.
Table 3. System rule-asset domains and execution paths.
Detection Pattern Types
Table 4 summarizes the four detection types and their downstream routes. Each type relies on different observations: absence of a required signal, textual matching, surrounding context, or syntactic structure. The downstream route is selected according to the decision requirements of each rule; this classification does not constitute an empirical reliability ranking among pattern types.
Table 4. Rule-matching pattern types and rule-specific downstream routes.
  • missing. After comments are removed, a review target is generated if a pattern required by the rule cannot be found. The scope field in each YAML rule specifies whether the check applies at file or project level. A file-scoped rule checks each file separately, whereas a project-scoped rule checks whether the signal occurs anywhere in the submitted source. The absence of a lexical signal alone does not establish that the required security behavior is absent.
  • regex. A compiled regular expression is applied to the original text. Each match is recorded as a separate occurrence with its line range. Rule-specific filters can exclude matches found in comments, identifier names, or irrelevant surrounding contexts. A textual match alone does not establish a violation; downstream processing follows the program-fact and evidence contract defined for that rule.
  • semantic. The configured expression is treated as an initial signal for contextual inspection rather than as a complete decision. For example, the presence of a term associated with a security function does not indicate whether the corresponding value is used in an actual cryptographic operation or only in test code. Semantic rules therefore combine rule-specific analysis logic with file- or project-level context. The resulting targets remain subject to the same evidence and decision-authorization contracts.
  • ast. AST-based rules traverse the syntax tree to inspect loop ranges, array indices, constant references, call order, argument propagation, and definition–use relationships. In the current implementation, 24 LEA rule identifiers are connected to 23 distinct checker functions because one checker is shared by two rules. A checker first extracts the structural facts relevant to a rule and then evaluates whether those facts satisfy its deterministic conditions. If the facts permit a deterministic disposition, the target is processed without L3. Incomplete facts produce hold; only residual cases requiring contextual interpretation can proceed to L3 after satisfying the evidence-authorization conditions.
Domain-Anchor-Based Review-Target Refinement
A domain anchor is a structural or contextual cue used to determine whether a file or code fragment falls within a rule’s intended scope. Because L1 generates review targets broadly to reduce the chance of omission, public constants, test vectors, simple wrappers, benchmark code, and mode-specific files may appear in the initial result. Domain anchors distinguish clear scope mismatches before downstream analysis.
Inclusion anchors provide evidence that a location belongs to the rule’s intended scope. For example, the co-occurrence of an LEA round loop, round-key access, rotation operations, and delta constants supports classifying a file as a possible LEA core implementation. Exclusion anchors identify structures that may not be direct targets of the rule, including public algorithm constants, known-answer test values, lookup tables, wrappers that only delegate processing, and benchmark-only plaintext values.
The refinement step uses both anchor types and records why a review target was retained or excluded. Anchors provide scope evidence rather than ground-truth violation labels. Their independent effect on false positives and recall has not been measured separately; they are therefore interpreted as part of the complete scope-refinement procedure rather than as an independently validated performance improvement. After refinement, verified program facts may authorize a deterministic disposition, incomplete facts produce hold, and only evidence-authorized residual cases become eligible for L3.
YAML Rule-Set Structure
Listing 2 shows an example YAML rule. The fields id, name, category, and scope identify the rule and its application range. The pattern_type field selects the detection path, while pattern stores the expression used by that path. The severity field supports review prioritization, and kcmvp_ref links the rule to related evidence.
Listing 2. YAML-structure summary of the COM-001 rule in the rule set.
Electronics 15 04428 i002
The implementation derives rule-specific program facts before semantic review. COM-001 checks whether a sensitive object is declared, initialized, used in an operation, and subsequently passed to a recognized erasure function. It also checks whether the called erasure function is an empty implementation. COM-004 traces data from a weak random source to a sensitive cryptographic consumer and records whether a strong random value overwrites it before use. When necessary, it also checks whether the relevant code belongs to the declared shipping target. CBC and CTR checkers inspect chaining-value updates, counter progression, and state reuse. Selected LEA checkers inspect round-key use, 128-bit block processing, and definition–use relationships involving required constants.
This stage evaluates relationships directly required for a rule disposition, such as declaration, initialization, call order, data flow, and state transition, rather than merely checking the presence of surrounding strings. It is therefore needed to distinguish superficially similar code from code satisfying the actual rule conditions and to identify cases that can be handled without an LLM call.
Table 5 summarizes the rule-specific program facts and their disposition conditions.
Table 5. Rule-specific program-fact analysis.
If the program facts required by a rule cannot be established sufficiently, the result is not finalized as a violation or non-violation and remains subject to reviewer assessment.

3.2.3. L2: Rule-Bound Evidence Retrieval

L2 proposes the context and evidence used in downstream review. Each L1 review target contains a rule identifier, detection type, source location, and generation reason. L2 extracts a bounded source-code or document region around that location.
Providing the complete source archive and all submitted documents can exceed the model’s input limit or allow irrelevant material to dilute the context directly related to a decision. The input range is therefore bounded. However, an excessively narrow range can omit relevant declarations, callers, conditions, or adjacent document provisions. The system consequently checks whether the selected region contains the declarations, use relationships, call paths, and neighboring provisions required by the rule. If this context-sufficiency condition is not met, L3 is not authorized and the target remains in hold.
Figure 3 summarizes the L2 context and evidence-retrieval flow.
Figure 3. L2 evidence-retrieval flow. Under the decision-authorization gate shown in Figure 1, only evidence-authorized targets enter L3; otherwise, they remain in hold.
L2 separates evidence lookup and ranking from decision authorization. A returned evidence unit is only a proposal until its source identity, digest, locator, rule linkage, applicability, and citation relationship have been validated. A textual or semantic similarity match cannot independently authorize a final disposition.
Retrieval recall measures whether relevant material appears in the retrieved set. Evidence-bundle acceptance measures whether the material satisfies the conditions required for decision use. Citation validity evaluates whether a final response correctly references authorized evidence, whereas disposition coverage and end-to-end accuracy concern the proportion and correctness of binary decisions. These quantities describe different stages and must therefore be evaluated separately; retrieval performance alone cannot establish final disposition performance.
Decision-Use Context Excerption
L3 receives a bounded contiguous region associated directly with the review target rather than the complete source or document. This limits input size while retaining local decision context. The selected range depends on the detection type and document structure. A direct textual match may require only the matching location and adjacent statements, whereas semantic or structural cases may require declarations, conditions, related calls, and data-flow context. A document case may similarly require neighboring provisions from the same section.
Within a fixed token budget, the system preferentially retains declarations, conditions, calls, and neighboring provisions required for the rule. The selected start and end positions and any truncation are recorded so that the reviewer can inspect possible context loss.
  • Code excerption. A contiguous source region is constructed around the line that generated the review target. A relatively narrow window is used for a regex target because its matched location is explicit. A broader window containing related declarations and calls is used for semantic and ast targets because their assessment depends on conditions, call order, and definition–use relationships.
  • Document excerption. The target provision and a bounded number of neighboring provisions from the same document type are combined. Per-section and total-length limits prevent one long section from consuming the complete input budget. If the selected region is truncated, the truncation metadata is included in the decision context.
Verified Evidence Lookup and Optional Retrieval Support
The current decision path first attempts to obtain an exact, verified evidence bundle associated with the rule from the sealed evidence registry. A bundle is accepted only when its registered source identity, digest, locator, rule mapping, authority class, and applicability conditions pass validation. If no verified bundle can be established, the default path fails closed and does not authorize an L3 disposition.
Direct rule mapping associates a rule identifier with a managed evidence unit derived from the cited KCMVP materials [1,23,24]. This explicit mapping makes the rule–evidence relationship inspectable. However, a managed evidence unit is not treated as an authoritative replacement for the original source or as an independently verified decision; its original location and applicability conditions must still be confirmed.
Vector and lexical retrieval are retained as optional exploratory or legacy support paths rather than as the default decision-authorizing sequence. Guidance passages may be divided into sections or paragraphs and stored in ChromaDB for similarity-based retrieval [25]. A lexical TF–IDF path can be used when vector retrieval is unavailable. Results from either path remain evidence proposals until they pass the same registry and applicability checks. Lexical or vector similarity does not establish the authority or semantic sufficiency of the evidence.
The retrieval support normalizes Korean and English expressions, numeric notation, hexadecimal values, rotation notation, affine expressions, and Greek letters so that equivalent expressions can be compared despite surface-form differences. Authorized evidence units are inserted into a prompt field separated from other inputs and are referenced by identifier in the structured result. The model cannot introduce an unregistered rule or evidence identifier.
Although Reciprocal Rank Fusion is a general method for combining heterogeneous ranked lists [26], the current decision-authorizing path does not apply RRF with a fixed value of k = 60 . Consequently, this study does not attribute the reported results to an RRF-based production ranking stage.
Prompt Construction
After context excerption and evidence lookup, a three-part template combines the decision frame, rule-specific review context, and output-safety checks. The separate contribution of each component has not been measured. The template is structured to restrict the model’s role, present rule-relevant facts and authorized evidence explicitly, and prevent the response from exceeding the permitted output format or evidence scope.
Table 6 summarizes the prompt configuration and output-validation roles.
Table 6. Configuration of the hierarchical hybrid prompting architecture.
The first part defines the decision frame. Fixed system instructions do not grant the model authority to create or reinterpret normative requirements; they assign only the bounded task of reviewing a target within the supplied rule and evidence. RAG fields are inserted as untrusted supporting context. The global code-flow summary compresses function definitions and call relationships from the perspective of project-level key lifecycles.
The second part is configured by rule. The current prompt assets include rule-specific boundary examples and structured checklists for selected rules. For rules depending on AST or symbol-graph information, the prompt explicitly supplies relevant types, initial values, parameters, call paths, and program facts. Because these prompt assets are maintained separately and can change with the rule repository, this section does not characterize the implementation using the earlier fixed counts of 18 few-shot rules and 23 checklist rules.
The model returns an integer from 0 to 100 in the existing confidence field. This study interprets that value as a self-reported routing score. It represents how clearly the model reports that it can assess the current input and evidence; it is not a statistically calibrated probability that the assessment is correct [16].
The third part performs output validation and conservative routing. The implementation uses detection-type- and rule-specific operating thresholds, but a routing score alone cannot authorize removal of a review target. Suppression additionally requires a consistent disposition, satisfied evidence and program-fact contracts, and deterministic post-processing validation. Otherwise, the target is retained or converted to hold.

3.2.4. L3: Constrained LLM-Assisted Review

L3 re-evaluates only residual review targets that have passed evidence-authorization conditions but cannot be disposed of using deterministic program facts alone. It is not an authority that creates new KCMVP requirements or replaces the authority of official requirements. Its role is restricted to structuring a review of residual cases within the rules, verified program facts, and authorized evidence supplied by the system.
L3 returns a schema-conforming, citation-bound assessment that can abstain when necessary. It cannot convert a case with incomplete evidence or unresolved program facts into an automatic violation or non-violation disposition.
Figure 4 summarizes restricted L3 review and post-decision validation.
Figure 4. L3 flow of the proposed prototype. Only evidence-authorized residual targets enter L3, and model confidence is interpreted solely as a self-reported routing score. The consistency reassessment is conditional, as described in Section 3.2.4.
LLM-Assisted Review
The L3 prompt combines the rule-specific decision criteria, rule-specific review context, verified program facts, and authorized evidence bundle. Its detailed structure and check order vary according to the detection type. Dynamic fields obtained from source code, documents, and retrieved evidence are placed in explicitly delimited untrusted-data regions so that instructions embedded in submitted material are not interpreted as new system commands.
The L3 implementation uses Gemini 2.5 Flash-Lite through the Gemini API [27]. The generation configuration sets the temperature to 0, uses 42 as the default seed, requests application/json responses, and applies a 60-s request timeout. Evaluation runners may supply another recorded seed while preserving the same seed within each paired comparison. The implementation does not override top-p, top-k, or the maximum output-token parameter, so the provider defaults apply to those settings. These controls standardize the reported execution conditions but do not imply that an external model service is mathematically deterministic.
For an authorized residual target, the model returns a structured assessment, rationale, and referenced evidence identifiers. An output is accepted as a review record only if it satisfies the required schema, identifier allowlist, and citation contract. If any check fails, the output is discarded and the target remains in hold for human review. This result is advisory support for KCMVP preliminary review rather than an official certification decision.
The implementation uses pattern- and rule-specific routing thresholds. The lower operating boundary is 25 for selected AST and semantic paths and 40 for regular-expression and other paths, but these values do not independently authorize suppression. In the isolated review path, an initial response that classifies a target as an issue and reports a score from 65 to 74 may be evaluated once more to inspect response consistency. This reassessment is conditional rather than a universal routing rule for every result in that interval.
The thresholds are operational routing boundaries established during prototype development rather than values optimized against independently adjudicated labels. Regardless of score, removing a review target requires consistent decision content, satisfied evidence and program-fact contracts, and successful validation of the output schema, citations, identifier allowlist, and deterministic post-processing conditions. Rules configured for mandatory retention are never removed solely by model output.
Conflicting evidence, missing context, an unverified rule–evidence mapping, or command-like input causes the system to abstain and assign hold rather than force a disposition. This selective abstention reduces the risk that a target will be incorrectly removed because of incomplete input or inconsistent model output [15].
Interpreting the routing score as a probability of correctness would require independently adjudicated labels and an evaluation of reliability intervals, expected calibration error, Brier score, threshold-specific miss rates, and the risk–coverage relationship. The available proxy labels do not establish this calibration. The score is therefore used only for routing and consistency checks [16].
Final Output and Review Record
Listing 3 shows the normalized review-record contract used to connect an assessment to the rule, evidence, applicability information, and verified program facts. This listing represents the normalized record consumed by the review pipeline rather than claiming that every field is returned directly in a single raw model response.
Listing 3. Normalized structured contract for the final review record.
Electronics 15 04428 i003
The prototype links each review target to its original code or document location and presents the applicable rule, verified program facts, supporting evidence, and disposition status in one review record. An L1 result is displayed as a location requiring additional review rather than as a confirmed violation. If the information or evidence required for a disposition is insufficient, the target remains in hold.
A code review record contains the source file and line range, the condition required by the rule, the program facts identified by the analysis, and the reason the review target was generated. A document review record contains the document type, page or section location, related requirement, and correspondence with source code. The recorded locations allow a reviewer to inspect the relevant original material and its surrounding context.
Each item presents the potential issue together with the rule and evidence used in the assessment. When a remediation direction is available, it is presented as a non-executable alternative for reviewer confirmation; the system does not automatically modify source code or submitted documents. If the program facts or evidence are incomplete, the system does not generate a definitive remediation instruction and instead presents the missing information and additional checks required.
The final report separately summarizes deterministic violation and non-violation dispositions, unresolved review targets, hold states, and the total number of review targets. Detailed results are grouped into common security, cryptographic algorithm, mode of operation, document, and traceability categories. Each item can include the review location, applicable rule, verified program facts, supporting evidence, disposition status, rationale, and a reviewable remediation direction. These outputs are preliminary-review records that support reviewer judgment rather than official KCMVP decisions.
Remediation Suggestions and Report Generation
The report presents the status of each review target, its evidence source, analysis limitations, and non-executable remediation suggestions. Automatically generated explanations and remediation directions are advisory information requiring human confirmation. The system does not predict certification outcomes or automatically apply patches to source code or documents.
For analysis paths to which the integrity contract is applied, a decision record includes the applicable rule, verified program facts, evidence identifiers, and generation metadata. Integrity protection detects whether a protected record was changed after generation; it does not guarantee that the disposition itself is correct. A reviewer can therefore trace not only the final summary but also the code information and evidence used in each protected record.

3.2.5. Traceability Verification

Traceability review compares source code with design, configuration-management, and test evidence. The prototype checks whether documented APIs appear in code, whether documented error codes correspond to declarations, and whether test evidence references configured public functions. These comparisons generate review targets for suspected cross-artifact inconsistencies and do not establish formal conformance.
Cross-Comparison Mechanism
For source code, regular expressions extract public-declaration candidates from header files. Functions marked static and identifiers following configured internal naming conventions are excluded from the public API set. The system links header declarations to source definitions and separately collects configured error-code macros. If a wrapper delegates an API operation to another file, symbol-graph definitions and call relationships supplement textual matching.
For documents, the system searches section text and tables for function names, error codes, test items, and references to APIs under test. These observations support the following three cross-artifact comparisons.
  • Design-document–header comparison. A header declaration not found in the design evidence is recorded as a target requiring confirmation of its documentation status. Conversely, an API described in the design document but not found in the header is recorded as a target requiring confirmation of its implementation status.
  • Design-document–source comparison. Error codes declared in source code and error codes described in documentation are compared in both directions. A value observed on only one side is recorded as a target requiring an additional consistency check between code and documentation.
  • Test-document–header comparison. Configured public functions corresponding to encryption, decryption, key setup, and initialization are compared with test items. A public function for which a reference cannot be found in the test evidence is recorded as a target requiring confirmation of whether it belongs to the tested scope.
The advisory report records the API names, error codes, and test items associated with each cross-artifact review target.
Traceability results remain review targets unless a rule-specific contract authorizes a deterministic disposition. The current extraction can fail to recover declarations generated by macros, declarations spanning multiple lines, function pointers, or API references embedded in tables. Therefore, the final determination of a cross-artifact inconsistency must be made by a human reviewer.

4. Evaluation

The proposed framework is evaluated in terms of rule-based candidate-generation regression agreement, program-fact decision coverage, and the control behavior of L2 evidence retrieval and L3 review. Evidence and untrusted-input stability tests and a case study involving a KCMVP-validated commercial cryptographic module are also presented. The evaluation examines whether the current implementation executes the intended review procedure and how each layer contributes to the decision process. Because the results rely on author annotations, source-derived mutations, and prespecified test vectors, they should be interpreted separately from independently verified KCMVP detection performance.

4.1. Evaluation Setup and Datasets

The evaluation uses five datasets that differ in purpose, construction, and labeling basis. Table 7 summarizes their composition, decision or labeling basis, and evaluation purpose. Because the experiments use different candidate cohorts and evaluation units, their counts are neither aggregated nor treated as directly comparable performance measurements.
Table 7. Summary of evaluation datasets, decision bases, and intended scope.
An L1 file–rule pair denotes an annotated expectation that a specific rule should be triggered at least once in a specific file. Multiple locations produced by the same rule in the same file constitute one file–rule pair but multiple candidate occurrences.
A program-fact or L3 output is accepted as a violation or non-violation only when the decision conditions defined for the corresponding layer are satisfied. A candidate is assigned hold when required code information or evidence is missing or conflicting, or when output-format or citation validation fails. Held candidates are excluded from binary-disposition counts but retained in the denominator for decision coverage and forwarded for human review.
The experiments involving an external LLM or OCR used only authorized research inputs. Source code and documents in actual submissions may contain sensitive information. Deployment on such material therefore requires verification of processing authority and data-retention policies, with local or isolated-network models and OCR preferred for sensitive environments.

4.2. L1 Regression Agreement

The L1 development corpus consists of seven code-archive and design-document sets derived from a C implementation of LEA and public validation material [8,28]. It covers code and document conditions related to CBC and CTR modes, key management, zeroization, random-number generation, the LEA key schedule, the round function, and the Monte Carlo test loop. The authors annotated the rule expected to apply to each file, producing 128 deduplicated positive file–rule pairs.
Table 8 compares the current rule-engine output with the 128 positive file–rule pairs. L1 matched 115 pairs and missed 13, resulting in a regression agreement of 89.84% against the positive author annotations.
Table 8. Current L1 agreement with the positive author annotations.
L1 also produced 41 file–rule pairs that were not included in the positive annotations. Multiple source locations were reported for some file–rule pairs, resulting in 233 location-level candidate occurrences. Because a rule may identify multiple locations within the same file, the 233 occurrences cannot be interpreted as mutually independent errors. The 41 additional pairs are also not classified as false positives because they have not been independently adjudicated.
Five of the 13 missed pairs occurred in Set 3, while Sets 6 and 7 each contained two. Forty of the 41 additional pairs were concentrated in Sets 2, 3, and 4. These results indicate that the detection variation was concentrated in particular code and document configurations rather than distributed uniformly across the seven sets. Determining the causes of individual misses and additional candidates requires further error analysis linking rule identifiers, parser paths, and candidate context.
The results show that L1 can broadly collect potential review locations but cannot independently resolve context-dependent cases. The subsequent layers complement this broad candidate scope by examining program facts, related documents, evidence applicability, and call relationships. Independent review by KCMVP experts is still required to determine which additional candidates are false positives and to establish priorities for rule refinement.

4.3. Program-Fact Decision Coverage

Program-fact evaluation was conducted to determine the extent to which binary dispositions could be generated from structural information extracted from source code. In a separate set of 256 L1 candidates, 65 corresponded to rule-specific evaluation paths that inspect API use, call relationships, definition–use relationships, initialization, or zeroization.
Without invoking an external LLM, program-fact evaluation produced 36 violation and 19 non-violation dispositions. Ten candidates remained hold. The resulting binary-decision coverage was therefore 55/65, or 84.62%.
Table 9 summarizes the program-fact dispositions and decision coverage.
Table 9. Program-fact decision coverage in the separate 256-candidate set.
The ten held candidates lacked sufficient code information or could not be fully evaluated using the extracted facts. They were retained for subsequent review rather than being arbitrarily classified as violations or non-violations. This result demonstrates that the evidence-gating policy permits automated disposition only when the required program facts are available and abstains when the available information is incomplete.
Each disposition is stored as an integrity-protected decision record containing the applied rule, verified program facts, evidence identifiers, and generation metadata. The HMAC detects post-generation modification and binds the record to its producer; it does not establish the substantive correctness or legal authority of the disposition. The record nevertheless enables a reviewer to trace the code information and evidence used for each decision.

4.4. L2 Evidence Retrieval and Fixed-Cohort RAG Experiment

To examine how additional official evidence affects L3 decisions and abstention behavior, nine source-derived CBC and CTR mutation groups were evaluated under two paired conditions. Each group was processed once without additional retrieved evidence and once with an improved official-evidence bundle. The model, paired generation-seed schedule, prompt contract, candidate order, and software configuration were held constant across both conditions.
Both conditions produced six binary dispositions and three holds. No paired disposition changed between conditions, and the exact McNemar test yielded p = 1.0 [29].
Table 10 reports the paired evidence-condition results.
Table 10. Fixed-cohort RAG experiment with nine mutation groups per condition.
The improved evidence bundle did not change the number of binary dispositions, and the three cases with insufficient information remained held in both conditions. The experiment therefore does not demonstrate that the provision of official evidence improves decision accuracy. It does, however, show that incomplete cases were not forced into binary dispositions in either condition, which is consistent with the conservative routing behavior of the evidence gate.
L2 is not an independent decision maker that determines compliance from retrieval results alone. It supplies source-identified evidence to L3 and permits its use only when the evidence identifier and citation relationship are valid. This design prevents a candidate from being automatically suppressed solely on the basis of a retrieval score or the LLM’s self-reported confidence. Although the experiment does not directly measure hallucination frequency, restricting L3 to evidence bundles linked to actual cryptographic-module rules reduces the possibility that untraceable content will be accepted as a decision basis. The present result therefore characterizes L2 as a control over the evidence scope and decision authority of LLM-assisted review rather than as a demonstrated source of accuracy improvement. Retrieval quality, faithful evidence use, and generation quality remain distinct RAG evaluation dimensions [30].

4.5. Evidence and Input Stability Evaluation

Evidence and input stability were evaluated to determine whether retrieved evidence remained correctly linked to its rules and whether untrusted content in externally submitted material could alter the decision procedure.
The retrieval test used three human-reviewed rule–evidence relationships as representative queries. Each query was repeated 20 times with the top three retrieved evidence units. Recall@3, which denotes the proportion of all relevant evidence included among the top three results, was 0.778. Relevant or prespecified oracle bundles passed citation validation in all trials, whereas irrelevant or conflicting evidence did not authorize a disposition.
The result confirms that the control procedure for checking evidence relevance and citation relationships operated as intended for the evaluated queries. However, three representative queries are insufficient to characterize retrieval performance across all 166 rule assets. Future evaluation should expand the query set across rule categories and separately measure retrieval recall, faithful evidence use, citation validity, and final decision performance [30].
Code, comments, OCR output, and retrieved passages can blur the boundary between data and instructions, creating an indirect prompt-injection risk [17,31]. The implementation separates dynamic fields extracted from source code and documents from fixed system instructions by representing them as untrusted JSON data. Rule and evidence identifiers are restricted to allowlists, and output format, evidence identifiers, and citation relationships are revalidated after each LLM call. Detected instruction-like content produces hold before the LLM is invoked. ZIP admission additionally enforces restrictions on paths, file types, normalized-name collisions, nesting depth, compression ratio, and total size.
Table 11 summarizes the prespecified attack tests and clean controls.
Table 11. Prespecified untrusted-input boundary tests.
The prompt-injection vectors placed attack text in source comments and strings, document text and tables, OCR-like text, retrieved evidence, archive names, Korean-language instructions, and hidden Unicode control characters. All nine vectors were blocked before the LLM was invoked, and requests intended to suppress candidates were converted to hold. All nine clean controls followed the normal analysis pipeline. The malicious-archive test rejected all nine vectors involving path manipulation, abnormal file types, excessive nesting, or abnormal compression ratios, while both clean ZIP archives proceeded to analysis.
These results show that the implemented boundary controls blocked the prespecified attack vectors while preserving the normal input flow. The tests do not establish security against untested adaptive attacks, other archive formats, all parser vulnerabilities, or arbitrary image-to-OCR attacks. Additional validation with broader attack types and file formats is therefore required before deployment.

4.6. Validated Commercial-Module Case Study

To examine behavior on actual code, the framework was applied to the analyzable source subset of a commercial cryptographic module that had obtained KCMVP validation. The dataset consisted of 59 C and header files totaling 14.5 KLOC, from which L1 generated 11 review candidates.
The candidates primarily concerned locations requiring further examination of residual-data clearing and code–document traceability. The analyzed material did not include the complete compiler configuration of the validation-target build, all conditional-compilation results, generated code, or binary-level behavior. Consequently, the analysis could not automatically determine whether a source-level store was preserved in the final binary or whether an equivalent security operation was performed through a wrapper function or macro. Traceability between the implementation and its supporting documents was also incomplete when the complete submission-document set was unavailable.
Regular-expression and name-based rules cannot always recover the semantics of implementations that use non-standard function names, macros, function pointers, or cross-file wrappers. The 11 candidates were therefore retained for additional review rather than being classified as false positives. The result indicates that L1 alone is insufficient for determining final conformity in an actual validated module and that subsequent review must consider build context and submission documents.
The commercial-module results suggest that the proposed system is best used as a pre-review tool for identifying code and document locations requiring further examination rather than as a mechanism for reassessing an existing validation decision. Future analysis should classify the generated candidates into concrete risk categories, such as potential residual sensitive data, possible removal of zeroization operations during optimization, API call-order violations, and code–document traceability gaps. This classification should distinguish actual security weaknesses or KCMVP requirement violations from candidates caused by static-analysis limitations or implementation-specific exceptions. Feeding these findings back into rule definitions for naming conventions, wrapper structures, and exception conditions can reduce unnecessary candidates and improve the practical value of pre-certification review.
The evaluation has four principal limitations. First, the development corpus and positive annotations are centered on LEA and author-constructed cases, limiting generalization to other KCMVP-approved algorithms such as AES, ARIA, and SEED. Second, the absence of independent KCMVP-expert judgments and a labeled clean-negative corpus prevents estimation of whole-system precision, specificity, and false-positive rate. Third, the commercial-module case did not include the complete validation-target build context required to resolve all effects of conditional compilation, generated code, and binary optimization. Finally, the paired RAG and retrieval experiments are too small to establish component-level performance differences. Future work should expand algorithm-specific and clean corpora, obtain independently adjudicated expert labels, and conduct component ablation and source–binary linkage analyses using identical candidate cohorts and build conditions.

5. Conclusions

This study designed and implemented an evidence-gated decision-support prototype for KCMVP pre-certification review. The pipeline links rule-based review-target generation, rule-bound evidence retrieval, program-fact verification, constrained LLM-assisted review, and an integrity-protected decision record. It applies these stages across code, submission documents, and traceability artifacts. Together, these stages avoid treating a rule-matched location as an immediate violation. Instead, the pipeline presents the relevant source evidence together with the operational relationships identified in the code, allowing reviewers to trace how a result was produced and which basis supported it. When the evidence or program facts required for a disposition are insufficient, the review target remains in hold rather than being removed. This reduces the possibility that incomplete information will cause a potentially relevant location to be omitted automatically.
Evaluation results show that L1 matched 115 of 128 author-annotated positive file–rule pairs, corresponding to 89.84% agreement with the positive author annotations, while program-fact analysis produced dispositions for 55 of 65 eligible review targets and held the remaining ten. These results indicate that rule-based analysis can capture a substantial portion of the predefined issue locations as initial review targets. They also show that program facts such as function calls and data flow can resolve some targets without an LLM call while routing uncertain cases to human review. The RAG comparison produced no difference in dispositions between the two conditions ( p = 1.0 ). Its observed role was therefore not an improvement in decision accuracy, but the provision of traceability and control by linking original evidence to rules and authorizing a disposition only when that evidence was verified. The security regressions also blocked all nine authored prompt-injection vectors and all nine malicious ZIP vectors while accepting the corresponding clean controls. This result indicates that untrusted-input separation and input validation operated as intended within the tested attack boundaries.
However, these evaluations are based mainly on LEA-centered cases, positive annotations created by the authors during development, a limited set of mutations, and prespecified attack vectors. No external expert labels or independently labeled clean-negative corpus were available, and cases not selected by L1 were not reviewed separately. The present results therefore do not establish precision, specificity, false-positive rate, routing-score calibration, or cross-algorithm generalization for the full KCMVP rule set. The security regressions demonstrate resistance only to the constructed cases and do not constitute a general security proof against unknown attacks. Accordingly, the prototype should be interpreted as an auxiliary pre-certification review tool that structures locations and related evidence for reviewer inspection, not as an automated decision system that replaces official KCMVP validation.
Future work will obtain labels from independent KCMVP experts through blinded assessment and measure inter-rater agreement to evaluate disposition accuracy and routing-score behavior. The rule set and analysis paths will also be extended beyond the C and C++ implementations emphasized in this study to other implementation languages used for cryptographic modules subject to validation in Korea. The coverage of cryptographic algorithms and submission documents will also be expanded. Where access is permitted, the evaluation corpus will be extended with source code and submission documents from validated cryptographic modules. The proposed system will then be evaluated in actual KCMVP pre-certification review and validation workflows in terms of detection accuracy, the appropriateness of held dispositions, review time, and support for reviewer judgment.

Author Contributions

Conceptualization, S.-B.C. and H.-J.S.; methodology, S.-B.C. and H.-J.S.; software, investigation, S.-B.C., D.-Y.P., D.-E.L., J.-H.K., S.-M.J. and Y.-L.H.; writing—original draft preparation, S.-B.C. and Y.-L.H.; writing—review and editing, S.-B.C. and Y.-L.H.; supervision, H.-J.S.; project administration, H.-J.S. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (No. RS-2025-02306395, Development and Demonstration of PQC-Based Joint Certificate PKI Technology, 50%) and this work was supported by the Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (No. RS-2025-25394739, Development of Security Enhancement Technology for Industrial Control Systems Based on S/HBOM Supply Chain Protection, 50%).

Data Availability Statement

The implementation and experimental artifacts are not publicly available. Commercial source code and confidential submission materials are not disclosed.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
APIApplication Programming Interface
ASTAbstract Syntax Tree
FNFalse Negative
FPFalse Positive
HMACHash-based Message Authentication Code
KCMVPKorean Cryptographic Module Validation Program
KISAKorea Internet & Security Agency
LEALightweight Encryption Algorithm
LLMLarge Language Model
LOOLeave-One-Out
RAGRetrieval-Augmented Generation
RRFReciprocal Rank Fusion
TNTrue Negative
TPTrue Positive

References

  1. National Intelligence Service (NIS). Guidelines for Cryptographic Module Testing and Validation. Technical Report, National Intelligence Service, 2025. Revision Dated 1 June 2025. Available online: https://www.ncsc.go.kr/ko/main/PageLink.html?token=MDEyMzQ1Njc4OWFiY2RlZrBvBnCd4cC_Aqu-HrYk1bggl92j52iiEKC6MeRpbdR99hekw1dn45EjcBXqmojIcgWkJfOAuXl5Bj6LZHG_H56x-vMhjY_8ixjMx_eaJQfK (accessed on 17 September 2026). (In Korean)
  2. Korea Internet & Security Agency (KISA). Cryptographic Module Validation Program: Overview. Official Program Webpage. 2026. Available online: https://seed.kisa.or.kr/kisa/kcmvp/EgovSummary.do (accessed on 14 September 2026). (In Korean)
  3. Korea Internet & Security Agency (KISA). Cryptographic Module Testing and Validation Procedure. Official Procedure Webpage. 2026. Available online: https://seed.kisa.or.kr/kisa/kcmvp/EgovProcedure.do (accessed on 14 September 2026). (In Korean)
  4. Presidential Decree of the Republic of Korea. Cyber Security Work Regulations, 2024. Presidential Decree No. 34287 (Partial Amendment 5 March 2024, Effective 1 January 2025). Available online: https://www.law.go.kr/LSW/lsInfoP.do?lsiSeq=261003&efYd=20250101&joNo=000900 (accessed on 17 September 2026). (In Korean)
  5. Presidential Decree of the Republic of Korea. Article 69 of the Enforcement Decree of the Electronic Government Act, 2025. Presidential Decree No. 35948 (Amendment by Other Law 30 December 2025, Effective 2 January 2026). Available online: https://www.law.go.kr/LSW/lsInfoP.do?lsiSeq=281535&efYd=20260102&joNo=006900 (accessed on 17 September 2026). (In Korean)
  6. KS X ISO/IEC 19790:2015; Information Technology—Security Techniques—Security Requirements for Cryptographic Modules, 2015. Korean Adoption of ISO/IEC 19790:2012. Official Amendment Notice No. 2015-0342. Korean Agency for Technology and Standards: Eumseong-gun, Republic of Korea, 2015. Available online: https://standard.go.kr/KSCI/ksNotification/getKsNotificationView.do?menuId=921&ntfcManageNo=2015-0342&topMenuId=502 (accessed on 17 September 2026).
  7. KS X ISO/IEC 24759:2015; Information Technology—Security Techniques—Test Requirements for Cryptographic Modules, 2015. Korean Adoption of ISO/IEC 24759:2014. Official Amendment Notice No. 2015-0342. Korean Agency for Technology and Standards: Eumseong-gun, Republic of Korea, 2015. Available online: https://standard.go.kr/KSCI/ksNotification/getKsNotificationView.do?menuId=921&ntfcManageNo=2015-0342&topMenuId=502 (accessed on 17 September 2026).
  8. Hong, D.; Lee, J.K.; Kim, D.C.; Kwon, D.; Ryu, K.H.; Lee, D.G. LEA: A 128-Bit Block Cipher for Fast Encryption on Common Processors. In Proceedings of the 14th International Workshop on Information Security Applications (WISA 2013); Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2014; Volume 8267, pp. 3–27. [Google Scholar] [CrossRef] [Scilit]
  9. Korea Internet & Security Agency (KISA). LEA: 128-Bit Block Cipher Specification and Resources. Official Algorithm Information Page with the Korean Specification. 2026. Available online: https://seed.kisa.or.kr/kisa/algorithm/EgovLeaInfo.do (accessed on 14 September 2026).
  10. Rahaman, S.; Xiao, Y.; Afrose, S.; Shaon, F.; Tian, K.; Frantz, M.; Kantarcioglu, M.; Yao, D. CryptoGuard: High Precision Detection of Cryptographic Vulnerabilities in Massive-Sized Java Projects. In Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security (CCS); ACM: New York, NY, USA, 2019; pp. 2455–2472. [Google Scholar] [CrossRef] [Scilit]
  11. Krüger, S.; Nadi, S.; Reif, M.; Ali, K.; Mezini, M.; Bodden, E.; Göpfert, F.; Günther, F.; Weinert, C.; Demmler, D.; et al. CogniCrypt: Supporting Developers in Using Cryptography. In Proceedings of the 32nd IEEE/ACM International Conference on Automated Software Engineering (ASE); IEEE: Piscataway, NJ, USA, 2017; pp. 931–936. [Google Scholar] [CrossRef] [Scilit]
  12. National Institute of Standards and Technology (NIST). Cryptographic Algorithm Validation Program (CAVP). Available online: https://csrc.nist.gov/projects/cryptographic-algorithm-validation-program (accessed on 31 March 2026).
  13. National Institute of Standards and Technology. OSCAL: Open Security Controls Assessment Language. Available online: https://pages.nist.gov/OSCAL/ (accessed on 13 September 2026).
  14. Fang, C.; Miao, N.; Srivastav, S.; Liu, J.; Zhang, R.; Fang, R.; Asmita; Tsang, R.; Nazari, N.; Wang, H.; et al. Large Language Models for Code Analysis: Do LLMs Really Do Their Job? In Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24), Philadelphia, PA, USA, 14–16 August 2024; pp. 829–846. Available online: https://www.usenix.org/conference/usenixsecurity24/presentation/fang (accessed on 17 September 2026).
  15. Geifman, Y.; El-Yaniv, R. Selective Classification for Deep Neural Networks. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30, Available online: https://proceedings.neurips.cc/paper/2017/hash/4a8423d5e91fda00bb7e46540e2b0cf1-Abstract.html (accessed on 17 September 2026).
  16. Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On Calibration of Modern Neural Networks. In Proceedings of the 34th International Conference on Machine Learning; Proceedings of Machine Learning Research: Cambridge, MA, USA, 2017; Volume 70, pp. 1321–1330. Available online: https://proceedings.mlr.press/v70/guo17a.html (accessed on 17 September 2026).
  17. Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; Fritz, M. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. arXiv 2023, arXiv:2302.12173. [Google Scholar] [CrossRef] [Scilit]
  18. Aho, A.V.; Lam, M.S.; Sethi, R.; Ullman, J.D. Compilers: Principles, Techniques, and Tools, 2nd ed.; Pearson Education: Boston, MA, USA, 2006; ISBN 978-0-321-48681-3. [Google Scholar]
  19. LLVM Project. Clang: A C Language Family Frontend for LLVM. 2010. Available online: https://clang.llvm.org/ (accessed on 31 March 2026).
  20. Bendersky, E. pycparser: Complete C99 Parser in Pure Python. Software Documentation. 2008. Available online: https://pypi.org/project/pycparser/ (accessed on 29 March 2026).
  21. Singer-Vine, J. pdfplumber. Software Documentation. 2016. Available online: https://pypi.org/project/pdfplumber/ (accessed on 29 March 2026).
  22. Artifex Software. PyMuPDF: Python Bindings for MuPDF. 2024. Available online: https://pymupdf.readthedocs.io/ (accessed on 29 March 2026).
  23. National Intelligence Service (NIS). Guide for Preparing Cryptographic Module Submissions. Technical reort, National Intelligence Service, 2022. Official Release Notice Dated 31 August 2022. Available online: https://seed.kisa.or.kr/kisa/Board/145/detailView.do (accessed on 17 September 2026). (In Korean)
  24. National Intelligence Service (NIS). Cryptographic Module Implementation Guide, Part 2: Cryptographic Algorithm Implementation Guide. Technical Report, National Intelligence Service, 2024. Official Release Notice Dated 15 March 2024. Available online: https://www.nis.go.kr/AF/1_7_3_5/view.do?currentPage=1&seq=115 (accessed on 14 September 2026). (In Korean)
  25. Chroma. Chroma: The AI-Native Open-Source Embedding Database. Software Documentation. 2023. Available online: https://docs.trychroma.com/docs/overview/introduction (accessed on 29 March 2026).
  26. Cormack, G.V.; Clarke, C.L.A.; Buettcher, S. Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods. In Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR); ACM: New York, NY, USA, 2009; pp. 758–759. [Google Scholar] [CrossRef] [Scilit]
  27. Google. Gemini 2.5 Flash-Lite Model Documentation. Official Gemini API Documentation. 2025. Available online: https://ai.google.dev/gemini-api/docs/models/gemini-2.5-flash-lite (accessed on 14 September 2026).
  28. National Intelligence Service (NIS). List of Validated Cryptographic Modules. Official KCMVP Module List. 2026. Available online: https://www.nis.go.kr/AF/1_7_3_3/list.do (accessed on 14 September 2026). (In Korean)
  29. McNemar, Q. Note on the Sampling Error of the Difference between Correlated Proportions or Percentages. Psychometrika 1947, 12, 153–157. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Es, S.; James, J.; Anke, L.E.; Schockaert, S. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, St. Julians, Malta, 17–22 March 2024; pp. 150–158. [Google Scholar] [CrossRef] [Scilit]
  31. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile; Technical Report NIST AI 600-1; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2024. [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.