1. Introduction
RSA moduli are ordinarily evaluated through the security problem they define: given a public modulus N = pq, can an adversary recover its prime factors? Once p and q are known, a different forensic question becomes possible: what, if anything, do the factors reveal about the procedure by which they were generated? Prime-candidate formatting, congruence filters, factor-size allocation, and pair-level acceptance rules can leave persistent arithmetic structure in the factors. Such structure may remain observable long after the original software, random seeds, and operational records have disappeared [
1].
The solved RSA Challenge moduli provide an unusual historical dataset for studying this question. Unlike ordinary deployed RSA public keys, the factors of the solved Challenge numbers are publicly available. Their provenance also includes two distinguishable populations: the original decimal-labelled Challenge numbers and the later bit-labelled numbers generated through a separately documented BSAFE process [
2,
3]. The dataset therefore permits direct factor-level analysis while requiring explicit provenance control. RSA-129 is excluded from the study population because it originated as an earlier 1977 cryptographic puzzle rather than as part of the formal RSA Factoring Challenge [
4].
A construction family is a class of prime-generation and factor-selection procedures that share observable constraints, such as candidate high-bit settings, congruence conditions, and factor-size rules, without necessarily sharing the same software implementation. A construction-family signature is a reproducible arithmetic property consistent with such a class. It supports family-level inference but does not, by itself, uniquely identify a library, algorithm, device, seed, or implementation. In this terminology, a feature is a directly computed quantity, a signature is a reproducible pattern associated with a construction class, and a fingerprint is reserved for evidence that has been validated as sufficiently discriminating for implementation-level attribution.
The first form of evidence examined here is high-bit conditioning. For a positive integer x, define the normalised binary mantissa by
where b(x) is the bit length of x. Every retained prime factor satisfies M
2(x) ≥ 3/2, which is equivalent to binary prefix 11. Writing
gives
Because both normalised mantissas are at least 3/2,
When this identity is combined with balanced factor-bit allocation, it explains the observed division between equal factor-bit lengths for even-bit moduli and adjacent factor-bit lengths for odd-bit moduli. The bit-parity pattern is therefore treated as a consequence of more fundamental construction conditions rather than as independent evidence.
The second form of evidence is the residue constraint documented for the original decimal-labelled Challenge population:
This rule made the public exponent e = 3 compatible with each factor and provides a historically documented positive control [
2]. It also produces a derived consequence. With A = (p + q)/2, the difference q − p is divisible by both 2 and 3, so
The modulo-9 condition is therefore not a second independent construction-family signature; it is an algebraic consequence of the documented modulo-3 rule.
The third form of evidence is factor-balance. The analysis separates discrete factor-bit balance from continuous magnitude balance. The factor-bit difference records whether the two factors occupy equal or adjacent bit intervals, while the ratio q/p describes their relative magnitudes. Normalised AM-GM quantities are treated as transformations of that ratio, not as independent evidence. This dependency-aware treatment is necessary because several visually distinct numerical patterns can arise from the same underlying size constraints.
The central premise is that the solved RSA Challenge factors preserve layered arithmetic evidence of their construction. High-bit conditioning identifies a broad prime-generation family; the modulo-3 constraint identifies a documented rule within the original decimal-labelled population; and factor-balance patterns describe the effects of factor-size allocation and pair selection. These observations support cautious family-level provenance analysis, but they do not uniquely identify an implementation or imply reduced factoring security.
The paper addresses three research questions:
RQ1. Which high-bit, congruence, and factor-balance properties occur in the prime factors of the 22-modulus RSA Challenge dataset?
RQ2. Which observed patterns are direct construction evidence, which reproduce documented generation rules, and which are algebraic consequences of other constraints?
RQ3. Which properties remain informative after comparison with controls matched for factor-bit lengths, public modulus labels, and relevant residue restrictions, and what level of provenance inference do they justify?
This paper is framed as a methodological case study in cryptographic forensics, not as a search for a previously unknown structure in the RSA Challenge moduli. It demonstrates how dependency-aware arithmetic analysis, provenance partitioning, and matched generative controls can be applied to a small historical corpus with publicly known factors. The high-bit result, recovery of the documented modulo-3 rule, and factor-balance analysis are therefore presented as illustrations of an evidential method rather than as newly discovered RSA Challenge construction rules.
The study makes five contributions:
A provenance-controlled dataset that separates the original decimal-labelled and later bit-labelled Challenge populations rather than treating all solved moduli as outputs of one homogeneous generator.
An exact high-bit result: all 44 retained factors begin with binary prefix 11, together with a derivation showing why this sufficient condition forces b(N) = b(p) + b(q).
A reconstruction of the documented rule that both factors are congruent to 2 modulo-3 in all 18 original decimal-labelled pairs, together with an explicit separation of that rule from its derived residue consequences.
A dependency-aware analysis of factor-balance that distinguishes discrete bit allocation from continuous factor imbalance and avoids counting algebraically equivalent metrics as separate signatures.
A bounded provenance interpretation in which matched controls are used to distinguish generic generation practices from historically associated constraints, without claiming exact implementation attribution or key weakness.
Scope boundary. This study analyses construction evidence, not cryptographic weakness. None of the reported high-bit, residue, or balance properties is claimed to reduce the practical difficulty of factoring the associated moduli. Correspondence with a construction family is not treated as unique attribution to a particular implementation. Exact implementation attribution would require substantially larger labelled datasets, prospectively specified classifiers, and out-of-sample validation.
The novelty lies in the study design and evidential separation, not in the individual high-bit or modulo-3 conditions. The paper provides a provenance-controlled, corpus-wide analysis of the solved formal RSA Challenge moduli, including the 44-of-44 high-bit observation and its comparison with controls matched to factor-bit lengths, public labels, and provenance-specific constraints. These comparisons provide model-relative evidence that a size-and-label-only construction model is inadequate; they do not establish a unique historical mechanism. The study’s contribution is therefore a dependency-aware framework that distinguishes exact arithmetic observations, recovery of documented construction metadata, algebraic consequences, and bounded construction-family inferences.
The remainder of the paper is organised as follows.
Section 2 reviews related work and positions the study.
Section 3 describes the materials and methods.
Section 4 presents the frozen comparison plan and statistical analysis, and
Section 5 reports the results.
Section 6 discusses the findings,
Section 7 outlines the limitations and priorities for future validation, and
Section 8 concludes the paper.
4. Statistical Analysis and Frozen Comparison Plan
This section records the analysis plan governing matched-control generation and inference. The historical Challenge factors and preliminary summaries were examined during development of the earlier draft; consequently, Protocol S1 is not a preregistration, H1 and H2 are model-adequacy checks whose directions were already visible, and Challenge-only observations are not represented as prospectively confirmed discoveries. Before any new matched control was generated or inspected, however, the protocol froze the population strata, primary variables, model hierarchy, statistics, directions, multiplicity treatment, sensitivity analyses, and permissible claim language. Protocol S1 contains these decisions in machine-readable form, versioned as 1.0 and frozen on 13 July 2026.
6. Discussion
The analysis distinguishes exact observations, independently documented rules, model-relative construction-family inferences, and algebraic restatements of the same underlying constraints. It supports two bounded findings and does not support a third: high-bit conditioning is required to explain the corpus-level all-11 pattern after size and public-label matching; the original list modulo-3 rule is recovered as a documented positive control; and the adjacent-bit pairing contrast does not establish a separate factor-balance signature after the documented constraints are modelled. This section interprets those findings at the evidential levels fixed in
Table 3.
6.6. Methodological Implications
Three methodological lessons follow. First, controls should match the public selection problem. Exact factor-bit lengths and Challenge labels materially affect leading-bit probabilities, so a naive independence calculation is inadequate as the principal baseline. Second, the modulus is the statistical unit. Resampling complete pairs preserves the dependence between p and q and avoids overstating precision through 44 nominally independent factor observations. Third, evidence should be dependency-aware. Product-bit parity, AM-GM bands, and Fermat-square residues remain useful explanatory quantities, but they do not multiply the amount of evidence when they are determined by primary features.
The H3 outcome also illustrates the value of a frozen analysis plan even when full preregistration is impossible. The historical factors had already been inspected, so the protocol does not turn the study into a prospective experiment. It does, however, prevent the newly generated controls from being searched across changing statistics, tails, and subgroup definitions until a preferred result appears. Recording H3 as unsupported, despite a lower observed median, is a substantive result of that discipline.
Finally, reproducibility is part of the evidential claim. The supplementary package records the exact tuples, validation code, model definitions, pool-generation code, accepted control pools, manifests, quality control output, inference results, protocol amendments, and hashes. The independently seeded rerun checks pseudo-dataset draw stability from the same pools. It does not replace an independently regenerated control archive, and the manuscript states that limitation directly.
The method has bounded relevance to critical infrastructure assurance. In authorised known-factor corpora or controlled key-generation test environments, dependency-aware factor analysis may support provenance assessment and validation of documented generation constraints. It does not provide a public-key scanner or an operational governance control.
High-bit conditioning is treated here as a generic construction-family feature consistent with documented RSA prime-generation practices, including FIPS 186-5 and OpenSSL [
5,
6]. The matched-model
p-values are conditional on the control families defined in the protocol and should not be interpreted as unconditional historical probabilities. The present analysis does not include a modern OpenSSL- or BSAFE-based full key-generation baseline. Such an implementation-specific sensitivity analysis remains future work and would require a versioned build and API path, public exponent, random-number-generation and seeding procedure, factor-size allocation, rejection behaviour, sample size, and archived outputs.
Overall, the study demonstrates a dependency-aware forensic method on a small historical dataset and retains only conclusions consistent with known or generic RSA construction practices. It neither claims previously unknown RSA Challenge structure nor provides a public-key scanner, governance control, weak-key detector, or implementation-attribution classifier. Such applications would require substantially larger labelled generator corpora and rigorous out-of-sample validation.
7. Limitations and Future Validation