Abstract
This paper studies Jaccard similarity under a fully coupled one-factor scoring policy in which all coordinates of an interval-valued fuzzy soft profile are governed by the same parameter. Each interval runs from confidence-discounted evidence to the value obtained when that evidence is accepted in full. In contrast to the usual independent coordinate envelope, the coordinates move together. We prove that the exact extrema can be found from finitely many coordinate crossings and that every possible change in pairwise ranking is determined by breakpoint evaluations and polynomial equations of degree at most two. The method is illustrated in a single-corpus six-character proof of concept based on a traceable positive evidence analysis of Elif Shafak’s The Saint of Incipient Insanities. The case study yields seven positive pair similarities and two exact ranking transitions. It also shows that some similarity values and rankings change when normalization, weighting, parameter selection, or coding decisions change. Thus, this paper provides an exact way to study similarity and ranking stability under a common scoring policy while keeping the literary conclusions within the limits of the analyzed corpus.
Keywords:
interval-valued fuzzy soft sets; Jaccard similarity; coherent scoring path; ranking stability; literary character analysis; evidence aggregation MSC:
03E72; 90C31; 90C32
1. Introduction
Literary characters emerge through many kinds of textual evidence: direct description, self-report, observable action, reported speech, and the attributions of other characters. Computational studies often represent these signals through mentions, dialogue, network position, sentiment, or predefined personality labels. Such approaches are valuable for large-scale character networks and personality-oriented modeling, but they do not by themselves show, in a traceable way, how much positive support a defined parameter receives across a character’s documented appearances [1,2,3]. This paper asks that narrower question. It does not attempt to separate counter-evidence, unresolved evidence, irrelevant passages, and passages in which the parameter could not reasonably be assessed.
Soft set theory describes objects relative to parameters without requiring a single universal membership function [4]. Fuzzy soft sets add graded membership to this parameterized setting [5], and the resulting framework has supported a wide range of decision models [6,7,8,9,10,11]. For literary analysis, the value of this framework is not that it turns a novel into a conventional ranking problem. Rather, it offers a clear way to represent character–attribute relations when the supporting evidence is graded, comes from different sources, and remains open to interpretation.
Several parts of this problem already have precedents in the fuzzy–soft literature. Multi-observer models combine more than one assessment [8]; bipolar fuzzy soft sets separate positive and negative evaluations [12]; and incomplete soft set models address unavailable entries [9,13]. More expressive soft set extensions have also been developed for intuitionistic, rough, neutrosophic, and hesitant information [14,15,16,17]. These precedents help keep the novelty claim in proportion: simply recording support and opposition would not amount to a new mathematical contribution. The methodological problem addressed here is how to connect established representations to an auditable literary protocol in which every number can be traced to a textual unit and narrative exposure, evidence source, and adjudication confidence remain distinct.
Research on narrative annotation reinforces this need. Richly annotated literary corpora depend on explicit layers, documented guidelines, and stable textual units [18]. Computational studies of emotion in drama likewise show that the unit of analysis must match the textual phenomenon being modeled [19]. For that reason, the present protocol is checked through auditability, computational repeatability, and sensitivity to the recorded decisions.
The present study uses an interval-valued fuzzy soft framework in a novel-length mathematical case study of Elif Shafak’s The Saint of Incipient Insanities [20]. The analytical narrative is delimited to printed pages 3–351 of the cited edition, excluding cover material, promotional text, copyright information, epigraphs, acknowledgments, and author notes. This delimitation is a corpus fact rather than a finding. After rule-assisted contextual adjudication, all character–parameter values, similarity envelopes, and robustness comparisons are computed deterministically.
The representation is intentionally conservative. Every retained record has an evidence strength grade s and a contextual adjudication confidence grade c. Collapsing them into a single product would hide the difference between a weak but clear passage and a strong passage whose interpretation is less certain. We therefore use the established interval-valued fuzzy soft set model [21]. In the interval , the lower endpoint discounts the evidence by the assigned confidence, whereas the upper endpoint fully accepts the same evidence strength. The interval covers these two scoring assumptions and the interpolations between them. It is not a probabilistic confidence interval, a possibility distribution, or a claim that an unknown “true” membership must lie within its endpoints. Aggregating these intervals by narrative exposure produces a character–parameter scoring range. Counter-evidence, missingness, and other coordinates are not introduced unless they have actually been observed and coded.
The study asks four connected questions. First, can evidence strength and contextual adjudication confidence be retained without merging them into one grade? Second, how do normalization and parameter weighting affect the profiles and pairwise similarities? Third, do the similarity patterns remain stable across all scoring realizations within the constructed intervals? Fourth, does one coherent scoring policy preserve the order of the character pairs? The mathematical contribution is neither a new fuzzy set species nor the first interval extension of Jaccard similarity. Established embedded set bounds serve as the independent coordinate benchmark. The additional construction is a coupled path along which every membership moves from its discounted value to its fully accepted value under the same parameter. Exact extrema and ranking transitions on this path reduce to finitely many breakpoint and polynomial root calculations. The empirical contribution is a traceable occurrence-level case study supported by binary, raw-count, log-normalized, alternative-denominator, equal-domain, and deletion comparisons. No claim about this corpus is generalized to other authors, genres, languages, or literary traditions.
The main contribution is the question we pose about Jaccard similarity under interval uncertainty. Existing embedded set methods allow the interval coordinates to vary independently and optimize over the resulting rectangular set. We instead require every coordinate to follow the same policy parameter. The word coherent refers to this common scoring rule; it does not imply that an empirical dependence structure has been estimated. Under this restriction, we identify a finite set of events that contains every exact similarity extremum and every possible change in the ranking of several pairwise profiles. These events include coordinate switches, isolated ties, persistent tie intervals, and the endpoints of the policy range. The differentiation and cross-multiplication used in the derivation are standard. What is new is the event-based Jaccard construction that brings them together into a complete finite ranking analysis, avoiding endpoint-only and grid-based comparisons.
2. Related Work
Soft sets represent uncertainty through parameterized families of subsets [4]; fuzzy soft sets replace crisp membership by graded membership [5], and early decision procedures developed this framework further [6,7]. Subsequent work has expanded the soft set family in many directions, including relation-based, intuitionistic, neutrosophic, hesitant, incomplete, and adjustable constructions [9,10,14,15,16,17,22]. These developments establish a broad representational background, but the present study does not introduce a new soft set species or add uncertainty coordinates that are not observed in the literary protocol.
The representation used here is the established interval-valued fuzzy soft set of Yang et al. [21], in which every object–parameter pair receives an interval in . Its subsequent literature includes level-soft, preorder, incomplete-data, and decision procedures [23,24,25], together with continuing fuzzy–soft applications such as multi-attribute evaluation [11]. For the present purpose, the interval retains two declared scoring realizations of the same positive-evidence record rather than serving as a new uncertainty ontology.
Similarity for interval-valued fuzzy information has been approached through distance, logarithmic, and related measures [26,27], while axiomatic work cautions that distance, dissimilarity, divergence, and similarity should not be treated as interchangeable concepts [28]. Related Jaccard-type and interval-valued fuzzy–soft comparisons have also been developed [24,25,29]. Most directly, Nguyen and Kreinovich [30] defined an interval of Jaccard values over the embedded type-1 fuzzy sets and Wu and Mendel [31] developed more efficient endpoint algorithms for the same general construction. Bazhenov and Telnova [32] addressed joint and non-joint interval data under a different generalized interval formulation.
The independent-coordinate benchmark used below is not claimed as new. Positive coordinate weights can be absorbed into the coordinates, and its endpoint calculation is a weighted fractional-programming rederivation of established embedded-set Jaccard optimization. Classical rectangular linear-fractional sensitivity is also established: Mészáros and Rapcsák [33] gave an algorithm, and the Dinkelbach transformation supplies the standard ratio optimization principle [34]. The new object studied here is narrower: all interval coordinates move with one common policy parameter. Jaccard’s coordinatewise min–max switches then create a finite set of policy breakpoints, and pairwise rankings can change only at explicitly determined tie events. The common path is thus a restricted admissible-set assumption, not an empirically recovered dependence model and not a special case claim about the signed joint/non-joint measure of Bazhenov and Telnova [32].
These comparisons concern different optimization objects. The rectangular linear-fractional problem of Mészáros and Rapcsák [33] varies several admissible inputs independently. Here a prescribed line restricts those inputs, but the Jaccard minimum and maximum still switch along it. Our single-path sweep is not a speedup over that classical bound; the additional task is to enumerate all mutual tie events of the reported Jaccard paths.
For proper intervals , with a positive hull width, the symmetric consistency index of Bazhenov and Telnova [32] is
Its numerator may be negative for disjoint intervals, using the signed width of a generalized intersection. By contrast, our coefficient compares non-negative membership realizations and remains in . For example, , give , whereas their one-coordinate common policy realizations give . Thus “joint/non-joint” concerns interval compatibility, not synchronization by a common parameter. The two measures are mathematically distinct; our dependence restriction is relative to the embedded type-1 box benchmark.
Computational literary research provides the second methodological background. Character network and personality-oriented studies show both the value and the interpretive difficulty of representing fictional characters [1,2]; sentiment and emotion profiling provide scalable alternatives [3]. Annotation-centered work, including ProppLearner and character emotion corpora, emphasizes explicit textual units, layered annotation, and provenance [18,19]. These studies motivate the frozen segmentation and occurrence-level registry used here, while the present small-corpus approach is not proposed as a replacement for scalable NLP.
Scholarship on The Saint of Incipient Insanities has emphasized identity, multilingualism, belonging, migrancy, transculturalism, home, and nation [35,36,37]. Later studies have examined food and cross-border identity [38] and the relation of marginality, nonconformity, and insanity discourse [39]. These works motivate the literary domains but are not treated as preassigned character labels.
In summary, the established literature supplies interval-valued fuzzy–soft profiles, embedded-set Jaccard bounds, fractional sensitivity tools, and annotation practices. The contribution developed next is the exact analysis of a fully coupled one-factor path through the interval box: its extrema, its finite tie events, and the induced ranking partition. The literary registry then supplies one auditable proof-of-concept input to that mathematical object.
3. Preliminaries
Only the established objects used in the subsequent Jaccard analysis are recalled. A soft set is a parameterized mapping [4], while a fuzzy soft set replaces each crisp subset by a fuzzy membership function [5,40]. The standard soft set calculus and pointwise fuzzy operations are not repeated because they are not used in the proofs below [41]. Interval-valued fuzzy sets replace a single grade by a closed interval in [42].
Definition 1
(Interval-valued fuzzy soft set). Let X be a finite universe, a finite parameter set, and the family of interval-valued fuzzy subsets of X. An interval-valued fuzzy soft set is a mapping
with [21]. Degenerate intervals recover ordinary fuzzy grades. Only interval endpoints and scalar min–max comparisons are needed below.
Definition 2
(Weighted generalized Jaccard coefficient). For non-negative vectors and positive weights ,
when the denominator is positive; by convention, . This is the non-negative vector intersection-over-union form of the classical Jaccard idea [43].
The independent coordinate benchmark later optimizes Equation (1) over an interval box. The standard fractional programming transformation replaces a ratio , , by ; at an extremal ratio, the corresponding auxiliary optimum is zero [34]. This established principle is used only for the benchmark. The one-factor path introduced immediately below has a different feasible set and admits a direct finite breakpoint characterization.
4. Methods
4.1. Fully Coupled One-Factor Jaccard Path: Core Theory
The main mathematical object is introduced before the corpus-specific construction. Suppose objects have interval profiles , , and let be fixed coordinate weights. Put
The same scalar is used for every record-derived coordinate. It defines a fully coupled one-factor policy path: uses all lower endpoints, uses all upper endpoints, and intermediate values remove the same proportion of every endpoint discount. This is an imposed policy coupling, not an estimate of empirical dependence.
For a pair , define
assuming positive union mass. Whenever , the coordinate order can change only at
so we define
with repeated values removed.
Theorem 1
(Exact one-factor coherent-path range). For a pair with positive union mass on , is continuous and piecewise linear-fractional and
Thus a continuum of fully coupled scoring realizations is solved by at most endpoint/breakpoint evaluations. A direct implementation is ; sorting the crossings and updating the active affine coefficients gives an sweep.
Theorem 2
(Finite characterization of ranking transitions). Let r and s be two pair paths with positive union mass. On every open interval between consecutive points of , write
All interior ties, and hence all possible strict rank reversals, are roots in that interval of
a polynomial of degree at most two. If , the two paths are tied throughout that subinterval. All ranking transitions can be identified by testing finitely many coordinate breakpoints, endpoints, linear/quadratic roots, and persistent tie intervals. A candidate event need not reverse a strict ranking; no grid over λ is required.
The proof ingredients are standard, consisting of fixed min–max patterns, differentiation of a linear-fractional branch, and cross-multiplication. The contribution claimed here is the Jaccard-specific event construction that combines coordinate switches, endpoint ties, isolated roots, persistent tie intervals, and the resulting exact finite ranking partition under one shared policy. Proofs, the finite-family bound, and an interior transition example are given after the application-specific profile and benchmark constructions.
4.2. Corpus and Analytical Units
The sole primary corpus is the 2004 Farrar, Straus and Giroux edition of The Saint of Incipient Insanities [20]. The analytical corpus is restricted to printed pages 3–351. Cover copy, publication data, epigraphs, acknowledgments, and author material are excluded because they do not belong to the narrative. The opening sequence begins with “Started Drinking Again” on printed page 3, and the four subsequent macro-sections, The Crow, The Stork, Birds of a Feather, and Destroying Your Own Plumage, begin on printed pages 31, 71, 271, and 301, respectively. Each boundary was checked directly against the page images rather than inferred from OCR.
The analysis concerns six recurrent characters: Omer, Abed, Piyu, Gail/Zarpandit, Alegre, and Debra Ellen. The focal set was fixed for the formal registry as a purposive case-study set: the six characters are recurrent, traceable individually across the narrative, and traceable collectively cover the five reporting domains. No frequency threshold, similarity result, or post hoc ranking criterion was used to admit or exclude characters. Surface forms are normalized only when the text establishes identity equivalence. Thus, Omer and Omar are retained as forms of the same character, as are Gail and Zarpandit. OCR output is used only to locate candidate passages. The protocol requires character identity, paragraph boundaries, and evidential decisions to be checked against the corresponding page image because optical character recognition does not reliably preserve diacritics, line breaks, or speaker attribution.
The intended annotation unit is a character observation paragraph: a narrative paragraph with an explicit focal character or an unambiguous referent. The recovered computational inventory, however, contains registered text units whose stored boundaries are not uniformly complete orthographic paragraphs; page-bounded fragments and extraction-based locators remain. For reproducibility, the present calculations retain the same 2282 unit identifiers and character exposure counts throughout. A multi-character unit can have a separate inventory record for each focal character. The source locator combines macro-section, printed page, and archived fragment index; the index is not guaranteed to equal visible paragraph numbering. The mathematical results are conditional on this registered-unit inventory, not on a newly verified full-paragraph census.
Corpus segmentation relies on two complementary controls. First, deterministic name and alias matching identifies high-precision candidates, while immediately adjacent referential continuations and possible continuations across page breaks are retained as separate candidate classes rather than accepted automatically. The protocol intends continuations of one orthographic paragraph to be joined, but the surviving merge dispositions do not establish that this was completed consistently in the stored inventory. No new segmentation or denominator correction is claimed in the present calculations. Second, the protocol calls for a page image review of the analytical corpus, including pages for which name matching returns no candidate. The recovered documentation does not establish completion of every human check; the available audit records and their limits are specified below. This page-level audit allows unambiguous pronoun and dialogue speaker references to be added and any correction to a paragraph boundary to be recorded. To control recall before manual adjudication, all name-free paragraphs on the remaining pages are also subjected to a full-corpus contextual screen. The screen records pronoun- and dialogue-bearing paragraphs situated within a fixed three-paragraph window of an explicit focal-character reference. Characters recovered from that window are treated only as review hypotheses, because conversational turns can reverse the apparent speaker–addressee relation. Each hypothesis therefore retains a paragraph digest, a short locating anchor, and a field for contextual resolution. These fields document the intended review workflow; their existence is not evidence that every referent and boundary received completed human adjudication. Review proceeds from single-character hypotheses to multi-character scenes. This ordering reduces avoidable ambiguity without treating a sole machine proposal as evidence of correct reference; embedded letters, journals, survey items, and alternating dialogue turns are checked explicitly before any provisional addition or page-boundary merge is recorded. In multi-character scenes, a nearby name is not by itself sufficient for inclusion. The contextual pass distinguishes recoverable pronominal referents, locally bounded dialogue participants, orthographic continuations, and unresolved group exchanges. The last category is retained in a separate ambiguity queue for extended-context review. In the corpus preparation pass, these scenes are resolved only when the visible page and adjacent scene establish the speaker, direct addressee, or referent; otherwise they are excluded rather than distributed across all nearby characters. The recovered page-audit and precision-audit forms contain 349 and 272 rows, respectively, but their human decision fields are blank. The precision form comprises 75 explicit-name candidates, 124 adjacent co-reference candidates, and 73 cross-page candidates. Because the earlier fixed-seed description is not supported by the recovered archive, we withdraw it as evidence of a completed reproducible audit: neither the seed nor a human detected-error rate can be recovered. The 69-row ambiguity file records operational dispositions of 54 ADD, 13 DO_NOT_ADD, and two MERGE_WITH_PREVIOUS cases. All 54 resolved ADD character assignments occur in the final inventory, and neither merge fragment appears as a separate inventory unit; however, the separate human confirmation field is blank for all 69 rows. We report these as recovered disposition records, not as independently documented human adjudications. They are distinct from the 1250-candidate parameter ledger (59 retained and 1191 rejected in the archived candidate ledger) and the targeted 30-case second-reader assessment. The analyzed candidate ledger contains 1252 decisions: 60 retained and 1192 rejected. The supplement preserves the available forms, counts, dispositions, and reconstruction code. No claim of a historical audit error rate or complete independent corpus verification is made.
4.3. Codebook and Annotation Procedure
The codebook combines themes established in peer-reviewed criticism of the novel with concepts verified in the primary text. Parameter construction proceeded from literature-guided themes and preliminary close reading to the 17 operational definitions and exclusion boundaries shown below. This codebook was fixed for the formal retained-evidence scoring pass and is used unchanged in all reported computational comparisons. The surviving project archive does not contain a dated protocol version or complete pilot-coding history, so we do not claim formal preregistration or that every conceptual label was fixed before all exploratory reading. The reproducible temporal boundary is the start of the formal registry-scoring pass, not the earlier exploratory reading stage. Seventeen operational parameters are indexed in five mutually exclusive reporting domains: identity, belonging, and migration-related language (); religious and cultural orientation (); psychological and self-destructive patterns (); behavioral regulation and avoidance (); and feminist participation and sexual identity positions (). Identity, belonging, multilingualism, migrancy, and transculturalism are supported by prior readings of the novel [35,36,37]; food-related coping and homemaking are informed by Atik [38]; and nonconformity is treated as a narrative and social category rather than a psychiatric diagnosis [39]. Table 1 fixes the operational meaning and principal decision boundary of each parameter. The labels describe modes of literary representation and are not clinical constructs. Exclusive placement is an accounting convention for weighting and reporting, not a claim of conceptual independence. In particular, social inhibition , represented affective oscillation , and represented suicidal behavior can co-occur in one narrative sequence while remaining distinct codebook decisions.
Table 1.
Operational character parameter codebook.
For the analyzed registry, the recurrence condition in is clarified at the registered-unit level: multiple acts or an explicit recurrence expression must occur within that unit and concern suicidal orientation. Surrounding context resolves identity and speaker, not a missing recurrence criterion. Repeated overdose without established suicidal orientation does not suffice. This unit-level rule was applied consistently to the records used in the analysis.
The empirical case study combines rule-assisted primary coding, a targeted second-reader assessment, and deterministic computation. A single primary coder selects the retained evidence in context, records the supporting rationales, and assigns strength and contextual-confidence grades. The targeted second-reader assessment examines character attribution, contextual interpretation, codebook fit, and grades in a 30-case subset. Its observations are kept separate from the primary registry; it is not a second complete coding of the corpus. The assessment scope and observed disagreements are reported below. Computational checks verify the resulting profiles and numerical outputs. Parameter-specific lexical rules first create a high-recall review queue. Each queued character–unit–parameter record is then judged against the frozen operational definition, character referent, and surrounding textual context. A lexical cue does not automatically determine acceptance. Retained evidence is listed in a decision registry; all remaining pairs receive the computational indicator . This zero means only “not retained as positive evidence under this codebook and coding pass.” It does not distinguish an irrelevant unit, no opportunity to assess , an unresolved unit, or counter-evidence. Accordingly, the model estimates positive-evidence density; it makes no claim about bipolar evidence or missingness. The unit identifier, page, paragraph digest, source type, strength, confidence, and copyright-safe paraphrase are preserved for every retained record. The deterministic aggregation rules reproduce the numerical matrices from this registry. This computational repeatability should not be confused with independent interpretive reliability. The complete coding records, provenance-aware decision ledger, derived profiles, sensitivity outputs, and reproducibility code are provided in the Supplementary Materials.
For each retained record, the procedure assigns the anchored grades
Strength indicates how directly the unit realizes the parameter, not the psychological severity of a trait. The four strength anchors denote minimally admissible indirect evidence, contextually clear but inferential evidence, direct and unambiguous evidence, and direct evidence that is both central and elaborated within the unit. Contextual adjudication confidence records the primary coder’s assessment of how securely the parameter assignment fits the passage in context; it does not repeat the strength judgment and is not a measure of agreement between the authors. Strength and confidence are left blank for nonretained records. The numerical anchors are a transparent rational coding convention, not a validated psychometric scale. Equal numerical spacing does not demonstrate equal interpretive distances. For a fixed strength, retains one half, three quarters, or all of that strength in the lower endpoint. These are scoring discounts, not estimated probabilities of a correct reading. The pre-adjudication primary registry used strength levels and confidence levels . The final adjudicated registry retains the same strength levels and additionally uses for one record; the declared anchor remains unpopulated. Consequently, the registry alone cannot validate every proposed level.
Two records in Table 2 illustrate the distinction. A00105/U00007 (Abed, , printed 4:13) has , contributing before exposure normalization; the retained linguistic evidence is assigned less than maximal strength without an additional confidence discount. A33009/U01942 (Debra Ellen, , printed 291:4) has , contributing . The same strength is retained, but contextual uncertainty reduces only the lower endpoint. After adjudication, A35465/U02087 (Omer, , printed 317:1) uses , illustrating the strongest confidence discount actually used in the final registry. These examples explain the registered convention; they do not establish agreement on the grades. Each record additionally preserves the printed page, macro-section, normalized character identity, observed surface form, evidence source (narration, self-report, behavior, or attribution by another character), and a copyright-safe paraphrase. The dataset does not reproduce extended passages from the novel.
Table 2.
Complete embedded registry of the 60 retained evidence decisions in the analyzed dataset.
4.4. Embedded Decision Registry
Table 2 reproduces every retained decision used in the calculations. The first field combines the stable annotation and inventory identifiers; “page:paragraph” gives the printed-page location and the archived within-page fragment index. Source abbreviations are Narr. (narration), Self (self-report), Beh. (observed behavior), and Other (attribution by another character). The rationales are copyright-safe paraphrases rather than quotations. Nonretained unit–parameter entries are not adjudicated negative observations. They are computational zeros outside this positive registry and cannot be used to infer counter-evidence or absence of the represented characteristic.
In the current registry, 14 records have : A00363, A00895, A04546, A04582, A05149, A08487, A08521, A11924, A11925, A13087, A23011, A24472, A33009, and A35465. The original pre-adjudication second-reader sample had deliberately included all ten records then carrying ; later adjudication changed several grades and removed A11926. Removing all 14 current records is reported as the strict-confidence sensitivity analysis below. This diagnostic is not a substitute for blinded recoding or corpus-wide inter-reader agreement.
The primary registry, the blinded second-reader form, and the source-adjudication history are preserved separately. The assessment contains 30 targeted cases: 20 of the original 59 retained records (all ten records then carrying , plus ten with ) and ten nonretained near-miss cases. The second reader completed the targeted judgments without access to the primary retain/nonretain decisions, original grades, or downstream numerical results. The archived first-pass form is retained separately from the source-adjudicated decisions. The verified comparison yields 28/30 decision matches: 18/20 among the originally retained sample and 10/10 among nonretained boundary controls. Among the 18 cases retained by both readers, exact strength agreement is 9/18, exact confidence agreement is 13/18, and the complete pair agrees in 8/18. These are descriptive results for a purposive enriched sample, not corpus-wide reliability estimates. No time-separated intra-rater repeat is available. Final source-based adjudication is reported in Section 5.5; computational verification remains distinct from literary adjudication.
Internal verification separates record consistency from interpretive verification. Every retained record carries an inventory key, source locator, codebook parameter, and rationale. Computational checks test key consistency, duplicate unit–parameter combinations, valid grade ranges, aggregation identities, and numerical outputs. A stored digest or valid identifier is not proof that the source interpretation has been independently verified. Second, the complete 2282 × 17 = 38,794 unit–parameter matrix, all 102 interval profiles, and all 15 pairwise ranges are computed from the decision registry rather than transcribed into the manuscript. The supplementary full ledger materializes all 38,794 cells while preserving provenance: 60 are retained positives, 1192 are recorded screened rejections, and 37,542 did not enter the recorded parameter candidate queue and are are labeled as not separately adjudicated rather than retrospectively rejected. Third, decision sensitivity is quantified by (i) excluding all records with , and (ii) deleting each of the 60 current retained records in turn and recomputing the complete pairwise analysis. These checks expose dependence on the analyst’s uncertain or influential decisions; they do not estimate how another reader would represent the novel.
4.5. Interval-Valued Fuzzy Soft Scoring Profiles
Let be the focal characters and the fixed parameter set. The representation used below is the established interval-valued fuzzy soft set recalled in Section 3; no new fuzzy set species is defined.
For character , let be the frozen set of character observation units. For parameter , let indicate whether unit u is retained as evidence under the codebook. When , the record has strength and contextual adjudication confidence . Its membership contribution is the interval
Thus, a unit outside the positive evidence registry contributes to this particular positive support score. That computational value does not mean that the unit contains counter-evidence or that the parameter was assessable and found absent. For a retained unit, the lower endpoint discounts the evidence strength by the stated confidence, whereas the upper endpoint asks what the same strength would contribute under full acceptance of that contextual decision. The interval is a scoring-scenario interval: its endpoints encode two declared scoring rules. The construction does not assert that an unknown true membership is bounded by and s.
The interval-valued fuzzy soft membership of under is
The denominator is registered character unit exposure. Consequently, measures evidence density among all frozen observation units for character ; it is not a claim that every such unit offers a parameter-specific diagnostic opportunity. A genuine parameter-opportunity denominator would require an additional independently adjudicated mask , with . No such mask was coded, so the application neither estimates nor imputes .
Proposition 1
(Admissibility, partition consistency, and replication). For every ,
If is a partition into nonempty blocks and is computed inside block , then
where interval addition and non-negative scalar multiplication are applied endpoint-wise. Uniformly replicating every observation unit leaves the interval unchanged.
Proof.
For every retained record, gives ; for a nonretained unit both endpoints are zero. Summation and division by prove admissibility. Splitting each endpoint sum over the disjoint blocks proves Equation (8). Uniform replication multiplies each endpoint numerator and its denominator by the same positive integer. □
The interval width has a direct algebraic interpretation:
It is zero precisely when every retained contribution has full confidence (or no evidence is retained). Evidence weakness and the analyst’s confidence discount are not conflated: s controls both endpoints, while controls only their separation. The word “confidence” names an input grade; it does not turn into a statistical uncertainty estimate.
4.6. Independent-Coordinate Jaccard Range
Let , where and . The primary application uses , treating each operational parameter as the same accounting unit in the absence of externally calibrated importance weights. This is an equal-parameter benchmark, not a claim that the five domains are equally represented or equally important. The domains contain and 3 parameters, so their total weights are . The alternative equal-domain specification assigns each domain mass :
For character , the set of interpolated scoring realizations inside its interval-valued fuzzy soft profile is the box
For and , define
whenever the denominator is positive. The lower-endpoint score is . The full similarity scenario range is
For a reporting domain , the separately reported domain-specific coefficient is
when the domain union is positive; otherwise, it is reported as not available. This coefficient conditions on one reporting domain and is not substituted for the overall coefficient in Equation (11). This definition differs from applying Jaccard separately to lower and upper endpoint vectors. The minimum or maximum ratio can combine different realizations across parameters.
The convention belongs to the mathematical coefficient but is not used as a literary claim. If the weighted intersection mass is and the union mass is , then . This is the situation for the eight zero-overlap pairs in the application: at least one profile is nonzero, but the two profiles share no positive coordinate. If two profiles were both identically zero, the mathematical convention would return 1; substantively, the study would report that case as “not comparable from positive evidence” rather than as maximum character similarity. The positive-union assumptions in the theorems exclude only this latter zero-union case.
For two coordinate intervals and , define the finite candidate set
Repeated points are ignored. For , put
Theorem 3
(Weighted embedded-set Jaccard range: fractional form). Assume that every scoring realization in has positive union mass. Then, and in Equation (11) are, respectively, the zeros of and on :
Both auxiliary functions are continuous, strictly decreasing, and piecewise linear; therefore, each has a unique zero in . Consequently, bisection evaluates the exact extremal characterization to tolerance ε in arithmetic time. Each auxiliary function evaluation examines at most six candidate points per coordinate, and the number of bisection iterations is .
Proof.
For a fixed coordinate j, the rectangle is divided by the diagonal into at most two polygons. On either polygon, is linear, so its minimum and maximum are attained at polygon vertices. These vertices are exactly the four rectangle corners together with the endpoints (when present) of the diagonal segment inside the rectangle; hence, they form .
For any , separability of the Cartesian box gives
and the analogous maximum equals . The standard fractional programming equivalence [34] states that the minimum ratio is the value of at which the minimum auxiliary objective is zero; the same argument with maxima gives the upper ratio.
It remains to establish uniqueness rather than only monotonicity. Let and , where . Compactness and the hypothesis imply . If and if minimizes the auxiliary objective at , then
If maximizes the auxiliary objective at , then
Thus, both functions are strictly decreasing. They are finite sums of pointwise extrema of finitely many affine functions, hence continuous and piecewise linear. Moreover, and , because . Existence and uniqueness of both zeros follow, and bisection gives the stated numerical complexity. □
Example 1
(Endpoint substitution can miss both extrema). Consider two equally weighted coordinates with
The lower-vector and upper-vector substitutions respectively give and . They are not the attainable bounds. Choosing and gives , whereas gives . Theorem 3 yields the exact range . This example explains why evaluating only the lower, upper, or midpoint profiles cannot replace the fractional optimization.
Corollary 1
(Scenario-robust ordering). For two character pairs and , if
then every scoring realization of the first pair has larger Jaccard similarity than every scoring realization of the second pair.
Proof.
Every scoring realization for the first pair is at least , and every scoring realization for the second is at most . □
Theorem 3 is a weighted fractional-programming rederivation of the embedded-set Jaccard range studied by Nguyen and Kreinovich [30] and Wu and Mendel [31]. It is retained because it supplies the broad independent-coordinate benchmark used in the empirical analysis. The characterization is exact; the reported decimal endpoints are bisection approximations with tolerance . Since positive weights can be absorbed into the coordinates, weighting is not claimed as a separate mathematical novelty.
4.7. Policy Interpretation, Coherence Gap, and Exact Ranking Procedure
The interval profiles in Section 4.5 arise from two record-level scoring policies. For a retained-record indicator , evidence strength , and contextual confidence , define
Aggregating over the frozen observation inventory gives exactly the path defined in Section 4.1:
Hence, the one-factor construction removes the same proportion of every contextual-confidence discount. It deliberately represents perfect positive policy dependence, and is neither a general dependence model nor the only possible coherent policy.
A source-group extension replaces in Equation (16) by . If is the recorded source class, put
Then, , and the common-policy path is exactly the diagonal of . This representation still couples records within each source class. The resulting four-parameter family allows source classes to respond differently. It is not estimated here; the one-factor path is retained as the simplest fully coupled sensitivity model. For a concrete restriction in the literary registry, narrator-based A13087 (Piyu, ) and other-character A33009 (Debra Ellen, ) both have . A common cannot fully accept one while keeping the other at its discounted endpoint. Thus any one-factor stability result covers only a diagonal slice of these source-specific policies.
Proof of Theorem 1.
Each difference is affine and either has no zero, is identically zero, or changes sign once at . Between consecutive points of , the argument selected by every coordinatewise minimum and maximum is fixed. Hence,
for which the derivative is and so has constant sign. Each branch is monotone or constant, so extrema occur at branch endpoints. Continuity and Equation (4) follow. Sorting at most n interior crossings costs ; updating the active affine coefficients yields the stated sweep complexity. □
Because the coherent path is a subset of the Cartesian product used for the independent-coordinate range, we have the following containment.
Corollary 2
(Containment and coherence gap). For every pair r,
Consequently,
measures only the additional width allowed by coordinatewise-incoherent scoring choices; it is not an endpoint or midpoint displacement measure.
Proposition 2
(Zero coherence-width gap). Let , and let and be the argmin and argmax sets of the Jaccard coefficient on . Then
Thus, the width gap vanishes exactly when both independent box extrema are attainable on the one-factor path.
Proof.
Corollary 2 gives nested closed intervals. Their widths are equal if and only if both endpoints coincide, which is equivalent to the path attaining the corresponding box minimum and maximum. □
Proof of Theorem 2.
Theorem 1 gives linear-fractional representations on each common subinterval. Positive denominators make equality equivalent to cross-multiplication, producing Equation (5). Since both paths are continuous, a strict order can change only through equality. Testing all subinterval endpoints and solving the resulting linear/quadratic equations finds every isolated transition; records a persistent tie interval. □
A ranking cell is a maximal open interval on which the complete weak order of the reported pair similarities is unchanged. Persistent ties belong to the weak order of a cell and do not split it; coordinate breakpoints and isolated tie roots are boundary events. Not every tie is a reversal: an isolated root can be tangential, and a coordinate switch can leave all rankings unchanged. The event construction first yields a finite refinement. Adjacent intervals are merged into maximal ranking cells only if their weak orders and the weak order at the intervening event agree. This separates computational breakpoints, transient ties, persistent ties, and genuine order reversals.
Let P be the number of reported pair profiles. If all unordered pairs of m objects are included, , and the number of distinct interior coordinate breakpoints satisfies .
Theorem 4
(Finite-family ranking partition and complexity). For P reported pairs on n coordinates with positive one-factor union mass, let be the number of distinct interior coordinate breakpoints. The complete ranking partition contains at most
open ranking cells. Its event set can be constructed in
arithmetic and sorting operations, where is the number of retained distinct breakpoints and comparison roots. Explicit evaluation and sorting of all P scores at each event and open cell adds operations. These are arithmetic-operation bounds, not bit-complexity bounds.
Proof.
There are open base intervals. On each, every comparison of two paths is governed by one polynomial of degree at most two by Theorem 2; hence, contributes at most two interior roots unless the paths are identical on that interval. Inserting those roots and the K breakpoints gives Equation (17). Sorting at most crossings costs , solving all pairwise equations costs , and sorting retained events costs . Equal breakpoints are updated as one batch, and multi-pair ties require no additional asymptotic term. □
The complexity bound becomes practically relevant when many objects or coordinates are compared because makes the pair-comparison term quadratic in P. The present literary example has only and , with no interior coordinate switches, so it is an illustration of exactness rather than a performance benchmark. For scale, all pairs of six objects give 15 paths and 105 path–path comparisons, whereas 100 objects give 4950 paths and 12,248,775 comparisons before subdivision into branches. These counts explain when the finite-event organization matters; they are not measured running times. The evaluation bound assumes that each path’s active affine coefficients have been stored or maintained by the sweep, so evaluating a branch is constant-cost in the arithmetic-operation model.
Example 2
(Interior extremum and two exact rank transitions). The following mathematical example is separate from the literary data. For two equally weighted coordinates, consider
The first coordinate changes order at , whereas the second does not cross on . Hence,
The exact coherent maximum is , larger than both endpoint values and . Comparing this path with the constant similarity yields exact ties at and , so the pair ordering reverses twice. This example activates both the coordinate-switch and exact rank-transition mechanisms of Theorems 1 and 2.
Figure 1 shows the two similarity paths and marks the coordinate breakpoint, the interior maximum, and the two exact tie points described above.
Figure 1.
Exact event structure in the mathematical example. The curves show the interior coordinate breakpoint at , the one-factor maximum, and the ranking ties at and . The marked points are exact event values, not empirical estimates. The blue solid curve represents , the orange horizontal line represents , and the gray dashed vertical line marks the coordinate breakpoint . The black markers indicate the exact event points.
4.8. Internal Verification and Robustness Analysis
Let , let
and define the discounted and fully accepted positive-support totals
For brevity, whenever a lower-endpoint sensitivity vector is reported. The lower-endpoint profile is compared with three transparent baseline representations:
These are, respectively, binary presence, raw retained-record count, and parameter-wise log-normalized support. The normalization places each parameter on relative to the largest retained-record count observed for that parameter; it does not alter the zero/nonzero pattern. Their generalized Jaccard coefficients use the same parameter set and equal parameter weights as the primary profile.
Because denominator choice can change the interpretation, it is tested separately. In addition to the primary character exposure profile , four computable alternatives are evaluated:
The active-unit denominators are computed from the frozen unit identifiers:
Thus, multiple retained records attached to the same character–unit identifier are counted once in , and multiple retained records attached to the same character–unit–domain combination are counted once in . CA and DA remain in because their upper numerators cannot exceed the corresponding distinct active-unit counts. By contrast, and can exceed one. They are therefore non-negative baseline support profiles used only with the generalized Jaccard formula, not interval-valued fuzzy soft memberships. CA and DA are active evidence unit normalizers, not estimates of unobserved parameter opportunity. They show how strongly the chosen exposure denominator affects the result without introducing an opportunity mask after the outcomes are known. Equation (10) supplies the separate weighting sensitivity.
The exact range in Equation (11) tests all interpolated scores between the two stated scoring rules. Leave-one-parameter-out analysis additionally reports
so that a deletion range is not mistaken for a probabilistic interval or a direct measure of parameter importance.
Deterministic checks verify the record identifiers, parameter codes, grade ranges, duplicate keys, profile dimensions, interval admissibility, and the expected number of pairwise outputs. The computational analyses were implemented in Python 3.10 or later; the supplied reproduction scripts use no third-party packages. Algorithm 1 gives the full similarity-range calculation. Robustness is then examined at four levels. Representation sensitivity compares binary, raw-count, log-normalized, and exposure-adjusted specifications. Normalization and weighting sensitivity apply Equations (10) and (19)–(22). Parameter sensitivity leaves out one parameter at a time. Decision sensitivity removes the 14 current records with contextual adjudication confidence below one as a strict-confidence analysis and separately deletes each retained record in turn. The last two diagnostics go beyond varying scores within the constructed intervals: they remove a parameter or an evidential decision from the registry altogether. None of these procedures is presented as a test of human reliability or external criterion validity.
| Algorithm 1: Exact characterization and numerical evaluation of . |
Input: , weights , and tolerance .
Guarantee: and . |
The corresponding one-factor range and exact ranking procedure is summarized in Algorithm 2.
| Algorithm 2: One-factor ranges and exact ranking stability. |
|
5. Results
5.1. Evidence Records and Interval-Valued Profiles
The primary calculations in this section use the source-adjudicated registry of 60 retained records. The registry assigns the religious statement at 146:4 to Piyu, applies the unit-level recurrence rule for , and keeps the first-pass second-reader judgments separate from the adjudicated values. This separation allows the agreement figures to describe the blinded comparison while the numerical analysis uses the source-adjudicated evidence.
The frozen inventory contains 2282 character-observation units: 655 for Omer, 510 for Gail/Zarpandit, 412 for Abed, 307 for Alegre, 243 for Piyu, and 155 for Debra Ellen. The 2282 × 17 = 38,794 unit–parameter entries include 60 retained positives and 38,734 computational zeros. These zeros do not distinguish irrelevant, unresolved, unassessed, or opposing evidence. The total confidence-discounted strength is , whereas total full strength is ; their difference is before exposure normalization. The corresponding corpus-level totals and per-entry values are summarized in Table 3.
Table 3.
Corpus-level construction of the final adjudicated scoring intervals.
All 102 character–parameter memberships are valid intervals in . Fourteen retained records have , while 46 have . The confidence levels actually used are 0.50, 0.75, and 1, while the strength levels used are 0.50, 0.75, and 1. These remain assigned scoring ranges rather than statistical confidence intervals. Source adjudication assigns one religious statement to Piyu rather than Abed and changes several interval widths; the analyzed support graph contains seven positive pairs.
5.2. Character–Parameter Profiles
Table 4 distinguishes narrative exposure, retained occurrences, and graded evidence density. After adjudication, Abed has eight retained records, Alegre five, Debra Ellen seven, Gail/Zarpandit 17, Omer 12, and Piyu 11. Seven of Piyu’s eleven records concern , which remains the largest individual lower membership. These counts do not measure psychological intensity or literary importance.
Table 4.
Character exposure and retained evidence after final literary adjudication. The last column lists the interval(s) with the greatest lower membership.
The example still illustrates the two endpoints. Alegre’s three records each have , so its membership is . Debra Ellen’s two records have total strength 1.75 and discounted strength 1.5625, yielding . The width is the gap between discounted and fully accepted positive evidence, not a probability statement or a nonmembership grade. The resulting two interval constructions are summarized in Table 5.
Table 5.
Construction of two interval-valued profiles in the final registry.
5.3. Pairwise Similarities and Exact Ranking Transitions
Seven of the fifteen character pairs have positive similarity; eight remain identically zero. Piyu shares with Abed and Alegre through two retained Piyu– records, A13689/U00806 and A08521/U00502. This does not establish that the remaining pairs are dissimilar in a broader literary sense.
The Alegre–Debra Ellen calculation provides a direct numerical check. At the lower endpoint, the shared mass is
while the union mass is
Hence, . Its coherent range equals its box range, . The high within-behavioral-domain value of 0.9694 concerns only , and should not be read as broad character similarity. The complete independent coordinate and coherent path ranges for the seven positive pairs are reported in Table 6.
Table 6.
Independent coordinate and coherent path ranges, ordered by similarity at .
A direct comparison of the coherent path and independent coordinate ranges for the seven positive pairs is shown in Figure 2.
Figure 2.
Independent coordinate and coherent path similarity ranges. The markers indicate arithmetic range midpoints; they are not observations, expectations, or similarities at . The horizontal segments show deterministic ranges rather than confidence intervals. Blue open-circle markers and horizontal segments denote the independent-coordinate box ranges, whereas orange filled-circle markers and segments denote the coherent-path ranges. When the two ranges coincide exactly, the markers and segments overlap.
No within-pair membership-order crossing occurs in , so each final pair path consists of one linear-fractional branch and its extrema occur at or 1. Four positive coherent ranges attain the width of their independent-coordinate box benchmark. Three are strictly narrower: Abed–Alegre, Alegre–Piyu, and Debra Ellen–Gail/Zarpandit.
The complete seven-pair ranking has two exact interior transitions. The three relevant paths can be written as
Cross-multiplication after removing common positive factors produces two crossing polynomials,
Their unique roots in are
At , Alegre–Piyu crosses Abed–Piyu; at , Alegre–Piyu crosses Gail/Zarpandit–Omer. Thus, the weak order is
where AD, AA, AO, GO, AP, LP, and DG correspond, in the same order, to the seven positive pairs listed in Table 6. Testing all positive-pair comparisons finds no other interior tie. The eight zero paths form a permanent bottom tie. This is an exact ranking partition, not a grid-based assertion.
The three paths involved in these changes of rank are shown in Figure 3, where the two vertical lines mark the exact tie locations.
Figure 3.
Exact ranking transitions in the final registry. The two vertical lines mark the tie locations obtained analytically from Equation (32); the curves are shown to make the changes in order visible. None of the three paths has an interior membership-order crossing.
5.4. Representation, Normalization, and Deletion Sensitivity
The results for the seven positive pairs under the alternative representations are reported in Table 7.
Table 7.
Seven nonzero pairs under alternative representations. All use the same retained registry.
Binary, raw-count, and log-normalized profiles share the seven-link support graph by construction: they transform the same non-negative retained support and do not create new positive coordinates. This agreement is not independent validation. Among the seven positive pairs, Spearman’s relative to the lower-endpoint profile is 0.556, 0.342, and 0.286 for binary, raw-count, and log-normalized representations, respectively; Kendall’s is 0.514, 0.195, and 0.143. The coefficients are descriptive, with midrank and tie corrections. All-pair correlations are omitted because the eight common zero ties would inflate apparent agreement. The normalization and weighting sensitivity results are summarized in Table 8.
Table 8.
Normalization and weighting sensitivity. CE is character exposure; D and are support profiles; CA and DA use character-active and domain-active denominators; EW gives equal mass to each reporting domain. Correlations use only the seven CE-positive pairs.
The character-active denominators are , respectively, for Abed, Alegre, Debra Ellen, Gail/Zarpandit, Omer, and Piyu. These are not parameter opportunity counts. Character-active normalization places Abed–Alegre narrowly first (0.1533) ahead of Alegre–Debra Ellen (0.1521). Equal-domain weighting also places Abed–Alegre first (0.1892), followed by Alegre–Debra Ellen (0.1513) and Abed–Omer (0.1417). Domain-active normalization instead places Alegre–Debra Ellen first (0.3173), Alegre–Piyu second (0.2105), and Abed–Omer third (0.1735). Thus normalization and weighting change both magnitudes and portions of the ordering.
Five links depend on a single shared parameter: Abed–Alegre, Alegre–Piyu, and Abed–Piyu depend on ; Alegre–Debra Ellen on ; and Debra Ellen–Gail/Zarpandit on . Removing the corresponding parameter sets each affected similarity to zero. Abed–Omer and Gail/Zarpandit–Omer remain positive under every single-parameter deletion. The complete changes are in Table 9; positive changes can result from reduction of the union denominator and do not imply causal parameter importance.
Table 9.
Complete leave-one-parameter-out changes . AD: Alegre–Debra Ellen; AA: Abed–Alegre; AO: Abed–Omer; LP: Alegre–Piyu; GO: Gail/Zarpandit–Omer; AP: Abed–Piyu; DG: Debra Ellen–Gail/Zarpandit.
The strict confidence and leave-one-record-out results are summarized in Table 10.
Table 10.
Strict confidence calculations and 60 leave-one-record-out calculations. Ranks use midranks among the seven current-positive pairs. Max. change is the largest absolute change of either box endpoint.
Retaining only the 46 final records with preserves seven positive pairs but substantially changes their ordering: Alegre–Debra Ellen remains first (0.1564), followed by Abed–Omer (0.1447) and Gail/Zarpandit–Omer (0.1149), whereas Abed–Alegre falls to fourth (0.1085). Across all 60 single-record deletions, there are 13 distinct weak orders of the seven final-positive pairs. Alegre–Debra Ellen ranks first in 56 runs and Abed–Alegre in four. All seven positive links survive every single-record deletion in this registry. In particular, deleting either A13689 or A08521 leaves the other Piyu– record in place. Both Piyu links still vanish when parameter is omitted. Thus single-record support persistence does not imply parameter deletion robustness or invariance of similarity magnitudes and rankings.
5.5. Targeted Second-Reader Assessment and Final Adjudication
The targeted comparison uses 30 purposively selected cases: 20 records that were retained in the original 59-record registry and ten nonretained boundary controls. The retained sample contained all ten records that had at selection time plus ten controls. The second reader worked from the frozen codebook and source context without access to the primary retain/nonretain decisions, original grades, or downstream numerical results. The archived first-pass form is preserved separately from the source-adjudicated dataset. On the verified comparison, retain/nonretain agreement is 28/30: 18/20 in the originally retained subset and 10/10 among the nonretained boundary controls.
Among the 18 records retained by both readers, exact strength agreement is 9/18, exact confidence agreement is 13/18, and the complete pair agrees in 8/18; ten records differ on at least one grade. These counts describe this enriched sample and are not population estimates of corpus-wide reliability. The disagreements are not treated as ten interchangeable errors. Source adjudication identified a recurring calibration issue in several confidence judgments: the primary grade had sometimes relied partly on information available elsewhere in the corpus rather than on the coded unit and its local context. The final adjudication tightened this unit-level/corpus-level boundary.
For the ten mutually retained grade disagreements, source adjudication keeps the primary pair in three cases, adopts the second-reader pair in four, and uses an intermediate or otherwise third pair in three. The machine-readable supplement preserves the first-pass judgments and the adjudicated values as separate tables. Recomputing the analysis under the alternative sampled grades and recurrence boundary decision changes the number and location of ranking transitions. Thus, exact conditional mathematics does not remove sensitivity to interpretive input; it makes the numerical consequences of those inputs traceable. The primary 60-record dataset has seven positive pairs and two rank transitions at approximately 0.840521 and 0.868205.
5.6. Overall Interpretation of the Results
The results lead to four main conclusions. First, seven character pairs share positive evidence in the final registry, whereas the other eight do not share a positive coordinate under the adopted representation. The first three positions remain unchanged along the common-policy path, but the middle of the ranking does not. Second, none of the seven positive pairs has an interior membership-order crossing, yet the comparison between pairs still produces two exact ranking transitions, at and . Looking only at the endpoints would miss part of the ranking behavior. Third, stability depends on what is changed. Removing any single retained record leaves all seven positive links in place, but five links disappear when their only shared parameter is removed. The leading or middle positions also change under alternative normalization and weighting choices, and the ranking changes markedly when records with are excluded. Finally, the second-reader exercise found stronger agreement on whether a record should be retained than on its exact grades. The similarity ranges and transition points are exact for the final registry, but their literary meaning still depends on the coding and modeling choices used to build that registry.
6. Discussion
The representational and mathematical contributions are best considered separately. Interval-valued fuzzy soft sets are already established [21], and this study does not rename or redefine them. The application combines evidence strength with contextual adjudication confidence to build exposure-normalized membership envelopes. Proposition 1 shows that these envelopes are admissible and can be recombined exactly across corpus partitions. The independent-coordinate Jaccard envelope is likewise known from earlier interval-valued fuzzy research [30,31]; Theorem 3 gives a convenient weighted fractional rederivation, not a priority claim. The main methodological and mathematical contribution is the coherent-path construction. Theorem 1 proves that its exact similarity range is attained on a finite set of coordinate crossings, while Theorem 2 reduces every possible ranking change to linear or quadratic roots. Proposition 2 records the attainability criterion for coincident coherent and independent widths, and Theorem 4 supplies a finite-family ranking-cell bound and total complexity. Example 2 demonstrates that an interior crossing can determine the range and create two rank reversals, while the final adjudicated literary application contains two interior rank transitions but no within-pair membership-order crossing. These are different events: rank transitions can occur between smooth linear-fractional similarity branches.
In the case study, the interval-valued fuzzy soft representation keeps the provenance, exposure, evidence strength, and contextual confidence width visible instead of reducing the literary reading to an opaque character score. Each focal character is concentrated in a different subset of the fixed parameters. Therefore, the nonzero pairwise envelopes point to specific textual intersections rather than broad personality resemblance.
Normalization places an important limit on interpretation. Dividing by character exposure separates positive-evidence density from raw narrative prominence, but also dilutes a parameter with registered units that may offer no real opportunity to assess it. The sensitivity results confirm that this is a major modeling choice: character-active normalization places Abed–Alegre narrowly first and equal-domain weights also place Abed–Alegre first, whereas domain-active normalization places Alegre–Debra Ellen first and Alegre–Piyu second. Because none of these alternatives supplies the missing opportunity mask, the primary magnitudes cannot be read as parameter opportunity rates.
The generalized Jaccard coefficient prevents a misleading form of inflation. A comparison conditioned only on jointly covered parameters could make two characters appear highly similar because of one shared parameter while discarding all parameters supported for only one character. Equation (11) retains unmatched memberships in the denominator for every scoring realization. The Alegre–Debra Ellen range of , despite a lower-endpoint domain value of , separates a close local correspondence from a broad profile match.
The exact ranges are narrow, but some overlap. In the final adjudicated registry, Alegre–Piyu crosses Abed–Piyu at and Gail/Zarpandit–Omer at . The first three positions remain fixed, but the middle order changes across the two events. Three coherent ranges are strictly narrower than their independent-coordinate benchmarks: Abed–Alegre, Alegre–Piyu, and Debra Ellen–Gail/Zarpandit; the other four positive pairs attain the box width.
The two objects answer different questions. The box shows what is possible when coordinates can vary freely, whereas the path shows what remains possible under one common perfectly coupled scoring policy. The path is a one-factor sensitivity model, not a general account of dependence among records or coordinates. Introducing source-group parameters would be a principled next step. Parameter deletion is more consequential: five relationships disappear when one shared parameter is omitted. They are best described as parameter-specific bridges. The two remaining relationships persist under every single-parameter deletion, but their low overall values preclude strong claims of general similarity. The reported correlations concern only the seven positive pairs and describe agreement between representations of the same registry, not external literary validity.
The Awareness Logic of Ambiguity distinguishes valuation, contextual alignment, and restraint toward commitment [44]. Its restraint coordinate is not our confidence grade. This comparison underscores a limitation of the present positive-only score: unretained, withheld, and unassessed states cannot be reconstructed from zero. No such additional state model is implemented.
Several limitations set the boundaries of the analysis. The registry is based on rule-assisted primary coding with a targeted second-reader assessment. The 30-case comparison provides useful boundary and calibration evidence but is not a second full-corpus annotation: verified retain/nonretain agreement is 28/30, while the complete pair agrees in only 8/18 mutually retained cases. Source-based adjudication resolved the sampled disagreements, but neither intra-rater stability nor external literary validity is established. Several confidence disagreements exposed a recurring calibration issue between evidence contained in the coded unit/local context and knowledge available elsewhere in the corpus. The archived first-pass ratings, sensitivity scenarios, and final adjudicated registry are kept distinct. The interval widths reflect assigned contextual grades rather than sampling uncertainty or guaranteed bounds on a true membership; indeed, removing the less-certain records substantially reorders several positive-pair ranks. The parameter system is theory-informed but partly tailored to one novel, one edition, and six focal characters, and macro-section and source-type deletion have not yet been tested. The registry also contains no counter-evidence or parameter-opportunity mask. Its zeros combine several unmodeled states, and cannot support absence claims. For the same reason, the active-unit denominators remain sensitivity proxies rather than repairs for the missing opportunity information. The retained inventory also has unresolved segmentation and referent-audit limitations described in the corpus subsection. Exact recomputation on fixed identifiers does not validate those identifiers as a complete orthographic-paragraph census or establish invariance to resegmentation. The findings should be read as a traceable positive-evidence analysis under one declared protocol, not as a reader-consensus or opportunity-normalized account of the novel.
7. Conclusions
This paper uses the established interval-valued fuzzy soft set structure to separate confidence-discounted positive-evidence density from the density obtained when the same evidence is fully accepted. The resulting interval aggregation is admissible, partition-consistent, and invariant under uniform replication. Because earlier embedded-set work already established extremal Jaccard analysis for interval-valued fuzzy sets, the Cartesian-box result is used here as a benchmark. The main methodological and mathematical contribution is the coherent scoring path. Its exact similarity extrema occur at finitely many coordinate crossings, and its exact rank transitions reduce to breakpoint tests and equations of degree at most two. A zero-gap attainability criterion, global ranking partition bound, and fully worked interior transition example extend this finite analysis to within-pair extrema and between-pair rank changes, which need not occur at the same policy values.
The application to six characters within a single novel represents a proof of concept, not external validation. Across 2282 character observation units and 17 parameters, the analyzed registry retains 60 records and produces seven low positive similarities. Three coherent ranges are narrower than the independent-coordinate benchmark, and the common-policy path has two exact rank transitions at approximately and . Five links depend entirely on one shared parameter. All seven links survive the 60 single-record deletions, including deletion of either of the two Piyu– records; the two Piyu links still depend on parameter . Alternative denominators and equal-domain weights change the leading and middle ranks. The targeted 30-case assessment by a second reader yields 28/30 decision matches but only 8/18 exact joint grade matches among mutually retained cases; final source-based adjudication is reported separately from the blind comparison. These results limit the scope of any robustness claim and preserve the distinction between exact conditional mathematics and interpretive inputs.
Overall, the paper contributes an analysis of similarity path and ranking stability together with a transparent testing structure. While we report the effects of normalization, weighting, and fragile coding decisions, the empirical scope remains limited to the positive evidence in the documented registry.
Supplementary Materials
The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/math14183273/s1. File S1: Machine-readable supplementary archive containing the coding records, provenance-aware decision ledger, derived profiles, sensitivity outputs, and reproducibility code supporting the analyses reported in this study.
Author Contributions
Conceptualization, A.Y.; methodology, A.Y.; software, A.Y.; validation, D.Ö. and M.A. (adjudication of the sampled disagreements) and A.Y. (verification of the computational results); formal analysis, A.Y.; investigation, D.Ö. (primary corpus coding, evidence selection, and initial grading) and M.A. (targeted second-reader assessment); data curation, D.Ö. and M.A.; writing—original draft preparation, A.Y.; writing—review and editing, D.Ö., M.A., and A.Y.; visualization, A.Y. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The original contributions presented in this study are included in the article and the Supplementary Materials. Further inquiries can be directed to the corresponding author.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Labatut, V.; Bost, X. Extraction and analysis of fictional character networks: A survey. ACM Comput. Surv. 2019, 52, 89. [Google Scholar] [CrossRef] [Scilit]
- Picca, D.; Pitteloud, J. Personality recognition in digital humanities: A review of computational approaches in the humanities. Digit. Scholarsh. Humanit. 2023, 38, 1646–1658. [Google Scholar] [CrossRef] [Scilit]
- Jacobs, A.M. Sentiment analysis for words and fiction characters from the perspective of computational (neuro-)poetics. Front. Robot. AI 2019, 6, 53. [Google Scholar] [CrossRef] [Scilit]
- Molodtsov, D. Soft set theory—First results. Comput. Math. Appl. 1999, 37, 19–31. [Google Scholar] [CrossRef] [Scilit]
- Maji, P.K.; Biswas, R.; Roy, A.R. Fuzzy soft sets. J. Fuzzy Math. 2001, 9, 589–602. [Google Scholar]
- Maji, P.K.; Roy, A.R.; Biswas, R. An application of soft sets in a decision making problem. Comput. Math. Appl. 2002, 44, 1077–1083. [Google Scholar] [CrossRef] [Scilit]
- Roy, A.R.; Maji, P.K. A fuzzy soft set theoretic approach to decision making problems. J. Comput. Appl. Math. 2007, 203, 412–418. [Google Scholar] [CrossRef] [Scilit]
- Alcantud, J.C.R. A novel algorithm for fuzzy soft set based decision making from multiobserver input parameter data set. Inf. Fusion 2016, 29, 142–148. [Google Scholar] [CrossRef] [Scilit]
- Alcantud, J.C.R.; Khameneh, A.Z.; Santos-García, G.; Akram, M. A systematic literature review of soft set theory. Neural Comput. Appl. 2024, 36, 8951–8975. [Google Scholar] [CrossRef] [Scilit]
- Yolcu, A.; Ozturk, T.Y.; Bayramov, S. A new approach to soft relations and soft functions. Comput. Appl. Math. 2024, 43, 335. [Google Scholar] [CrossRef] [Scilit]
- Yolcu, A.; Öztürk, T.Y. Multi-criteria decision making for smart city implementation using fuzzy soft MAUT method. J. Anal. Uncertain. 2025, 1, 35–45. [Google Scholar] [CrossRef] [Scilit]
- Naz, M.; Shabir, M. On fuzzy bipolar soft sets, their algebraic structures and applications. J. Intell. Fuzzy Syst. 2014, 26, 1645–1656. [Google Scholar] [CrossRef] [Scilit]
- Zou, Y.; Xiao, Z. Data analysis approaches of soft sets under incomplete information. Knowl.-Based Syst. 2008, 21, 941–945. [Google Scholar] [CrossRef] [Scilit]
- Yolcu, A.; Smarandache, F.; Öztürk, T.Y. Intuitionistic fuzzy hypersoft sets. Commun. Fac. Sci. Univ. Ank. Ser. Math. Stat. 2021, 70, 443–455. [Google Scholar] [CrossRef] [Scilit]
- Yolcu, A.; Benek, A.; Öztürk, T.Y. A new approach to neutrosophic soft rough sets. Knowl. Inf. Syst. 2023, 65, 2043–2060. [Google Scholar] [CrossRef] [Scilit]
- Karataş, E.; Yolcu, A.; Öztürk, T.Y. Effective neutrosophic soft set theory and its application to decision-making. Afr. Mat. 2023, 34, 62. [Google Scholar] [CrossRef] [Scilit]
- Yolcu, A.; Öztürk, T.Y. Neutrosophic hesitant fuzzy soft sets: Algebraic structure, topology, and application in decision making. J. Ambient Intell. Humaniz. Comput. 2026, 17, 185–203. [Google Scholar] [CrossRef] [Scilit]
- Finlayson, M.A. ProppLearner: Deeply annotating a corpus of Russian folktales to enable the machine learning of a Russian formalist theory. Digit. Scholarsh. Humanit. 2017, 32, 284–300. [Google Scholar] [CrossRef] [Scilit]
- Dennerlein, K.; Schmidt, T.; Wolff, C. Computational emotion classification for genre corpora of German tragedies and comedies from the 17th to early 19th century. Digit. Scholarsh. Humanit. 2023, 38, 1466–1481. [Google Scholar] [CrossRef]
- Shafak, E. The Saint of Incipient Insanities; Farrar, Straus and Giroux: New York, NY, USA, 2004. [Google Scholar]
- Yang, X.-B.; Lin, T.Y.; Yang, J.-Y.; Li, Y.; Yu, D.-J. Combination of interval-valued fuzzy set and soft set. Comput. Math. Appl. 2009, 58, 521–527. [Google Scholar] [CrossRef] [Scilit]
- Yolcu, A.; Ozturk, T.Y. An adjustable method for decision making problems with neutrosophic soft sets. Int. J. Syst. Assur. Eng. Manag. 2025, 16, 613–621. [Google Scholar] [CrossRef] [Scilit]
- Feng, F.; Li, Y.; Leoreanu-Fotea, V. Application of level soft sets in decision making based on interval-valued fuzzy soft sets. Comput. Math. Appl. 2010, 60, 1756–1767. [Google Scholar] [CrossRef] [Scilit]
- Qin, H.; Wang, Y.; Ma, X.; Wang, J. A novel approach to decision making based on interval-valued fuzzy soft set. Symmetry 2021, 13, 2274. [Google Scholar] [CrossRef] [Scilit]
- Ali, M.; Kılıçman, A. On interval-valued fuzzy soft preordered sets and associated applications in decision-making. Mathematics 2021, 9, 3142. [Google Scholar] [CrossRef] [Scilit]
- Lu, Z.; Ye, J. Logarithmic similarity measure between interval-valued fuzzy sets and its fault diagnosis method. Information 2018, 9, 36. [Google Scholar] [CrossRef] [Scilit]
- Rico, N.; Huidobro, P.; Bouchet, A.; Díaz, I. Similarity measures for interval-valued fuzzy sets based on average embeddings and its application to hierarchical clustering. Inf. Sci. 2022, 615, 794–812. [Google Scholar] [CrossRef] [Scilit]
- Díaz-Vázquez, S.; Torres-Manzanera, E.; Díaz, I.; Montes, S. On the search for a measure to compare interval-valued fuzzy sets. Mathematics 2021, 9, 3157. [Google Scholar] [CrossRef] [Scilit]
- Liu, P.; Munir, M.; Mahmood, T.; Ullah, K. Some similarity measures for interval-valued picture fuzzy sets and their applications in decision making. Information 2019, 10, 369. [Google Scholar] [CrossRef] [Scilit]
- Nguyen, H.T.; Kreinovich, V. Computing degrees of subsethood and similarity for interval-valued fuzzy sets: Fast algorithms. In Proceedings of the 9th International Conference on Intelligent Technologies (InTech’08), Samui, Thailand, 7–9 October 2008; pp. 47–55. [Google Scholar]
- Wu, D.; Mendel, J.M. Efficient algorithms for computing a class of subsethood and similarity measures for interval type-2 fuzzy sets. In Proceedings of the 2010 IEEE International Conference on Fuzzy Systems, Barcelona, Spain, 18–23 July 2010; pp. 1–7. [Google Scholar]
- Bazhenov, A.N.; Telnova, A.Y. Generalization of Jaccard index for interval data analysis. Meas. Tech. 2023, 65, 882–890, Correction in Meas. Tech. 2023, 66, 288. https://doi.org/10.1007/s11018-023-02223-8. [Google Scholar] [CrossRef] [Scilit]
- Mészáros, C.; Rapcsák, T. On sensitivity analysis for a class of decision systems. Decis. Support Syst. 1996, 16, 231–240. [Google Scholar] [CrossRef] [Scilit]
- Dinkelbach, W. On nonlinear fractional programming. Manag. Sci. 1967, 13, 492–498. [Google Scholar] [CrossRef] [Scilit]
- Görümlü, Ö. Elif Shafak’s The Saint of Incipient Insanities: An issue of identity. Selçuk Üniv. Edeb. Fak. Derg. 2009, 21, 269–279. [Google Scholar]
- Öğüt Yazıcıoğlu, Ö. Who is the other? Melting in the pot in Elif Shafak’s The Saint of Incipient Insanities and The Bastard of Istanbul. Litera J. Lang. Lit. Cult. Stud. 2011, 22, 53–70. [Google Scholar]
- Öztabak-Avcı, E. Elif Şafak’s The Saint of Incipient Insanities as an “international” novel. ARIEL Rev. Int. Engl. Lit. 2007, 38, 83–99. [Google Scholar]
- Atik, E. Eating and cooking beyond the borders in Elif Shafak’s The Saint of Incipient Insanities. Textual Pract. 2024, 38, 1089–1104. [Google Scholar] [CrossRef] [Scilit]
- Sultana, M.F. The politics of (in)sanity: Elif Shafak’s The Saint of Incipient Insanities and 10 Minutes 38 Seconds in This Strange World. Rajshahi Univ. J. Arts Law 2024, 51, 107–117. [Google Scholar] [CrossRef] [Scilit]
- Zadeh, L.A. Fuzzy sets. Inf. Control 1965, 8, 338–353. [Google Scholar] [CrossRef] [Scilit]
- Maji, P.K.; Biswas, R.; Roy, A.R. Soft set theory. Comput. Math. Appl. 2003, 45, 555–562. [Google Scholar] [CrossRef] [Scilit]
- Zadeh, L.A. The concept of a linguistic variable and its application to approximate reasoning—I. Inf. Sci. 1975, 8, 199–249. [Google Scholar] [CrossRef] [Scilit]
- Jaccard, P. Étude comparative de la distribution florale dans une portion des Alpes et du Jura. Bull. Soc. Vaudoise Sci. Nat. 1901, 37, 547–579. [Google Scholar] [CrossRef] [Scilit]
- Edalatpanah, S.A. The Awareness Logic of Ambiguity (ALA): Triadic Interpretive States and Their Structure-Preserving Operational Core. J. Fuzzy Ext. Appl. 2026, 7, 661–702. [Google Scholar] [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.


