Highlights
Please indicate how your work links to systems science via your contributions to systems practice, theory, and/or methodology.
- Integrates topic discovery, sentiment analysis, exploratory clustering of complaint records, temporal monitoring, and module-specific uncertainty within a human-reviewed analytical workflow for large-scale consumer complaint analysis.
- Introduces a conceptual stock–flow and feedback formulation that distinguishes observed complaint records from the latent stock of unresolved consumer issues and situates analytical evidence within a potential organizational feedback process.
What are the main findings and/or the implications of the main findings?
- BERTopic assigned 80.46% of 99,434 CFPB complaint records to 89 non-outlier topics, while 19.54% remained explicitly unassigned; Cardiff RoBERTa showed greater agreement with human labels than FinBERT on the prediction-stratified benchmark.
- The exploratory K-means groupings showed weak geometric separation, and temporal patterns were descriptive rather than causal, supporting human-reviewed screening and monitoring rather than autonomous complaint prioritization or decision-making.
Abstract
Consumer complaint narratives are an underused source of business and regulatory evidence. This study presents a proof-of-concept analytical framework for multi-dimensional complaint analysis and uncertainty-aware evidence synthesis, with potential use in human-in-the-loop decision support, and applies it to 99,434 Consumer Financial Protection Bureau (CFPB) complaint records spanning from March 2015 to March 2026. The empirical workflow integrates BERTopic topic discovery, two pretrained sentiment checkpoints (ProsusAI/finbert and cardiffnlp/twitter-roberta-base-sentiment-latest), exploratory clustering of complaint records, and descriptive temporal summaries. A conceptual stock–flow and feedback representation links observed complaints, analytical alerts, response capacity, and unresolved issues; these relations are not causally estimated or simulated. BERTopic produced 89 non-outlier topics. A probability-based outlier-reassignment stage assigned 80.46% of records to these topics, while 19.54% remained unassigned and were retained as an uncertainty queue. A FinBERT-prediction-stratified benchmark of 600 complaint records was independently coded by two annotators (raw agreement = 87.5%; Cohen’s κ = 0.758). Against Annotator 1 as the prespecified reference, Cardiff RoBERTa aligned better with the labels on the prediction-stratified benchmark (κ = 0.395; sample accuracy = 67.3%) than FinBERT (κ = 0.195; sample accuracy = 46.3%), although neither model supports autonomous use. The four-group K-means solution is reported as exploratory because the highest observed Silhouette coefficient was only 0.106 and favored K = 2. Temporal patterns coincided with selected external events, but no causal effect is claimed. The contribution is therefore a transparent, uncertainty-aware analytical architecture whose potential organizational uses require human review and operational validation.
1. Introduction
Consumer financial protection has attracted increasing attention from researchers, regulators, and financial-service organizations [1,2]. Consumer complaints are also a valuable source of business intelligence because, unlike structured survey responses, they describe operational problems and lived experiences in consumers’ own words. These narratives combine domain terminology, informal language, emotion, and procedural detail. The growth of digital complaint channels, including the Consumer Financial Protection Bureau (CFPB) Consumer Complaint Database [3], has increased both the volume of available text and the need for transparent, scalable analysis [4].
Interest in linking systems thinking with consumer-behavior analysis reflects a desire to move from one-off description toward monitored, feedback-aware organizational learning [5]. Traditional complaint analysis relies on keyword search, manual coding, and descriptive statistics, which may miss semantic similarity and contextual differences. Recent natural language processing (NLP) methods can structure large collections of free text, but an analytical pipeline should not be equated with an implemented decision-support system unless its decision rules, feedback mechanisms, and organizational outcomes have been validated.
This study addresses three narrower gaps. First, complaint-mining studies often report topic and sentiment outputs separately, with limited attention to how their uncertainties should be combined. Second, finance-domain and stylistically contrasting sentiment checkpoints are rarely compared on the same independently annotated complaint benchmark. Third, temporal complaint summaries are often interpreted without explicitly separating observed complaint volume from the underlying stock of unresolved issues or from time-varying reporting and intake processes. The present study responds with an integrated but explicitly proof-of-concept architecture; it does not claim a deployed or adaptively validated system.
Building on these gaps, this study is guided by the following four research questions:
RQ1. What dominant themes and topic-coverage limitations emerge from transformer-based topic modeling of large-scale financial complaint narratives?
RQ2. How do a finance-domain checkpoint (FinBERT) and a stylistically contrasting social-media checkpoint (Cardiff RoBERTa) compare against human sentiment judgments on complaint narratives?
RQ3. What exploratory complaint-record groupings emerge when semantic representations are combined with sentiment features, and what do internal validation indicators imply about their stability and interpretation?
RQ4. How do complaint volume, topic prevalence, and model-assigned sentiment vary descriptively over time, and how can these observations be situated within a conceptual systems-feedback formulation without making causal claims?
The study analyzes 99,434 CFPB complaint records covering March 2015 to March 2026. It combines (1) BERTopic for topic discovery, (2) dual sentiment inference using FinBERT [6] and Cardiff RoBERTa [7], evaluated against human annotations, (3) an exploratory clustering of complaint records using PCA-reduced embeddings and K-means, and (4) temporal summaries at monthly, quarterly derived, and annual resolutions. The methodological contribution is not a new learning algorithm or an implemented decision-support system. Rather, relative to complaint-mining studies that examine analytical components largely in isolation, it lies in the coordinated integration of topic discovery, comparative sentiment evaluation, exploratory clustering, temporal description, and module-specific uncertainty within a single reproducible workflow that explicitly preserves human review. The system-dynamics component serves as a conceptual interpretation of how such analytical evidence could be situated within a future organizational feedback process; it is not an empirically calibrated or simulated model.
2. Literature Review
2.1. Consumer Complaint Analysis
The study of consumer complaints has evolved from rule-based and manually coded approaches toward statistical and machine-learning methods. Early work, such as Singh and Wilkes [8], organized complaint behavior around dimensions including severity, attribution, and expected resolution. Digital platforms then increased the scale and diversity of complaint records, creating demand for methods that can summarize large volumes of unstructured text while preserving interpretability.
Topic modeling is frequently used to analyze complaint and review data. Latent Dirichlet Allocation (LDA) [9] remains a common baseline, whereas embedding-based approaches such as BERTopic [10] combine contextual document representations, UMAP dimensionality reduction [11], HDBSCAN clustering [12], and class-based TF–IDF representations. This process allows identifying more detailed topics and semantically meaningful clusters without defining number of topics in advance. Recent comparative evaluations showed that BERTopic outperformed other topic modeling techniques such as LDA and NMF [13].
Within financial services, prior studies have applied text analysis to banking complaints, credit-reporting disputes, and related consumer-protection data [1,2,14,15]. The literature demonstrates the value of scalable complaint analytics, but methodological components are often assessed in isolation, and operational decision claims are not always matched by validation of the underlying classification or clustering modules.
Recent studies have extended BERTopic to consumer and service domains, including hospitality, e-commerce, and regulated health applications [16,17]. Generative approaches such as TopicGPT use large language models to propose or refine topic labels [18]. Such methods may improve label expressiveness, but they also introduce prompt sensitivity, provider dependence, and reproducibility concerns. BERTopic was retained here because the embedding, reduction, and clustering configuration can be reported explicitly; LLM-assisted labeling remains a relevant comparative extension rather than an unexamined omission.
2.2. Sentiment Analysis of Financial Text
Sentiment analysis of financial text is difficult because formal terminology, regulatory references, hedging, and procedural descriptions can coexist with emotional language. A model trained in financial news may capture topical vocabulary but still mismatch consumer-authored complaints. Conversely, a social-media model may better represent informal affect while mismatching the institutional context. Model selection therefore requires direct evaluation of the intended text type rather than an assumption that either domain or stylistic similarity is sufficient.
FinBERT is based on BERT and was further pre-trained in large collections of financial news and analyst reports to better capture domain-related sentiment signals. It has shown improved results on financial phrase-bank datasets compared to general-purpose models [19]. Unlike financial news, complaint text is typically written by non-experts and often contains emotional expressions. Moreover, formal language and personal storytelling are mixed resulting in a distinct stylistic category.
The Cardiff Roberta which is used here is a RoBERTa-base model pretrained on approximately 124 million tweets and fine-tuned for three-class sentiment analysis [20,21]. Its informal and affect-rich pretraining data provide a useful stylistic contrast to FinBERT. A comparison of these two checkpoints can test whether stylistic similarity may matter for this dataset, but it cannot establish a general rule that style is always more important than domain [22].
2.3. Complaint-Profile Clustering
Traditional segmentation of consumers complaints is based on demographics and product categories. More recent approaches move toward behavioral segmentation derived directly from complaint content. By using text embeddings, consumers are grouped based on similar experiences and expressed sentiment [23]. This type of semantic seg-mentation provides a more detailed understanding of consumer needs and dissatisfaction patterns compared to simple demographics grouping [24].
2.4. System Modeling Perspective
System dynamics examines how accumulations, flows, delays, and feedback relations generate behavior over time. In complaint management, Reinwald [25] developed a system-dynamics decision-support model linking complaint handling and repurchase behavior, while causal-loop applications have also been used to reason about customer satisfaction and public-service performance [26]. These studies go beyond placing modules in a circular diagram: they define system boundaries, variables, polarities, and feedback mechanisms [27]. The present work adopts this conceptual discipline but does not calibrate or simulate a quantitative system-dynamics model.
Our system contribution is therefore deliberately bounded. Figure 1 distinguishes the observed analytical workflow from a conceptual stock–flow and feedback formulation. Empirically, topic, sentiment, complaint-profile, and temporal outputs are combined through an uncertainty gate and human review. Conceptually, observed complaint records are treated as an imperfect measurement of issue inflow rather than as the full stock of unresolved consumer problems. This framing produces testable hypotheses for future calibration while avoiding the claim that the current pipeline is an implemented adaptive system.
Figure 1.
Proof-of-concept analytical workflow and conceptual systems framing. Panel (A) summarizes the empirical analytical workflow. Panel (B) provides a conceptual stock–flow and feedback interpretation discussed in Section 5.3; its dashed relations were not estimated, simulated, or operationally validated in the present study. B1 denotes the conceptual balancing feedback loop, through which organizational responses may contribute to reducing the stock of unresolved consumer issues.
3. Materials and Methods
The empirical analysis comprised the following four complementary modules: topic discovery, dual-model sentiment inference, exploratory clustering of complaint records, and descriptive temporal analysis. Outputs from these modules were retained with module-specific uncertainty information to support subsequent human interpretation. Figure 1A summarizes this empirical workflow, whereas the conceptual feedback representation in Figure 1B is discussed separately in Section 5.3 and was not used as an empirical or simulated component of the analysis.
Uncertainty was tracked separately for each analytical module rather than collapsed into a single confidence score. Residual BERTopic non-assignment indicates uncertainty in topic membership only and does not invalidate independently computed sentiment predictions, embedding-based complaint groupings, or aggregate temporal counts. When outputs are interpreted jointly, topic-assignment status (HDBSCAN-core, probability-reassigned, or residual-unassigned) is retained alongside the other analytical outputs so that uncertain or missing topic evidence remains explicit. The framework does not claim a calibrated cross-module confidence score.
3.1. Dataset
This study uses the CFPB Consumer Complaint Database, a public record of complaints about financial products and services. The fields used were date received, product, sub-product, issue, consumer complaint narrative, company, state, consumer-disputed status, and timely response status. The unit of analysis is a complaint record; the database does not provide a verified unique-consumer identifier for this purpose.
The database snapshot was downloaded from the CFPB Consumer Complaint Database on 20 April 2026. Records were first restricted to non-missing narratives longer than 50 characters and valid receipt dates. From the eligible records, a simple random sample of 100,000 rows—not a stratified sample—was drawn with pandas DataFrame.sample and random_state = 42. Text cleaning and the exclusion of narratives with fewer than 10 cleaned tokens removed 566 records, leaving 99,434. The final sample spans from March 2015 to March 2026 and encompasses 21 product categories. Because sampling was unstratified, no exact year-by-product preservation constraint was imposed, and the resulting temporal and product composition should be interpreted as descriptive of the sampled snapshot. Table 1 summarizes the main characteristics of the final analytical sample and the sampling procedure.
Table 1.
Dataset descriptive statistics.
3.2. Preprocessing
The preprocessing function removed sequences of two or more X characters used for CFPB redaction, numeric strings of four or more digits, and punctuation; it then collapsed whitespace and converted text to lowercase. Narratives with fewer than 10 cleaned tokens were excluded. Exact-text duplicates were not removed in the original record-level analysis, so repeated or attachment-only narratives may receive additional weight; this is addressed in the limitations.
3.3. Topic Modeling with BERTopic
Topic modeling used BERTopic [10] with document embeddings from the sentence-transformers/all-mpnet-base-v2 checkpoint, which maps text to 768-dimensional vectors [28,29]. Embeddings were generated in batches of 512. The vectorizer used English stop words, unigrams and bigrams, min_df = 15, max_df = 0.90, and max_features = 12,000.
UMAP was configured with n_neighbors = 30, n_components = 10, min_dist = 0.0, cosine distance, and random_state = 42. HDBSCAN used min_cluster_size = 80, min_samples = 10, Euclidean distance, excess-of-mass cluster selection, and prediction_data = True. Topic terms were represented with class-based TF–IDF, KeyBERT-inspired extraction, and maximal marginal relevance (diversity = 0.2). After the initial fit, BERTopic.reduce_topics(docs, nr_topics = 90) merged similar clusters. Because the outlier label −1 is separate from non-outlier topics, the requested target produced 89 non-outlier topics plus the outlier class. A second assignment stage was then applied to records labeled −1 using BERTopic.reduce_outliers with strategy = “probabilities” and threshold = 0.05. The fitted embeddings, topic representations, and original non-outlier assignments were not altered.
The probability-based stage assigned 23,508 previously unassigned records, yielding 80,009 topic-assigned records (80.46% coverage) and leaving 19,425 records unassigned (19.54%). The 89 topic representations were retained, none of the original non-outlier assignments changed, and topic-volume ranks remained similar to the pre-reassignment solution (Spearman’s ρ = 0.994; total-variation distance = 0.1182). Because BERTopic membership probabilities are not calibrated measures of correctness, record-level outputs distinguish HDBSCAN-core, probability-reassigned, and residual-unassigned cases. The latter two categories are treated as lower-certainty evidence for human review rather than equivalent to core cluster membership.
3.4. Dual Sentiment Analysis
Sentiment classification is performed using the following two pre-trained transformer models applied independently to each complaint narrative:
FinBERT (ProsusAI/finbert): a BERT-based three-class checkpoint fine-tuned for financial sentiment [6]. It was included as the finance-domain comparator.
Cardiff RoBERTa (cardiffnlp/twitter-roberta-base-sentiment-latest): a RoBERTa-base—not RoBERTa-large—checkpoint pretrained on approximately 124 million tweets and fine-tuned on TweetEval sentiment data [20,21]. It was included as a stylistically contrasting comparator. The two checkpoints were selected deliberately to provide a focused contrast between a finance-domain model and a model trained on informal, affect-rich text. The comparison was not intended as an exhaustive benchmark of available sentiment models, and conclusions regarding domain versus stylistic alignment are therefore restricted to these two checkpoints.
Both checkpoints were applied independently with truncation at 512 tokens, batch size 64, and normalized labels (negative, neutral, and positive). The analytical workflow was executed in a retained Conda environment using Python 3.10.20, BERTopic 0.16.0, umap-learn 0.5.12, hdbscan 0.8.42, scikit-learn 1.7.2, sentence-transformers 2.7.0, transformers 4.40.0, PyTorch 2.5.1 with CUDA 12.1 support, pandas 2.3.3, and NumPy 2.2.6. Computation was performed on an NVIDIA A40 GPU (NVIDIA Corporation, Santa Clara, CA, USA). During revision, the retained analysis environment was exported to document the complete dependency configuration; the resulting environment specification is provided as Supplementary File S1.
3.5. Human Annotation and Model Validation
A record-level benchmark of 600 complaints was created by stratifying on FinBERT’s predicted class: 200 records each were sampled from the predicted negative, neutral, and positive groups with random_state = 42. This design increased coverage of the model’s rare predicted classes and was not intended to estimate population sentiment prevalence. The benchmark was therefore designed as a diagnostic comparison set rather than as a prevalence-representative test sample. Because both checkpoints were evaluated on exactly the same 600 records, the paired comparison assesses their relative agreement with human reference within this deliberately balanced sample. However, the resulting accuracy, macro-F1, and κ values are sample-specific and should not be generalized as population-performance estimates for the full CFPB corpus. Annotator 1 assigned negative, neutral, or positive labels using written guidance as follows: negative denoted dissatisfaction, harm, accusation, or adverse experience; neutral denoted primarily factual or procedural text without clear valence; and positive denoted praise, relief, successful resolution, or clearly favorable experience. Annotator 2 independently and blindly coded all 600 records without access to Annotator 1 or model outputs. One free-text label note was normalized to negative before agreement calculation. The annotators achieved 87.5% raw agreement and Cohen’s κ = 0.758. The reported model metrics use Annotator 1 as the prespecified reference; Annotator 2 was used to quantify reference-label reliability, and disagreements were not adjudicated for the reported model comparison. Annotator 1’s class composition was 407 negative, 126 neutral, and 67 positive records.
3.6. Exploratory Clustering of Complaint Records
Complaint records were explored by reducing the 768-dimensional embeddings to 50 principal components, L2-normalizing and scaling them by 2.0, and concatenating two standardized ordinal sentiment features scaled by 0.3. K-means used k-means++ initialization, n_init = 10, and random_state = 42. K values from 2 to 8 were examined with an elbow plot and a Silhouette calculation on a random sample of 5000 observations in the 50-dimensional PCA representation using labels obtained from the composite feature space. The Silhouette coefficient favored K = 2 and was low (maximum = 0.106). K = 4 was retained only as a richer exploratory descriptive partition suggested by the elbow inspection. Because the Silhouette values were low across all examined K values, the four-group solution is not interpreted as evidence of four naturally occurring or statistically well-separated complaint types. The resulting groups are used only to summarize variation in semantic content and model-assigned sentiment and should not be interpreted as validated segments. A retained-output feature-weight sensitivity analysis repeated K = 4 for embedding/sentiment scale ratios of 1.0, 2.0, 3.3, 4.0, 6.7 (reported), 10.0, 16.7, and 20.0; adjusted Rand index relative to the reported partition and the sampled Silhouette coefficient were compared.
3.7. Temporal Analysis
Temporal analysis was descriptive and used three resolutions. Complaint volume was aggregated monthly. Sentiment proportions were summarized annually. For topic prevalence, the revised topic assignments were paired with receipt dates rounded to calendar quarters and passed to BERTopic.topics_over_time with nr_bins = 20, yielding a coarser 20-bin summary across the study period. Residual topic −1 records were preserved but excluded from named-topic interpretation. September 2017 (Equifax breach) and March 2020 (COVID-19 onset) were added only as contextual reference markers on the monthly volume plot. The earlier unsupported “CFPB Rule Update 2023” marker was removed. The first and last calendar years are partial, and no event-study, structural-break, or causal identification procedure was performed.
4. Results
4.1. Topic Modeling Results
After topic reduction and probability-based outlier reassignment, the analysis retained 89 non-outlier topics and assigned 80,009 of 99,434 records (80.46%). The remaining 19,425 records (19.54%) retained topic −1 and form an explicit uncertainty queue. Prior studies have likewise examined the generalizability and application of BERTopic across heterogeneous, multi-domain, and domain-specific short-text settings [30,31]. The reassignment stage added 23,508 records without changing any original non-outlier assignment or refitting the topic representations. Topic-volume ordering remained highly similar to the pre-reassignment solution (Spearman’s ρ = 0.994), although the topic-share total-variation distance of 0.1182 indicates a meaningful redistribution. Reassignments were concentrated in broad existing themes: five recipient topics absorbed 83.23% of reassigned records. Accordingly, reassigned cases are flagged separately and should not be treated as having the same evidential status as HDBSCAN-core members. Table 2 and Figure 2 and Figure 3 use the revised assignments consistently. All 99,434 records continue to contribute to sentiment inference, exploratory complaint-record grouping, and aggregate temporal counts because these modules do not require a BERTopic assignment. For the 19,425 residual-unassigned records, however, topic membership remains undefined and is retained as a separate uncertainty and provenance flag in any combined interpretation.
Table 2.
Ten highest-volume complaint topics, ranked consistently with Figure 2.
Figure 2.
Top 20 BERTopic topics by complaint-record volume after probability-based outlier reassignment (threshold = 0.05). Residual unassigned records (topic −1; 19.54%) are excluded from the bars but retained in the other analytical modules.
Figure 3.
FinBERT-assigned sentiment distribution within the 20 highest-volume topics after probability-based outlier reassignment. Percentages are normalized within topics.
The highest-volume revised topic was Topic 1 (alleged debt/late payments/reporting act; n = 12,521), followed by Topic 0 (credit reporting/consumer reporting/reporting act; n = 8023). Other high-volume themes concerned fraud reporting, account closure, mortgage and lien disputes, bureau-specific reporting, debt collection, and fraudulent inquiries. These labels summarize recurring narrative patterns; they should not be interpreted as verified root causes or as complete coverage of the complaint stream.
Figure 2 shows a long-tailed distribution of the 80,009 topic-assigned records. Several high-volume topics contain regulatory or legal language that FinBERT frequently labels as neutral. This descriptive association motivated the model comparison in Section 4.2; it does not by itself establish systematic model bias without reference labels.
The revised topic assignments show substantial variation in FinBERT-labeled sentiment across the 20 highest-volume topics. Topics involving bureau disputes and alleged violations include relatively high negative shares, whereas several debt- and reporting-related topics remain predominantly neutral under FinBERT. Because these labels are model-dependent and 19.54% of records remain topic-unassigned, the concentrations are treated as screening signals rather than validated severity measures.
4.2. Sentiment Analysis Results
Table 3 presents the overall sentiment distribution for both models across the full dataset. The two models exhibit substantially different sentiment distributions, reflecting their distinct training domains.
Table 3.
Sentiment distribution comparison FinBERT vs. Cardiff RoBERTa.
FinBERT labeled 61.9% of records neutral, whereas Cardiff RoBERTa labeled 83.0% negative. These contrasting distributions are consistent with differences in pretraining and fine-tuning data, but they do not establish the reason for the discrepancy. Human-reference evaluation is therefore required before drawing conclusions about comparative alignment.
Figure 4 shows substantial inter-model disagreement. Among complaints labeled neutral by FinBERT, 77% were labeled negative by Cardiff RoBERTa. This pattern indicates sensitivity to checkpoint choice; it should not be described as a confirmed error pattern until compared with human labels. Both checkpoints produced very few positive predictions in the full complaint sample.
Figure 4.
FinBERT versus Cardiff RoBERTa sentiment comparison. The right-hand heatmap is row-normalized by FinBERT label; each row sums to 100%. Overall exact label agreement is 49.9%.
4.3. Human Annotation Validation
Table 4 reports performance on the 600-record FinBERT-prediction-stratified benchmark using Annotator 1 as the reference. The reference-label composition was 407 negative, 126 neutral, and 67 positive. Annotator 2 was used for reliability assessment (raw agreement = 87.5%; κ = 0.758), not as an adjudicated consensus label.
Table 4.
Model performance against Annotator 1 on the prediction-stratified benchmark (n = 600).
On this benchmark, Cardiff RoBERTa showed higher alignment with Annotator 1 than FinBERT across κ, accuracy, and macro-F1. FinBERT predicted neutral for 182 of 407 reference-negative records. Cardiff RoBERTa correctly identified 329 of the 407 reference-negative records, although its performance for the smaller neutral class remained weak. These estimates characterize the deliberately prediction-stratified benchmark and should not be interpreted as population-weighted performance estimates for the full CFPB database. Their role is comparative: both checkpoints were evaluated against the same human reference labels on the same records, allowing differences in error structure and paired correctness to be examined within this diagnostic sample.
Figure 5 illustrates different error structures. FinBERT frequently mapped reference-negative records to neutral, while Cardiff RoBERTa recovered more negative records but did not resolve the difficulty of neutral classification. For this dataset, the comparison suggests that stylistic similarity may contribute to better alignment; a two-model comparison does not establish that stylistic similarity generally dominates domain similarity.
Figure 5.
Confusion matrices using Annotator 1 as the reference label. FinBERT: κ = 0.195; Cardiff RoBERTa: κ = 0.395. The benchmark contains 407 negative, 126 neutral, and 67 positive reference labels. Darker shades indicate higher cell counts, while lighter shades indicate lower counts.
A McNemar test on paired correctness outcomes found that Cardiff RoBERTa was correct on 159 records that FinBERT missed, compared with 33 in the reverse direction (χ2 = 81.38, p < 0.001). Statistical significance does not imply operational adequacy: the stronger checkpoint still reached only fair agreement (κ = 0.395). Accordingly, sentiment predictions are retained only as uncertain evidence for human review, not as autonomous routing or prioritization decisions.
4.4. Exploratory Complaint-Record Groupings
K-means produced four exploratory complaint-record groupings in the composite semantic and sentiment feature space. Given the weak internal separation, these groupings are treated as descriptive partitions of the analytical feature space rather than as validated clusters, behavioral segments, or naturally occurring complaint types. Because records cannot be linked reliably to unique individuals, the groupings also should not be interpreted as consumer segments. Table 5 summarizes their descriptive composition.
Table 5.
Descriptive Characteristics of the Four Exploratory Complaint-Record Groups.
Dominant issues per complaint-record grouping:
- Group 0 (Fraud and Inquiry): CFPB/FTC fraud reports and fraudulent or unauthorized-account narratives;
- Group 1 (Bureau Violation): alleged-debt, late-payment, and bureau-inaccuracy narratives;
- Group 2 (Debt and Mortgage): fraudulent-charge, account-closure, mortgage, lien, and fraud-department narratives;
- Group 3 (Regulatory Action): credit- and consumer-reporting-law language, CFPB reports, and reporting-agency narratives.
The Debt and Mortgage grouping is the largest (n = 36,289), followed by Bureau Violation (n = 28,648), Fraud and Inquiry (n = 26,669), and Regulatory Action (n = 7828). The groupings differ descriptively in dominant terms and model-assigned sentiment. These differences are not evidence that the groups are well-separated behavioral types or that group membership predicts organizational outcomes.
Figure 6 provides a two-dimensional view of the four-group pattern and the relative group sizes. The visible overlap is consistent with the low Silhouette coefficient and supports an exploratory, not confirmatory, interpretation.
Figure 6.
Exploratory complaint-record groupings: two-dimensional PCA projection of a 5000-record sample and group-size distribution. The substantial overlap is consistent with the low Silhouette coefficients; the visualization is descriptive and does not demonstrate four well-separated or validated clusters in the full feature space.
Clustering Validation and Feature-Weight Sensitivity
K values from 2 to 8 were examined. The highest reported Silhouette coefficient was 0.106 at K = 2, indicating weak geometric separation even for the favored coarse partition. At K = 4, the Silhouette coefficient was 0.0847, the Davies–Bouldin index was 2.7446, and the Calinski–Harabasz index was 751.0. K = 4 had the lowest Davies–Bouldin value among K = 2–8, but it was not favored by either Silhouette or Calinski–Harabasz. It was retained because the elbow inspection suggested a four-group descriptive summary and because it yielded interpretable descriptive groupings. This mixed internal evidence supports an exploratory partition, not a uniquely optimal or stable segmentation solution. Accordingly, no inferential or operational conclusions are drawn from membership in these four groups; their role in the present study is limited to descriptive summarization of heterogeneous complaint narratives.
The reported composite feature weights (embedding scale 2.0; sentiment scale 0.3; and ratio 6.7) were examined against alternative ratios. Nearby embedding-dominant ratios of 10.0, 16.7, and 20.0 produced adjusted Rand indices of 0.982, 0.965, and 0.971 relative to the reported solution, with Silhouette coefficients of 0.087, 0.088, and 0.088. Greater sentiment influence changed the partition substantially as follows: ratios of 4.0, 2.0, and 1.0 yielded adjusted Rand indices of 0.524, 0.116, and −0.002 and Silhouette coefficients of 0.081, 0.152, and 0.434, respectively. The high equal-weight Silhouette reflects a fundamentally different, sentiment-dominated partition rather than validation of the reported groupings. The reported solution is therefore locally stable under nearby embedding-dominant weights but globally sensitive to feature weighting. This parameter sensitivity is not a resampling or external-validity test. No downstream outcome prediction or operational A/B test was available; group-based actions remain unvalidated.
4.5. Descriptive Temporal Analysis
Monthly complaint volume and coarser topic-prevalence summaries vary across the 2015–2026 sample. Increases around September 2017 and March 2020 coincide with the Equifax breach and the onset of COVID-19, respectively. These event markers provide context only. The analysis does not control for reporting propensity, product composition, intake-policy changes, delayed publication, or other confounders, and therefore does not estimate event effects.
Figure 7 shows increasing sampled complaint volume after 2023, a peak in early 2025, and a decline near the partial 2026 endpoint. Because the 2025 change is concentrated in credit-reporting records and may also reflect changes in product composition, reporting behavior, intake practices, or publication processes, the series should be interpreted descriptively rather than as evidence of a specific external effect. No difference-in-difference, regression-discontinuity, interrupted-time-series, or formal structural-break test was conducted. The plot is descriptive and cannot support causal or forecasting claims.
Figure 7.
Monthly complaint-record volume with two contextual event markers. The markers do not establish causal effects. March–December 2015 and January–March 2026 are partial-year periods.
Annual model-assigned sentiment proportions are comparatively stable within each checkpoint, but this does not establish temporal model stability. No period-specific relabeling, recalibration, or drift benchmark against fresh human annotations was performed. A deployment-oriented extension should monitor changes in input composition and model error by time period and trigger retraining or review when predefined thresholds are exceeded.
5. Discussion
5.1. Implications for Business Management
The framework converts complaint records into topic, sentiment, clustering, and temporal summaries that may support human analysis. Potential applications include identifying recurring themes, organizing review queues, monitoring changes in complaint composition, and allocating analyst attention. In practical decision-support terms, descriptive temporal monitoring may be used to flag unusual changes in complaint volume, topic prevalence, or sentiment composition for analyst review and to identify areas that may warrant closer investigation or resource attention. Such flags are screening signals only and should not be interpreted as evidence that a particular external event caused the observed change. These are proposed uses, not validated decision outcomes. Topic non-assignment, model disagreement, and low confidence should increase—not bypass—human review.
Bureau-specific and credit-reporting clusters show that complaints can preserve institution- and issue-specific language. The co-occurrence of changes in relevant complaint topics with the 2017 Equifax breach illustrates how the framework can contextualize known events retrospectively. It does not demonstrate prospective event detection or earlier regulatory response. Operational early-warning performance would require predefined alarms, prospective evaluation, and measured false-positive and false-negative costs.
5.2. Model Selection for Complaint Sentiment Analysis
On the present benchmark, Cardiff RoBERTa aligned better with Annotator 1 than FinBERT. This result suggests that stylistic similarity between training data and complaint narratives may be relevant for this dataset. It does not establish that stylistic similarity generally outweighs domain similarity, because only two checkpoints were compared and the benchmark was stratified by FinBERT prediction rather than sampled for population prevalence.
The modest agreement of both checkpoints with the human reference—slight for FinBERT (κ = 0.195) and fair for Cardiff RoBERTa (κ = 0.395)—rules out standalone automated classification. A defensible use is an uncertainty-aware human-in-the-loop process: cases with model disagreement, topic non-assignment, or low confidence are routed for review; automated labels remain advisory; and performance is periodically re-estimated on newly annotated records. Complaint-specific fine-tuning may improve alignment, but that proposition requires a larger, de-duplicated, adjudicated benchmark and prospective evaluation. Accordingly, the comparison supports checkpoint selection for subsequent exploratory analysis within this study, but it does not establish the expected accuracy of either model under the natural class distribution of future CFPB complaints.
5.3. Conceptual System Interpretation
The temporal results do not demonstrate a functioning dynamic feedback system. Instead, Figure 1B offers a conceptual systems hypothesis: unresolved issues may generate complaint inflow; observed complaints may trigger analytical alerts and organizational responses; and effective resolution may reduce the unresolved stock after a delay. At the same time, observed complaint volume depends on reporting propensity and intake coverage. This distinction helps explain why a volume spike cannot be attributed automatically to a change in underlying consumer harm.
To formalize this conceptual distinction, let Ut denote the latent stock of unresolved consumer issues at time t, It the inflow of new issues, and Rt the resolution outflow, such that Ut+1 = Ut + It − Rt. Observed complaint records Ct are treated as an imperfect measurement of issue inflow, Ct = qtIt, where qt represents reporting propensity and intake coverage. These relations are included only to clarify the hypothesized system boundary and feedback logic; their coefficients, delays, and counterfactual behavior were not estimated or simulated in the present study.
The system formulation adds three bounded contributions. First, it separates a latent stock of unresolved issues from the observed complaint stream. Second, it identifies an explicit balancing loop and the direction of its hypothesized relations. Third, it places uncertainty and human review between analytics and action.
5.4. Contextual Considerations of the CFPB Data Landscape
The CFPB data landscape changes over time. The 2025 increase is concentrated in credit-reporting narratives, so aggregate volume may reflect product-composition change rather than a broad increase across financial services. Changes in product composition, reporting behavior, intake practices, and publication processes may also alter qt in the conceptual formulation. The absence of a formal structural-break test, rolling human-validation sample, or retraining experiment limits temporal generalizability. Any operational extension should monitor year-by-product distributions and revalidate model performance after material drift.
5.5. Limitations
Several limitations materially constrain interpretation. First, CFPB complainants are self-selected and do not represent all consumers. Second, the 100,000-record sample was simple random rather than stratified, and the source distribution was not preserved by an explicit year-by-product allocation rule. Third, the 600-record annotation benchmark was stratified by FinBERT prediction, used Annotator 1 as the reference without adjudication, and contains repeated or boilerplate narratives because the record-level workflow did not de-duplicate exact text; its metrics are not population-prevalence estimates. Fourth, probability-based outlier reassignment increased final topic coverage to 80.46%, but 19.54% remained unassigned. The membership scores are not calibrated probabilities of correctness, the reassigned cases were not independently human-validated, and 83.23% of reassignments entered five broad topics. HDBSCAN-core and probability-reassigned records are therefore retained as distinct provenance categories. Fifth, the four exploratory complaint-record groupings have weak internal separation (maximum Silhouette = 0.106), show only local stability under nearby embedding-dominant feature weights, and have no resampling, external, or operational validation. Sixth, temporal analyses are descriptive, the endpoint years are partial, and no causal, structural-break, retraining, or rolling-stability test was conducted. Seventh, although the retained Conda environment and complete dependency specification are now provided, exact computational reproducibility may still depend on hardware, low-level libraries, and stochastic behavior in dimensionality-reduction and clustering procedures.
6. Conclusions
This study presents a proof-of-concept analytical framework for 99,434 CFPB complaint records spanning from March 2015 to March 2026. It integrates BERTopic, two sentiment checkpoints, exploratory complaint-record groupings, descriptive temporal summaries, and a conceptual stock–flow and feedback representation. The revised probability-based topic-assignment stage retained 89 non-outlier topics and achieved 80.46% coverage, while 19.54% of records remained unassigned and were preserved as an uncertainty queue. Sentiment outputs remained substantially sensitive to checkpoint choice.
Cardiff RoBERTa aligned better with Annotator 1 than FinBERT on the 600-record prediction-stratified benchmark, which suggests that stylistic similarity may matter for this complaint dataset. However, the stronger model achieved only fair agreement, and the four-group exploratory partition showed weak geometric separation and is therefore interpreted only as a descriptive summary rather than as validated clustering structure. Neither component supports autonomous decision-making or validated group-based intervention.
The framework’s value lies in organizing heterogeneous evidence while preserving uncertainty and human oversight. Topic non-assignment, model disagreement, and temporal non-stationarity are retained as distinct review flags rather than collapsed into a single confidence measure or hidden as failure modes. Future work should use a pinned environment, a larger de-duplicated and adjudicated annotation set, topic and cluster stability analyses, prospective drift monitoring, calibrated system-dynamics models, and evaluation within an actual organizational workflow.
Supplementary Materials
The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/systems14101266/s1, Supplementary File S1 (Software Environment Specification).
Author Contributions
Conceptualization, M.E.C., A.B. and A.C.; methodology, M.E.C., A.B. and A.C.; software, M.E.C., M.K., K.V. and G.M.; validation, M.E.C., A.B. and N.T.; formal analysis, M.E.C., A.B. and N.T.; investigation, M.E.C., M.K., K.V. and G.M.; resources, M.E.C., A.B. and G.M.; data curation, M.E.C. and K.V.; writing—original draft preparation, M.E.C., N.T. and M.K.; writing—review and editing, M.E.C. and A.B.; visualization, A.B. and G.M.; supervision, A.B. and A.C.; project administration, A.B. and A.C.; funding acquisition, A.B. and A.C. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Data Availability Statement
The source data are publicly available from the CFPB Consumer Complaint Database at https://files.consumerfinance.gov/ccdb/complaints.csv.zip (accessed on 20 April 2026). The analyzed snapshot was downloaded on 20 April 2026. Eligible records had non-missing narratives longer than 50 characters and valid dates; a simple random sample of 100,000 records was drawn with random_state = 42, and 566 records with fewer than 10 cleaned tokens were removed. The final analytical sample contained 99,434 complaint records. Model checkpoints and analytical parameters are reported in Section 3. The retained Conda environment was exported during revision, and the complete dependency specification is provided as Supplementary File S1.
Acknowledgments
During the preparation of this work, the authors used ChatGPT (GPT-5.5, OpenAI) solely to improve the English language, readability, and clarity of the manuscript. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| NLP | Natural Language Processing |
| LLMs | Large Language Models |
| LDA | Latent Dirichlet Allocation |
| UMAP | Uniform Manifold Approximation and Projection |
| CFPB | Consumer Financial Protection Bureau |
| HDBSCAN | Hierarchical Density-Based Spatial Clustering of Applications with Noise |
| NMF | Non-negative Matrix Factorization |
References
- Siering, M. Explainability and Fairness of RegTech for Regulatory Enforcement: Automated Monitoring of Consumer Complaints. Decis. Support Syst. 2022, 158, 113782. [Google Scholar] [CrossRef] [Scilit]
- Vasudeva Raju, S.; Kumar Bolla, B.; Nayak, D.K.; Kh, J. Topic Modelling on Consumer Financial Protection Bureau Data: An Approach Using BERT Based Embeddings. In Proceedings of the 2022 IEEE 7th International Conference for Convergence in Technology (I2CT 2022), Pune, India, 7–9 April 2022; IEEE: Piscataway, NJ, USA, 2022. [Google Scholar] [CrossRef] [Scilit]
- Consumer Financial Protection Bureau Consumer Complaint Database. Available online: https://www.consumerfinance.gov/data-research/consumer-complaints/ (accessed on 20 April 2026).
- Bhattacharya, A. Consumer, Bank, and Stock Market Reaction to CFPB’s Complaint Data Disclosure. J. Financ. Serv. Mark. 2023, 28, 128–145. [Google Scholar] [CrossRef] [Scilit]
- Galli, C.; Cusano, C.; Meleti, M.; Donos, N.; Calciolari, E. Topic Modeling for Faster Literature Screening Using Transformer-Based Embeddings. Metrics 2024, 1, 2. [Google Scholar] [CrossRef] [Scilit]
- Araci, D. FinBERT: Financial Sentiment Analysis with Pre-Trained Language Models. arXiv 2019, arXiv:1908.10063. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; Stoyanov, V.; et al. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv 2019, arXiv:1907.11692. [Google Scholar] [CrossRef] [Scilit]
- Singh, J.; Wilkes, R.E. When Consumers Complain: A Path Analysis of the Key Antecedents of Consumer Complaint Response Estimates. J. Acad. Mark. Sci. 1996, 24, 350–365. [Google Scholar] [CrossRef] [Scilit]
- Jelodar, H.; Wang, Y.; Yuan, C.; Feng, X.; Jiang, X.; Li, Y.; Zhao, L. Latent Dirichlet Allocation (LDA) and Topic Modeling: Models, Applications, a Survey. Multimed. Tools Appl. 2019, 78, 15169–15211. [Google Scholar] [CrossRef] [Scilit]
- Grootendorst, M. BERTopic: Neural Topic Modeling with a Class-Based TF-IDF Procedure. arXiv 2022, arXiv:2203.05794. [Google Scholar] [CrossRef] [Scilit]
- McInnes, L.; Healy, J.; Melville, J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv 2018, arXiv:1802.03426. [Google Scholar] [CrossRef] [Scilit]
- Malzer, C.; Baum, M. A Hybrid Approach to Hierarchical Density-Based Cluster Selection. In Proceedings of the 2020 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems, Karlsruhe, Germany, 14–16 September 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 223–228. [Google Scholar] [CrossRef] [Scilit]
- Turan, S.C.; Yildiz, K.; Büyüktanir, B. Comparison of LDA, NMF and BERTopic Topic Modeling Techniques on Amazon Product Review Dataset: A Case Study. In Computing, Internet of Things and Data Analytics: Selected Papers from the International Conference on Computing, IoT and Data Analytics (ICCIDA); García Márquez, F.P., Jamil, A., Ramirez, I.S., Eken, S., Hameed, A.A., Eds.; Studies in Computational Intelligence; Springer: Cham, Switzerland, 2024; Volume 1145, pp. 23–31. [Google Scholar] [CrossRef] [Scilit]
- Humphreys, A.; Wang, R.J.H. Automated Text Analysis for Consumer Research. J. Consum. Res. 2018, 44, 1274–1306. [Google Scholar] [CrossRef] [Scilit]
- Vaishnav, D.; Neethinayagam, M.; Khaire, A.; Woo, J. Predictive Analysis of CFPB Consumer Complaints Using Machine Learning. arXiv 2024, arXiv:2407.06399. [Google Scholar] [CrossRef] [Scilit]
- Mishra, M. A Holistic Review of Customer Experience Research: Topic Modelling Using BERTopic. Mark. Intell. Plan. 2025, 43, 802–820. [Google Scholar] [CrossRef] [Scilit]
- Uncovska, M.; Freitag, B.; Meister, S.; Fehring, L. Rating Analysis and BERTopic Modeling of Consumer versus Regulated MHealth App Reviews in Germany. npj Digit. Med. 2023, 6, 115. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pham, C.M.; Hoyle, A.; Sun, S.; Resnik, P.; Iyyer, M. TopicGPT: A Prompt-Based Topic Modeling Framework. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2024, Mexico City, Mexico, 16–21 June 2024; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; Volume 1, pp. 2956–2984. [Google Scholar] [CrossRef] [Scilit]
- Huang, A.H.; Wang, H.; Yang, Y. FinBERT: A Large Language Model for Extracting Information from Financial Text. Contemp. Account. Res. 2023, 40, 806–841. [Google Scholar] [CrossRef] [Scilit]
- Camacho-Collados, J.; Rezaee, K.; Riahi, T.; Ushio, A.; Loureiro, D.; Antypas, D.; Boisson, J.; Espinosa-Anke, L.; Liu, F.; Martínez-Cámara, E.; et al. TweetNLP: Cutting-Edge Natural Language Processing for Social Media. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations; Che, W., Shutova, E., Eds.; Association for Computational Linguistics: Abu Dhabi, United Arab Emirates, 2022; pp. 38–49. [Google Scholar] [CrossRef] [Scilit]
- Loureiro, D.; Barbieri, F.; Neves, L.; Espinosa Anke, L.; Camacho-Collados, J. TimeLMs: Diachronic Language Models from Twitter. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations; Basile, V., Kozareva, Z., Stajner, S., Eds.; Association for Computational Linguistics: Dublin, Ireland, 2022; pp. 251–260. [Google Scholar] [CrossRef] [Scilit]
- Karanikola, A.; Davrazos, G.; Liapis, C.M.; Kotsiantis, S. Financial Sentiment Analysis: Classic Methods vs. Deep Learning Models. Intell. Decis. Technol. 2023, 17, 893–915. [Google Scholar] [CrossRef] [Scilit]
- Saha, R. Influence of Various Text Embeddings on Clustering Performance in NLP. arXiv 2023, arXiv:2305.03144. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Liu, Y.; Yu, M. Consumer Segmentation with Large Language Models. J. Retail. Consum. Serv. 2025, 82, 104078. [Google Scholar] [CrossRef] [Scilit]
- Reinwald, D. Complaint Management and Repurchase Behavior: A Decision Support Approach Using System Dynamics. In Proceedings of the Fifteenth Americas Conference on Information Systems (AMCIS), San Francisco, CA, USA, 6–9 August 2009; pp. 1–14. [Google Scholar]
- Aryani, R.P. Causal Loop Diagram of Customer Satisfaction with Public Services: A Case Study of Sanggau Regency, Indonesia. Eruditio 2022, 2, 56–72. [Google Scholar] [CrossRef] [Scilit]
- Sterman, J.D. Business Dynamics: Systems Thinking and Modeling for a Complex World; Irwin/McGraw-Hill: Boston, MA, USA, 2000. [Google Scholar]
- Song, K.; Tan, X.; Qin, T.; Lu, J.; Liu, T.-Y. MPNet: Masked and Permuted Pre-Training for Language Understanding. Adv. Neural Inf. Process. Syst. 2020, 33, 16857–16867. [Google Scholar]
- Sentence Transformers. all-mpnet-base-v2 Model Card. Available online: https://huggingface.co/sentence-transformers/all-mpnet-base-v2 (accessed on 19 August 2026).
- De Groot, M.; Aliannejadi, M.; Haas, M.R. Experiments on Generalizability of BERTopic on Multi-Domain Short Text. arXiv 2022, arXiv:2212.08459. [Google Scholar] [CrossRef] [Scilit]
- Medvecki, D.; Bašaragin, B.; Ljajić, A.; Milošević, N. Multilingual Transformer and BERTopic for Short Text Topic Modeling: The Case of Serbian. In Disruptive Information Technologies for a Smart Society; Trajanović, M., Filipović, N., Zdravković, M., Eds.; Lecture Notes in Networks and Systems; Springer: Cham, Switzerland, 2024; Volume 872, pp. 161–173. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.






