Next Article in Journal
Research on a Differential Game Considering Endurance and Marketing Effort for eVTOL Under a Subsidy Policy
Next Article in Special Issue
Forecasting Systemic Reconfiguration in Concentrated Global Supply Networks for Economic Resilience: A Systems-Theoretic Hypergraph-Structured Temporal Decision-Support Framework
Previous Article in Journal
Conflict-Driven Action Boundary Generation for Emergency Group Decision-Making: A Constructed Association-Proxy Approach
Previous Article in Special Issue
Students’ Perceptions of the Use of Artificial Intelligence Tools in Educational Activities
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Evaluating Software Quality in Cryptocurrency Wallet Applications: A User-Centered Approach Using ISO/IEC 25010:2011 Quality Model

1
Division of Artificial Intelligence and Data Science, Korea Cyber University, Seoul 03051, Republic of Korea
2
School of Interdisciplinary Studies, Dongguk University-Seoul, Seoul 04620, Republic of Korea
*
Author to whom correspondence should be addressed.
Systems 2026, 14(9), 1078; https://doi.org/10.3390/systems14091078
Submission received: 31 July 2026 / Revised: 26 August 2026 / Accepted: 31 August 2026 / Published: 2 September 2026

Highlights

Please indicate how your work links to systems science via your contributions to systems practice, theory, and/or methodology.
This study employs an integrated methodology that combines BERTopic, ISO/IEC 25010:2011, and dissatisfaction-weighted prioritization to transform unstructured user feedback into structured software quality priorities.
The proposed framework supports systems practice by linking user experience data with standardized software quality characteristics and explicitly accounting for topics associated with multiple quality characteristics through fractional allocation.
What are the main findings and/or the implications of the main findings?
Security, Performance Efficiency, and Reliability were identified as the highest-priority quality characteristics for cryptocurrency wallet applications under fractional allocation.
Sensitivity analysis showed that the topic-allocation rule can affect the resulting priority rankings, highlighting the importance of addressing duplicated contributions when aggregating multi-mapped user concerns.

Abstract

This study proposes a systematic methodology integrating Bidirectional Encoder Representations from Transformers (BERT)-based topic modeling (BERTopic) with the ISO/IEC 25010:2011 product quality model to evaluate and prioritize software quality characteristics in cryptocurrency wallet applications using user review data. Using 68,790 preprocessed reviews, BERTopic identified 15 major topics representing user-reported quality concerns. The proposed approach combines topic frequency with user dissatisfaction to provide clear and prioritized insights for software quality improvement. Specifically, topic frequency is logarithmically transformed to reduce the disproportionate influence of highly frequent topics, while dissatisfaction is derived from the average star rating associated with each topic. These measures are combined to calculate topic-level weighted values, which are subsequently mapped to the corresponding ISO/IEC 25010:2011 quality characteristics. Because a single topic may be associated with multiple quality characteristics, fractional allocation is applied to distribute its weighted value across the mapped characteristics and avoid duplicated contributions. The results identify Security as the highest-priority quality characteristic, followed by Performance Efficiency and Reliability. A sensitivity analysis comparing fractional and full allocation further shows that the allocation rule can materially affect the resulting priority structure. These findings establish clear priorities for software quality improvement and highlight the importance of accounting for multi-label topic mappings when aggregating user feedback. Overall, the proposed framework provides a transparent and user-centered approach for transforming large-scale review data into structured software quality priorities.

1. Introduction

With the rapid growth of the cryptocurrency market, users increasingly rely on cryptocurrency wallet applications for digital asset management. These wallets serve as a primary interface between users and blockchain networks and play an important role in supporting secure asset management and transactions [1]. As cryptocurrency adoption has expanded, wallet applications have incorporated increasingly diverse functions, making software quality an important consideration for user trust and adoption [2]. Nevertheless, cryptocurrency wallets continue to face challenges related to privacy protection, security mechanisms, and complex user interfaces that can hinder transactions and user experience (UX) [3].
Although recent user-centered research has examined how security and usability influence cryptocurrency wallet selection and use [4], large-scale review-based studies that systematically evaluate and prioritize wallet software quality remain limited. User reviews provide direct evidence of problems encountered during actual application use and can reveal software strengths, weaknesses, and user requirements [5]. However, frequently discussed issues do not necessarily represent the areas requiring the greatest improvement, as highly frequent topics may be associated with relatively high satisfaction, whereas less frequent topics may reflect substantial dissatisfaction.
To address this issue, this study integrates Bidirectional Encoder Representations from Transformers Topic modeling (BERTopic) with the ISO/IEC 25010:2011 product quality model. BERTopic is used to identify major topics from cryptocurrency wallet reviews, and Generative Pre-trained Transformer (GPT) o1-mini generates preliminary topic labels based on the extracted keywords. The labels are subsequently reviewed and refined by the authors, who manually map the identified topics to the corresponding ISO/IEC 25010:2011 quality characteristics.
Based on this mapping, the study combines topic frequency and user dissatisfaction to calculate weighted values for quality prioritization. Topic frequency captures the prevalence of each user concern, while dissatisfaction reflects the average rating associated with the topic. Considering both dimensions enables the analysis to distinguish frequently discussed issues from those associated with stronger dissatisfaction. Because a single topic may correspond to multiple quality characteristics, fractional allocation is applied to prevent duplicated contributions. A sensitivity analysis comparing fractional and full allocation is also conducted to examine the effect of the allocation rule on the resulting rankings.
Based on the proposed framework and the research gap outlined above, the study addresses the following research questions:
RQ1. What major software quality issues can be identified from cryptocurrency wallet user reviews, and how are these issues associated with the ISO/IEC 25010:2011 quality characteristics?
RQ2. Which ISO/IEC 25010:2011 quality characteristics receive the highest priorities when topic frequency and user dissatisfaction are jointly considered?
RQ3. How sensitive are the quality-priority rankings to the allocation method used for topics mapped to multiple quality characteristics?
Addressing these research questions, the results identify Security as the highest-priority quality characteristic under fractional allocation, followed by Performance Efficiency and Reliability. The sensitivity analysis further demonstrates that the allocation rule can affect the resulting priority structure, highlighting the importance of explicitly accounting for multi-label topic mappings in software quality prioritization.
Beyond these empirical findings, the study makes three main contributions. First, it provides a user-centered software quality assessment by identifying major UX and software quality issues from large-scale cryptocurrency wallet reviews and mapping them to ISO/IEC 25010:2011 quality characteristics. Second, it introduces a dissatisfaction-weighted prioritization approach that jointly considers topic frequency and user ratings. Third, it incorporates fractional allocation and sensitivity analysis to account for topics associated with multiple quality characteristics and to evaluate the robustness of the resulting rankings. Overall, the proposed framework provides a transparent and user-centered basis for identifying relative software quality priorities from large-scale review data.
The remainder of this manuscript is organized as follows. Related Work reviews previous studies on user review-based quality assessment and applications of the ISO/IEC 25010:2011 product quality model. Materials and Methods describes the research procedure, including data collection, BERTopic-based topic modeling, quality-characteristic mapping, weighted prioritization, and sensitivity analysis. Results present the identified topics and quality priorities, followed by a discussion of their implications, limitations, and future research directions. The manuscript concludes by summarizing the main findings and contributions.

2. Related Work

This section reviews previous studies that provide the conceptual and methodological foundation for this research. The literature is organized into two areas: (1) user review-based approaches to software and service quality assessment and prioritization, and (2) applications of the ISO/IEC 25010:2011 product quality model in software quality evaluation. Together, these research streams provide the basis for integrating user-derived quality concerns with a standardized software quality framework.

2.1. User Review-Based Quality Assessment and Prioritization

User reviews provide direct evidence of UX, satisfaction, and problems encountered during actual service use. Large-scale review analysis can reveal both explicit evaluations and latent patterns that may not be captured through aggregate ratings or conventional surveys alone.
Previous studies have demonstrated the value of review data across various service domains. Ye et al. [6] found that positive reviews significantly increased hotel bookings, while Archak et al. [7] showed that product attributes extracted from review text influenced consumer decisions and that negative information could have particularly strong predictive value. Chintagunta et al. [8] similarly demonstrated the influence of online reviews on movie box-office performance. Other studies have focused on the informational value and credibility of reviews. Cheung et al. [9] showed that argument quality, source credibility, and review consistency influence users’ perceptions of review usefulness and trust. Chen and Xie [10] emphasized the role of consumer reviews as a distinct source of user-generated product information, while Hu and Liu [11] demonstrated that feature-based review mining can systematically identify product-specific strengths and weaknesses.
More recent research has extended review mining toward software engineering and digital financial services. Through a systematic review of 167 studies, Wang et al. [12] identified review text, ratings, sentiment, and domain-specific information as important sources for extracting actionable knowledge from application reviews. In the mobile financial-service context, Yi et al. [13] applied BERTopic and generative artificial intelligence (AI) to mobile trading application reviews and showed that topic frequency and rating-related information can jointly reveal major drivers of user satisfaction and areas requiring improvement.
Recent studies have also moved beyond issue identification toward systematic prioritization. Mihany et al. [14] proposed a data-driven framework for prioritizing software requirements from application reviews, demonstrating the potential of user-generated feedback to support systematic software development prioritization. However, review frequency alone may overemphasize commonly discussed issues even when users are relatively satisfied, whereas ratings or sentiment alone may overlook how widely an issue occurs. Effective prioritization therefore requires consideration of both the prevalence of a concern and the degree of dissatisfaction associated with it.
Building on these insights, this study analyzes large-scale cryptocurrency wallet reviews to identify user-reported software quality concerns and determine their relative priorities. Topic frequency and user dissatisfaction are jointly incorporated into the prioritization process, providing a user-centered basis for assessing both the prevalence and severity of software quality concerns.

2.2. Applications of ISO/IEC 25010:2011 in Software Quality Assessment

The ISO/IEC 25010:2011 product quality model defines eight quality characteristics and provides a structured framework for evaluating and improving software quality [15]. Its applicability has been demonstrated across various software domains, where the framework has been used to identify strengths and weaknesses in system quality and support improvement decisions.
Pratama and Mutiara [16] applied ISO/IEC 25010:2011 to evaluate the Halodoc telemedicine application using black-box testing, stress testing, and user surveys. The application achieved an overall score of 4.515 out of 5, indicating generally high software quality, while Security and Portability were identified as areas requiring further improvement. Similarly, Canlas et al. [17] assessed a Faculty Research Productivity Monitoring and Prediction System based on characteristics including Functional Suitability, Performance Efficiency, and Usability. The system achieved an average evaluation score of 3.87, suggesting that it was suitable for deployment while still presenting opportunities for further quality enhancement.
In the healthcare domain, Oliveira and Peres [18] evaluated an electronic nursing documentation system with 37 evaluators, including nurses and information technology specialists. Most quality characteristics received positive evaluations above 70%; however, response time, data interoperability, and error recovery were identified as areas requiring improvement. These findings illustrate the usefulness of the ISO/IEC 25010 framework for identifying specific quality limitations even when the overall evaluation of a system is favorable.
Keibach and Shayesteh [19] applied the framework to five software tools used in climate adaptation planning. Their evaluation showed that the tools provided different functional strengths but shared limitations related to interoperability, Performance Efficiency, and data processing. The study also emphasized the importance of Usability and Functional Suitability in supporting effective software use. In another context, Karnouskos et al. [20] applied software quality criteria to the integration of industrial automation systems with software agents. Their assessment highlighted Reliability, Maintainability, and Compatibility as particularly important characteristics for successful system integration.
Collectively, these studies demonstrate that ISO/IEC 25010:2011 can be applied across diverse software environments to systematically identify quality strengths and areas requiring improvement. However, previous applications have commonly relied on system testing, surveys, or expert-based assessments. The present study extends this line of research by deriving quality concerns from large-scale user reviews and mapping the identified topics to ISO/IEC 25010:2011 quality characteristics. These review-derived characteristics are then combined with information on topic frequency and user dissatisfaction to establish relative priorities for software quality improvement.

3. Materials & Methods

This study is organized into several key phases: data collection and preprocessing, topic modeling using BERTopic, mapping of extracted topics to ISO/IEC 25010:2011 quality characteristics, computation of dissatisfaction-weighted values, and prioritization of quality characteristics using fractional allocation and sensitivity analysis (Figure 1).
Prior to detailing each phase, the versions of the main Python 3.12 libraries utilized are provided in Table 1.

3.1. Data Collection

User reviews were collected from the Google Play Store for cryptocurrency wallet applications with over 500,000 downloads. The search was conducted using the query “cryptocurrency wallet,” and only seven applications met the inclusion criterion at the time of data access. The selected applications were Exodus, MetaMask, SafePal, Trust Wallet, Guarda Wallet, Bitcoin.com Wallet, and Bitget Wallet. Reviews were collected exclusively in English using the Python google-play-scraper library. The scraper collected all available review data accessible through Google Play for the selected applications, covering the period from 28 June 2017 to 26 January 2025. The dataset initially comprised 154,169 reviews, including review content, star ratings, and review dates. Duplicate reviews were removed, while edited reviews were not separately processed because the collected dataset reflected the review records available at the time of scraping. The data were accessed on 10 February 2025.
This study focused on Google Play reviews because large-scale review collection from the Apple App Store is more restricted. Apple’s public Really Simple Syndication (RSS)-based review access generally provides only a limited number of recent review pages and requires country-specific collection, which limits the construction of a comparable large-scale longitudinal dataset. The distribution of star ratings across the collected reviews is illustrated in Figure 2.

3.2. Preprocessing

The preprocessing procedure consisted of the following steps. First, special characters and non-English text were removed, followed by the elimination of unnecessary whitespace. Next, stopwords were removed, and word normalization was performed, including handling of singular/plural forms, past tense, and synonyms. Subsequently, words with fewer than two characters, except for key terms such as “UI” and “UX,” as well as empty strings, were filtered out. Lastly, reviews with fewer than 15 characters were discarded. Following these preprocessing steps, the final experimental dataset consisted of 68,790 reviews.

3.3. Topic Modeling Process with BERTopic

This study utilized BERTopic to extract topics from the user review dataset. BERTopic is a topic modeling method that integrates transformer-based embeddings, dimensionality reduction, and density-based clustering, enabling the generation of coherent topics [21]. Unlike traditional methods such as Latent Dirichlet Allocation (LDA), BERTopic leverages semantically rich embeddings, making it particularly effective in analyzing large-scale textual data where contextual meaning is crucial.
The process began by generating pre-trained sentence embeddings through the SentenceTransformers library. These embeddings represented each review as a high-dimensional vector encapsulating its semantic meaning [22]. To facilitate clustering, the vectors were reduced in dimensionality using the Uniform Manifold Approximation and Projection (UMAP) algorithm, preserving essential information while simplifying the vector space [23].
Subsequently, Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) was applied to identify clusters within the reduced vector space, with each cluster corresponding to a distinct topic derived from user review [24]. Afterwards, the CountVectorizer was employed to convert the text into a token count matrix, filtering out common English stop words and retaining the most relevant tokens based on their document frequency. Additionally, overly common or infrequent tokens were excluded in this step by setting min_df and max_df hyperparameters.
Then the Class TfidfTransformer was applied to enhance the significance of informative terms by adjusting their frequencies based on class relevance, improving the accuracy and relevance of the subsequent topic clustering. Finally, the data underwent analysis using BERTopic to identify the top ten words within each cluster. Detailed specifications of the hyperparameters utilized in the experiment are presented in Table 2.
The choice of BERTopic as the topic modeling approach was informed by a comprehensive and systematically established framework of inclusion and exclusion criteria. The chosen method was required to effectively capture semantic relationships by leveraging transformer-based embeddings. Additionally, it needed to demonstrate scalability for handling a large-scale user review dataset and support the clustering of short, user-generated, and often noisy textual data. Traditional topic modeling approaches, such as LDA and Non-negative Matrix Factorization (NMF), were excluded due to their dependence on frequency-based representations, which inadequately capture contextual meaning. Similarly, embedding-based methods like Word2Vec and Doc2Vec were not adopted, as they lack the capability to represent semantic information at the sentence level with the precision offered by transformer-based models. Based on these methodological characteristics, BERTopic was selected as an appropriate approach for identifying contextually meaningful topics from the short and heterogeneous user reviews analyzed in this study.
Following the extraction of clusters through BERTopic, topic labeling was conducted using GPT o1-mini [25]. Figure 3 presents the fixed instruction template used for topic labeling. The same instruction was applied across all topics, while the topic-specific input was programmatically inserted as a dictionary containing the top ten keywords and their corresponding relevance scores for each BERTopic cluster. Thus, the figure presents the fixed instructional component of the prompt, whereas the keyword–score pairs varied according to the topic being labeled.
This step generated concise labels derived from the top ten keywords associated with each topic. The GPT o1-mini-generated labels were independently reviewed by the two authors for semantic adequacy based on the extracted keywords and cluster content. Labels considered unclear or insufficiently representative were revised through discussion until consensus was reached. Because topic labeling does not involve a single predefined correct answer, the validation focused on the semantic appropriateness of the proposed labels rather than exact label agreement. Manual inspection has been commonly applied in topic modeling studies to enhance semantic accuracy and domain relevance [26,27]. Recent research on large language model (LLM)-assisted qualitative coding similarly emphasizes the continued role of human interpretive oversight in reviewing and refining model-generated classifications [28].

3.4. Topic Mapping with the ISO/IEC 25010:2011 Quality Characteristics

The ISO/IEC 25010:2011 product quality model provides a structured framework for evaluating software quality and defines eight key quality characteristics. These characteristics support the systematic assessment of software products across multiple dimensions and have been widely applied in both academic and practical software evaluation contexts. Table 3 presents the definitions of the eight quality characteristics used in this study.
In this phase, the two authors independently assessed the mapping of the BERTopic-derived topics to the eight ISO/IEC 25010:2011 quality characteristics based on the topic labels, extracted keywords, and the definitions of the quality characteristics. Each topic–characteristic pair was coded as applicable or not applicable, and inter-rater agreement was assessed using Cohen’s kappa (κ = 0.82) [29]. Disagreements were subsequently resolved through discussion and consensus. Because a topic could correspond to more than one quality characteristic, multiple mappings were permitted. The resulting mappings were used to calculate weighted values and determine relative priorities for quality improvement.

3.5. Weighted Value Calculation for Prioritizing Quality Improvements

Building upon the mapping of user feedback to the ISO/IEC 25010:2011 quality characteristics, this section describes the method used to assess the relative priority for improvement of each quality characteristic. The approach combines two indicators: the frequency of occurrence of each topic in user reviews and the average star rating associated with that topic. Topics that appear frequently and have low average ratings indicate stronger user dissatisfaction and therefore higher priority for improvement. These weighted values are used to determine the relative priority of quality characteristics requiring improvement.

3.5.1. Frequency Normalization

To address the wide variability in document frequencies across topics, a logarithmic transformation was applied to attain normalized frequency in Equation (1):
T i = log ( 1 + F i )
where T i represents the normalized frequency, and F i denotes the number of reviews associated with topic i . This transformation helps to reduce the impact of extreme variations in term frequency, ensuring that highly frequent topics do not disproportionately influence the analysis. This approach is well-documented in the information retrieval literature for term weighting [30] and has been employed to stabilize data distributions [31,32,33]. Textual data typically exhibit a long-tail distribution characterized by a small number of very frequent terms and a large number of rare terms. Applying a logarithmic transformation to raw frequencies mitigates the disproportionate influence of highly frequent topics, thereby preventing dominant terms from skewing the analysis [34]. This normalization enhances statistical stability by reducing variance and improving the balance of term weighting across topics.

3.5.2. Dissatisfaction Score Calculation

The dissatisfaction score was derived from user review ratings. Since review scores range from 1 to 5, the dissatisfaction score was computed by subtracting the average rating from the maximum score in Equation (2):
D i = 5 A v e r a g e   S c o r e i
where D i is the dissatisfaction score for topic i , and A v e r a g e   S c o r e i represents the average star rating for topic i . This method effectively captures the gap between user expectations and their actual experiences, providing a quantitative measure of user dissatisfaction. Topics with both frequent mentions and low average ratings highlight areas of heightened user dissatisfaction, indicating a pressing need for improvement. In this study, higher D i values signify greater user dissatisfaction and thus higher urgency for addressing those topics.

3.5.3. Weighted Value Calculation

The final weighted value for each topic was calculated by multiplying the normalized frequency with the dissatisfaction score. These weighted values were then aggregated to obtain the final weighted value for each quality factor. Equations (3) and (4) are as follows:
W e i g h t e d   V a l u e i = T i × D i
W e i g h t e d   V a l u e q = i = 1 n I i q · W e i g h t e d   V a l u e i m i
where W e i g h t e d   V a l u e i is the weighted value of topic i , W e i g h t e d   V a l u e q represents the aggregated weighted value of quality factor q , I i q is an indicator variable that equals 1 when topic i is mapped to quality factor q , and 0 otherwise. m i denotes the total number of quality characteristics associated with topic i .
This fractional allocation prevents a topic mapped to multiple quality characteristics from contributing its full weighted value repeatedly to each factor. Accordingly, the total contribution of each topic remains equal to its original weighted value regardless of the number of quality characteristics to which it is mapped. The resulting quality-factor weighted values integrate topic frequency and user dissatisfaction, with higher values indicating a greater relative priority for quality improvement. The same fractional allocation principle was also applied to topic frequencies when summarizing the descriptive review distribution across quality characteristics, ensuring that the total allocated review count remained equal to the number of reviews assigned to the identified topics.

3.5.4. Mapping Sensitivity Analysis

To assess the sensitivity of the prioritization results to the treatment of topics mapped to multiple ISO/IEC 25010:2011 quality characteristics, two allocation approaches were compared. Under full allocation, the entire weighted value of each topic was assigned to every quality characteristic associated with that topic. Under fractional allocation, the weighted value was equally divided among the associated quality characteristics, as specified in Equation (4). The resulting rankings were compared using Spearman’s rank correlation coefficient [35].

4. Results

4.1. Topic Modeling Results and ISO/IEC 25010:2011 Quality Characteristic Mapping

This section presents the results of topic modeling, along with the mapping of the extracted topics to the relevant ISO/IEC 25010:2011 quality characteristics (Table 4). Of the 68,790 preprocessed reviews, 41,203 were assigned to 15 topic clusters, while 27,587 were classified as outliers by HDBSCAN and excluded from the topic-level analysis. Topic labeling was performed using GPT o1-mini, and the identified topics were subsequently mapped to the corresponding ISO/IEC 25010:2011 quality characteristics. Table 4 summarizes the resulting topic labels, representative keywords, and quality-characteristic mappings. To provide further clarity, each topic is briefly described based on the extracted keywords and its relationship with the corresponding quality characteristics [36].
The first topic, labeled “User Interface & Experience,” emphasizes simplicity, ease of navigation, and user-friendliness. Keywords such as “Simple,” “User,” “Interface,” “Navigate,” and “Friendly” highlight the importance of intuitive interaction, while “Fast,” “Smooth,” “Seamlessly,” and “Convenient” indicate users’ expectations for efficient and effortless use of the application. These characteristics are closely associated with Usability and Functional Suitability in the ISO/IEC 25010:2011 model.
The second topic, labeled “Security & Trustworthiness,” focuses on users’ concerns regarding data protection and confidence in the application. Keywords such as “Security,” “Safe,” and “Trusted” reflect the importance users place on protecting sensitive information and maintaining system integrity. Additional terms such as “Trustworthy” and “Recommend” suggest that perceptions of security are closely related to users’ confidence in the application. This topic was therefore mapped to Security.
The third topic, labeled “Application Installation & Loading,” highlights issues related to system responsiveness, installation processes, and device compatibility. Keywords such as “Download,” “Loading,” and “Installation” indicate delays or difficulties encountered when setting up or accessing the application, while “Slow” and “Fix” further reflect performance-related concerns. Terms such as “Mobile,” “Phone,” and “Desktop” also indicate expectations for consistent operation across different platforms. Accordingly, this topic was mapped to Performance Efficiency, Maintainability, Compatibility, and Portability.
The fourth topic, labeled “Transaction Fees & Gas Costs,” focuses on concerns related to financial transactions within the application. Keywords such as “Gas,” “Fee,” and “Transaction” indicate users’ sensitivity to transaction costs and processing efficiency, while terms such as “Charge,” “Pay,” and “Withdrawal” reflect concerns regarding financial operations. Keywords including “Convert,” “Token,” “Swap,” and “Tether United States Dollar (USDT)” further indicate expectations for smooth asset exchanges and transaction functions. These characteristics were mapped to Usability and Functional Suitability.
The fifth topic, labeled “Application Responsiveness,” highlights user concerns regarding response time, delays, and overall application performance. Keywords such as “Open,” “Opening,” and “Time” indicate delayed responses during application startup or navigation, while “Slow,” “Hanging,” and “Long” further reflect performance-related frustration. In contrast, “Fast” represents users’ expectations for improved responsiveness, and “Fix” suggests recurring issues requiring technical correction. These concerns are strongly linked to Performance Efficiency, Maintainability, and Reliability.
The sixth topic, labeled “Fraud & Scam Concerns,” focuses on security threats and fraudulent activities. Keywords such as “Fraud,” “Scam,” and “Hack” reflect concerns about potential risks to personal and financial security, while “Money,” “Fund,” and “Lost” indicate fears of financial loss or unauthorized transactions. Terms such as “Transaction,” “Transfer,” and “Coin” further suggest that these concerns are closely associated with financial operations within cryptocurrency wallet applications. This topic was therefore mapped to Security and Maintainability.
The seventh topic, labeled “Update & Bug Fix,” highlights problems related to application updates and unresolved technical issues. Keywords such as “Update,” “Bug,” and “Fix” emphasize recurring problems affecting application stability, while “Crash,” “Version,” and “Latest” indicate that failures may occur during or after software updates. Terms such as “Updating” and “Open” also reflect difficulties associated with updated versions of the application. These concerns were mapped to Maintainability and Reliability.
The eighth topic, labeled “Loading & Bugs Issues,” addresses recurring delays and technical malfunctions. Keywords such as “Lagging,” “Loading,” and “Slow” indicate prolonged waiting times and sluggish performance, while “Bug,” “Glitch,” and “Problem” reflect persistent technical errors. Terms such as “Forever” and “Worst” further emphasize severe dissatisfaction associated with these problems. This topic was mapped to Performance Efficiency, Reliability, and Maintainability.
The ninth topic, labeled “Device Compatibility,” highlights concerns regarding multi-platform support and operation across different devices. Keywords such as “Android,” “iPhone,” and “Device” reflect users’ expectations for consistent application performance across operating systems and hardware configurations. Additional terms such as “Samsung,” “Apple,” “Windows,” and “Desktop” indicate the importance of reliable functionality across both mobile and desktop environments. These characteristics were mapped to Portability, Compatibility, and Functional Suitability.
The tenth topic, labeled “Network & Connectivity Issues,” focuses on unstable network connections and related technical problems. Keywords such as “Connect,” “Network,” and “Internet” indicate disruptions that may interfere with application use, while “Server,” “Problem,” and “Issue” reflect recurring network-related errors. The term “Abnormal” suggests irregular network behavior, whereas “Fix” indicates the need for resolution of such problems. This topic was mapped to Performance Efficiency and Reliability.
The eleventh topic, labeled “Account Security & Recovery,” highlights concerns regarding secure account access and recovery processes. Keywords such as “Password,” “Recovery,” and “Login” indicate difficulties in managing credentials and recovering access, while “Phrase,” “Secret,” and “Fingerprint” reflect users’ expectations for secure authentication mechanisms. Terms such as “Lost” and “Recover” further emphasize the importance of reliable account recovery. These concerns were mapped to Security and Usability.
The twelfth topic, labeled “Trading & Transaction Experience,” focuses on user experiences associated with financial operations within the application. Keywords such as “Exchange,” “Trading,” and “Transaction” indicate frequent interaction with trading and transaction functions, while “Smooth,” “Seamlessly,” and “Experience” reflect expectations for convenient and uninterrupted processes. Additional terms such as “Swap,” “Transfer,” and “Exchanging” indicate the importance of flexible asset conversion and transfer functions. This topic was mapped to Functional Suitability and Usability.
The thirteenth topic, labeled “Functional Suitability & Performance,” emphasizes concerns related to core application functions and performance efficiency. Keywords such as “Working,” “Smoothly,” and “Properly” reflect expectations that essential functions operate consistently without interruption. Additional terms such as “Functional,” “Functions,” and “Performing” further emphasize the importance of correct operation of application features. This topic was mapped to Functional Suitability and Performance Efficiency.
The fourteenth topic, labeled “Performance & Speed,” highlights concerns related to processing speed and system stability. Keywords such as “Speed,” “Fast,” and “Time” indicate expectations for efficient operation, whereas “Crash,” “Freeze,” and “Sluggish” reflect problems involving system failures and unresponsiveness. Additional terms such as “Slowly” and “Poor” further indicate dissatisfaction with application performance. These characteristics were mapped to Performance Efficiency and Reliability.
The fifteenth topic, labeled “Token & Balance Updates,” focuses on accurate and timely updates of token balances and asset-related information. Keywords such as “Update,” “Token,” and “Coin” indicate issues related to balance updates and token management, while “Updating,” “List,” and “Add” reflect expectations regarding the integration of new tokens and functions. Terms such as “Support” and “Chains” further suggest the importance of supporting multiple blockchain networks and digital assets. This topic was mapped to Functional Suitability and Maintainability.

4.2. Prioritization of Quality Characteristics Through Weighted Values

Table 5 summarizes the topics associated with each ISO/IEC 25010:2011 quality factor and the corresponding review counts after fractional allocation. Because several topics were mapped to more than one quality factor, assigning the full number of reviews to every mapped factor would result in duplicate counting. Therefore, the review count of each topic was equally distributed across the quality characteristics to which it was mapped. As a result, the fractionally allocated counts across all quality characteristics sum to 41,203, corresponding to the total number of reviews assigned to the 15 topic clusters. Functional Suitability shows the highest allocated review count, followed by Security, Usability, and Maintainability. However, review frequency alone does not fully represent the urgency of quality improvement; the level of dissatisfaction associated with each topic must also be considered.
Table 6 presents the average rating for each topic. Lower average ratings indicate higher levels of user dissatisfaction. Topics 2, 3, 9, 10, 11, and 14 show comparatively low average ratings, with Topic 3 (“Application Installation & Loading”) and Topic 9 (“Device Compatibility”) receiving particularly low scores of 2.175084 and 2.223176, respectively. In contrast, Topics 1 and 8 show relatively high average ratings of 4.637310 and 4.601473. These topic-level ratings were combined with topic frequency to calculate the weighted values used for quality-factor prioritization.
Table 7 presents the weighted values and priority rankings of the ISO/IEC 25010:2011 quality characteristics under both fractional and full allocation. The primary analysis uses fractional allocation, in which each topic-level weighted value is equally divided among the quality characteristics to which the topic is mapped. This approach prevents topics associated with multiple quality characteristics from contributing their full weighted value repeatedly.
Under fractional allocation, Security has the highest weighted value (35.0166), followed by Performance Efficiency (27.9102), Reliability (26.5742), Maintainability (23.0377), Functional Suitability (22.3119), and Usability (18.0978). Compatibility and Portability have identical weighted values of 13.0784 and share the lowest rank. Accordingly, Security emerges as the highest-priority quality factor for improvement, followed by Performance Efficiency and Reliability.
To assess the sensitivity of the prioritization results to the topic-to-factor allocation rule, the fractional allocation results were compared with those obtained using full allocation, in which the entire topic-level weighted value was assigned to each associated quality factor. The comparison shows notable changes in the priority structure. Under full allocation, Performance Efficiency ranks first, followed by Maintainability and Reliability, whereas under fractional allocation, Security ranks first, followed by Performance Efficiency and Reliability. Security shows the largest rank change, moving from fifth under full allocation to first under fractional allocation. Maintainability moves from second to fourth, whereas Reliability remains third under both approaches.
The Spearman rank correlation between the two allocation approaches was 0.663 (p = 0.073), indicating moderate rank correspondence, although the association was not statistically significant at the 0.05 level. Nevertheless, the observed changes in individual rankings, particularly the shift in Security from fifth to first, demonstrate that the prioritization results are substantively sensitive to the allocation rule.
The results indicate that full allocation can amplify the influence of topics mapped to multiple quality characteristics because the same topic-level weighted value is repeatedly assigned to each associated factor. Fractional allocation was therefore adopted as the primary specification because it preserves the total contribution of each topic while avoiding duplicate weighting across multiple quality characteristics.

5. Discussion

This study proposes a user-centered approach for evaluating and prioritizing software quality characteristics in cryptocurrency wallet applications based on user review data. By integrating BERTopic with the ISO/IEC 25010:2011 quality model, the proposed framework transforms latent user concerns into structured quality characteristics and combines topic frequency with dissatisfaction scores to determine relative priorities for improvement. Fractional allocation further addresses the potential duplication that occurs when a single topic is mapped to multiple quality characteristics.
The results indicate that Security is the highest-priority quality characteristic. Security-related topics reflected concerns regarding fraud, scams, unauthorized access, account recovery, loss of funds, and protection of sensitive credentials. These concerns are particularly important in cryptocurrency wallet applications because users directly manage digital assets and authentication information through the application. Consequently, perceived security problems may have substantial implications for users’ confidence in wallet services and their willingness to continue using them.
Performance Efficiency ranked second. User feedback frequently reflected issues related to application responsiveness, loading delays, transaction processing, and network connectivity. These findings suggest that users expect wallet applications to provide timely and stable interactions, particularly when conducting transactions or accessing rapidly changing financial information. Although Performance Efficiency was not the highest-ranked characteristic under fractional allocation, its consistently high priority indicates that application responsiveness remains an important component of the cryptocurrency wallet user experience.
Reliability ranked third. Topics associated with crashes, unstable updates, connectivity problems, and inconsistent application behavior indicate that users are sensitive to disruptions that may interfere with access to or management of financial assets. Reliability is particularly relevant in financial software because users expect transaction-related functions and account access to operate consistently. Previous studies have similarly emphasized the importance of reliability in financial service environments [37,38,39].
Maintainability ranked fourth and was mainly associated with bug fixes, application updates, recurring technical problems, and the need for continued software improvement. Cryptocurrency wallet applications operate in an environment characterized by frequent technical changes, evolving blockchain ecosystems, and emerging security threats. Therefore, effective maintenance remains important even though its relative priority was lower than Security, Performance Efficiency, and Reliability.
An important methodological finding concerns the treatment of topics mapped to multiple quality characteristics. Under full allocation, Performance Efficiency ranked first, followed by Maintainability and Reliability. In contrast, fractional allocation resulted in Security ranking first, followed by Performance Efficiency and Reliability. Security showed the largest change, moving from fifth to first. The Spearman rank correlation between the two allocation approaches was 0.663 (p = 0.073), indicating moderate correspondence between the rankings but also substantial changes in several individual positions. These findings show that the allocation rule can materially influence the resulting priority structure.
Fractional allocation was therefore used as the primary approach because it preserves the total weighted contribution of each topic. Under full allocation, a topic mapped to several quality characteristics contributes its entire weighted value to each characteristic, potentially increasing the influence of characteristics associated with a larger number of multi-mapped topics. Fractional allocation reduces this duplication by distributing the topic-level weighted value equally across its mapped characteristics. This approach provides a more conservative basis for interpreting relative quality priorities while retaining the multi-dimensional nature of user concerns.
More broadly, the findings demonstrate that review frequency alone is insufficient for identifying improvement priorities. A highly frequent topic may represent generally positive experiences, whereas a less frequent topic associated with low ratings may indicate a more serious quality concern. Combining topic frequency and dissatisfaction therefore provides additional information for prioritizing software quality characteristics from a user-centered perspective.

5.1. Managerial and Practical Implications

The findings provide practical guidance for cryptocurrency wallet developers, product managers, and service operators. The priority structure suggests that Security, Performance Efficiency, and Reliability warrant particular attention when allocating development resources based on user-reported concerns.
For Security, wallet providers should focus on strengthening account protection, fraud prevention, authentication, and recovery processes. User concerns related to scams, lost funds, passwords, recovery phrases, and unauthorized access indicate that security mechanisms should be accompanied by clear and usable guidance. Recovery phrase management may be informed by established wallet specifications such as Bitcoin Improvement Proposal 39 (BIP-39), which defines mnemonic codes for deterministic wallet generation [40]. Authentication and account-management practices may also draw on current digital identity guidance such as National Institute of Standards and Technology Special Publication 800-63-4 (NIST SP 800-63-4) [41].
For Performance Efficiency, developers can focus on reducing loading times, improving application responsiveness, optimizing transaction-processing workflows, and providing clearer feedback during network congestion or delayed transaction confirmation. These improvements may reduce user frustration in situations where timely interaction is particularly important.
Reliability-related priorities include preventing crashes, improving stability after updates, strengthening error handling, and providing clear information when transactions or network connections fail. Continuous monitoring and systematic failure detection may assist developers in identifying recurring problems before they substantially affect user experience.
Maintainability also remains relevant because user reviews repeatedly referred to updates, bugs, and unresolved technical problems. Stable version management, timely bug fixes, and systematic maintenance procedures can support the sustained performance and reliability of wallet applications as technical requirements evolve.
From a managerial perspective, the proposed approach enables development teams to distinguish between frequently discussed issues and issues associated with stronger dissatisfaction. This can support more informed prioritization of limited development resources. The framework can also be reapplied periodically as new review data become available, allowing organizations to examine whether user concerns and relative quality priorities change over time. However, the present analysis does not directly estimate effects on return on investment, user retention, regulatory compliance, market competitiveness, or causal changes in user trust. Such outcomes should therefore be regarded as potential managerial implications rather than empirically demonstrated effects.

5.2. Limitations and Future Work

This study has several limitations. First, the dataset includes only English-language Google Play reviews from seven cryptocurrency wallet applications with more than 500,000 downloads. This limits generalizability across platforms, regions, languages, and smaller or newer applications.
Second, the reviews span from 28 June 2017 to 26 January 2025. Some complaints may therefore reflect issues in earlier application versions that were subsequently resolved. Future research could conduct longitudinal analyses to examine changes in quality priorities over time.
Third, manually mapping BERTopic-derived topics to ISO/IEC 25010:2011 quality characteristics involves interpretive judgment. Future studies could use multiple independent coders and report inter-rater agreement to strengthen mapping reliability.
Fourth, the sensitivity analysis showed that the allocation rule for multi-mapped topics can influence the final rankings. Although fractional allocation prevents duplicate weighting, it assumes equal contribution across mapped quality characteristics. Future research could examine alternative allocation schemes.
Finally, the dissatisfaction measure was based on the difference between the maximum rating and the average topic rating. Alternative measures, such as negative-review proportions or review-level sentiment, could be examined to further assess the robustness of the prioritization results.

6. Conclusions

This study integrated BERTopic with the ISO/IEC 25010:2011 product quality model to prioritize software quality characteristics in cryptocurrency wallet applications based on user reviews. Fifteen topics were manually mapped to quality characteristics, and topic frequency and user dissatisfaction were combined to calculate weighted priorities.
Using fractional allocation, Security was identified as the highest-priority quality characteristic, followed by Performance Efficiency and Reliability. The sensitivity analysis showed that the allocation rule can affect the resulting rankings, indicating the importance of accounting for multi-label topic mappings when aggregating user-review-based quality indicators.
The proposed framework provides a transparent and user-centered approach for identifying software quality priorities from review data. Future research could extend the analysis to additional platforms and languages, longitudinal settings, and alternative dissatisfaction and allocation measures.

Author Contributions

Conceptualization, H.S.J. and H.L.; methodology, H.S.J.; software, H.S.J.; validation, H.S.J. and H.L.; formal analysis, H.S.J.; investigation, H.S.J.; resources, H.L.; data curation, H.S.J.; writing—original draft preparation, H.S.J.; writing—review and editing, H.L.; visualization, H.S.J.; supervision, H.L.; project administration, H.L. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Dongguk University Research Fund under Grant S-2025-G0001-00139. This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Korean government (MSIT) (RS-2026-25481681). This research was supported by the “Regional Innovation System & Education (RISE)” through the Seoul RISE Center, funded by the Ministry of Education (MOE) and the Seoul Metropolitan Government (2026-RISE-01-018-01).

Institutional Review Board Statement

Not applicable. This study analyzed publicly available Google Play review data and did not involve human subjects intervention or the collection of personally identifiable information.

Informed Consent Statement

Not applicable.

Data Availability Statement

The aggregated and derived data supporting the findings of this study are available from the corresponding author upon reasonable request. The complete dataset containing verbatim Google Play review text is not publicly available because the reviews constitute third-party user-generated content and their redistribution may be subject to platform terms and privacy considerations. De-identified derived data and analysis materials necessary to support the reported results may be made available upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Eyal, I. On cryptocurrency wallet design. In 3rd International Conference on Blockchain Economics, Security and Protocols (Tokenomics 2021); Schloss Dagstuhl—Leibniz-Zentrum für Informatik: Wadern, Germany, 2022. [Google Scholar]
  2. Albayati, H.; Kim, S.K.; Rho, J.J. A study on the use of cryptocurrency wallets from a user experience perspective. Hum. Behav. Emerg. Technol. 2021, 3, 720–738. [Google Scholar] [CrossRef] [Scilit]
  3. Halpin, H. Holistic privacy and usability of a cryptocurrency wallet. arXiv 2021, arXiv:2105.02793. [Google Scholar]
  4. Yu, Y.; Sharma, T.; Das, S.; Wang, Y. “Don’t Put All Your Eggs in One Basket”: How Cryptocurrency Users Choose and Secure Their Wallets. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA, 11–16 May 2024; Article 353; Association for Computing Machinery: New York, NY, USA, 2024; pp. 1–17. [Google Scholar] [CrossRef] [Scilit]
  5. Jung, H.S.; Kim, J.H.; Lee, H. Refining the prediction of user satisfaction on chat-based AI applications with unsupervised filtering of rating text inconsistencies. R. Soc. Open Sci. 2025, 12, 241687. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Ye, Q.; Law, R.; Gu, B. The impact of online user reviews on hotel room sales. Int. J. Hosp. Manag. 2009, 28, 180–182. [Google Scholar] [CrossRef] [Scilit]
  7. Archak, N.; Ghose, A.; Ipeirotis, P.G. Deriving the pricing power of product features by mining consumer reviews. Manag. Sci. 2011, 57, 1485–1509. [Google Scholar] [CrossRef] [Scilit]
  8. Chintagunta, P.K.; Gopinath, S.; Venkataraman, S. The effects of online user reviews on movie box office performance: Accounting for sequential rollout and aggregation across local markets. Mark. Sci. 2010, 29, 944–957. [Google Scholar] [CrossRef] [Scilit]
  9. Cheung, C.M.Y.; Sia, C.L.; Kuan, K.K. Is this review believable? A study of factors affecting the credibility of online consumer reviews from an ELM perspective. J. Assoc. Inf. Syst. 2012, 13, 2. [Google Scholar] [CrossRef] [Scilit]
  10. Chen, Y.; Xie, J. Online consumer review: Word-of-mouth as a new element of marketing communication mix. Manag. Sci. 2008, 54, 477–491. [Google Scholar] [CrossRef] [Scilit]
  11. Hu, M.; Liu, B. Mining and summarizing customer reviews. In Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Seattle, WA, USA, 22–25 August 2004; pp. 168–177. [Google Scholar]
  12. Wang, X.; Zhang, T.; Tan, Y.; Shang, W.; Li, Y. How to Effectively Mine App Reviews Concerning Software Ecosystem? A Survey of Review Characteristics. J. Syst. Softw. 2024, 213, 112040. [Google Scholar] [CrossRef] [Scilit]
  13. Yi, J.; Oh, Y.K.; Kim, J.M. Unveiling the drivers of satisfaction in mobile trading: Contextual mining of retail investor experience through BERTopic and generative AI. J. Retail. Consum. Serv. 2025, 82, 104066. [Google Scholar] [CrossRef] [Scilit]
  14. Mihany, F.A.; Galal-Edeen, G.H.; Hassanein, E.E.; Moussa, H. Data-Driven Requirements Prioritization Framework for App Reviews. Inventions 2026, 11, 33. [Google Scholar] [CrossRef] [Scilit]
  15. ISO/IEC 25010:2011; Systems and Software Engineering—Systems and Software Quality Requirements and Evaluation (SQuaRE)—System and Software Quality Models. International Organization for Standardization; International Electrotechnical Commission: Geneva, Switzerland, 2011. Available online: https://www.iso.org/standard/35733.html (accessed on 25 August 2026).
  16. Pratama, A.A.; Mutiara, A.B. Software quality analysis for Halodoc application using ISO 25010:2011. Int. J. Adv. Comput. Sci. Appl. 2021, 12, 383–392. [Google Scholar] [CrossRef] [Scilit]
  17. Canlas, R.B.; Piad, K.C.; Lagman, A.C. An ISO/IEC 25010:2011-based software quality assessment of a faculty research productivity monitoring and prediction system. In Proceedings of the 2021 9th International Conference on Information Technology: IoT and Smart City, Guangzhou, China, 22–25 December 2021; pp. 238–242. [Google Scholar]
  18. Oliveira, N.B.D.; Peres, H.H.C. Evaluation of the functional performance and technical quality of an electronic documentation system of the nursing process. Rev. Lat.-Am. Enferm. 2015, 23, 242–249. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Keibach, E.; Shayesteh, H. BIM for landscape design improving climate adaptation planning: The evaluation of software tools based on the ISO 25010 standard. Appl. Sci. 2022, 12, 739. [Google Scholar] [CrossRef] [Scilit]
  20. Karnouskos, S.; Sinha, R.; Leitão, P.; Ribeiro, L.; Strasser, T.I. Assessing the integration of software agents and industrial automation systems with ISO/IEC 25010:2011. In Proceedings of the 2018 IEEE 16th International Conference on Industrial Informatics (INDIN), Porto, Portugal, 18–20 July 2018; pp. 61–66. [Google Scholar]
  21. Grootendorst, M. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv 2022, arXiv:2203.05794. [Google Scholar]
  22. Reimers, N. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. arXiv 2019, arXiv:1908.10084. [Google Scholar]
  23. McInnes, L.; Healy, J.; Melville, J. UMAP: Uniform manifold approximation and projection for dimension reduction. arXiv 2018, arXiv:1802.03426. [Google Scholar]
  24. McInnes, L.; Healy, J.; Astels, S. HDBSCAN: Hierarchical density-based clustering. J. Open Source Softw. 2017, 2, 205. [Google Scholar] [CrossRef] [Scilit]
  25. Jaech, A.; Kalai, A.; Lerer, A.; Richardson, A.; El-Kishky, A.; Low, A.; Helyar, A.; Madry, A.; Beutel, A.; Carney, A.; et al. OpenAI o1 system card. arXiv 2024, arXiv:2412.16720. [Google Scholar]
  26. Jung, H.S.; Lee, H.; Woo, Y.S.; Baek, S.Y.; Kim, J.H. Expansive data, extensive model: Investigating discussion topics around LLM through unsupervised machine learning in academic papers and news. PLoS ONE 2024, 19, e0304680. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Yang, Z.; Wu, Q.; Venkatachalam, K.; Li, Y.; Xu, B.; Trojovský, P. Topic identification and sentiment trends in Weibo and WeChat content related to intellectual property in China. Technol. Forecast. Soc. Change 2022, 184, 121980. [Google Scholar] [CrossRef] [Scilit]
  28. Dunivin, Z.O. Scaling Hermeneutics: A Guide to Qualitative Coding with LLMs for Reflexive Content Analysis. EPJ Data Sci. 2025, 14, 28. [Google Scholar] [CrossRef] [Scilit]
  29. Cohen, J. A Coefficient of Agreement for Nominal Scales. Educ. Psychol. Meas. 1960, 20, 37–46. [Google Scholar] [CrossRef] [Scilit]
  30. Salton, G.; Buckley, C. Term-weighting approaches in automatic text retrieval. Inf. Process. Manag. 1988, 24, 513–523. [Google Scholar] [CrossRef] [Scilit]
  31. Kaplan, S.; Garrick, B.J. On the quantitative definition of risk. Risk Anal. 1981, 1, 11–27. [Google Scholar] [CrossRef] [Scilit]
  32. Huber, W.; Von Heydebreck, A.; Sültmann, H.; Poustka, A.; Vingron, M. Variance stabilization applied to microarray data calibration and to the quantification of differential expression. Bioinformatics 2002, 18, S96–S104. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  33. Osborne, J. Notes on the use of data transformations. Pract. Assess. Res. Eval. 2002, 8, 6. [Google Scholar]
  34. Zhan, Z.; Zhao, J.; Zhang, Y.; Gong, J.; Wang, Q.; Shen, Q.; Zhang, L. Grabbing the long tail: A data normalization method for diverse and informative dialogue generation. Neurocomputing 2021, 460, 374–384. [Google Scholar] [CrossRef] [Scilit]
  35. Spearman, C. The Proof and Measurement of Association between Two Things. Am. J. Psychol. 1904, 15, 72–101. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Jung, H.S.; Lee, H.; Kim, J.H. Unveiling cryptocurrency conversations: Insights from data mining and unsupervised learning across multiple platforms. IEEE Access 2023, 11, 130573–130583. [Google Scholar] [CrossRef] [Scilit]
  37. Korda, A.P.; Snoj, B. Development, validity and reliability of perceived service quality in retail banking and its relationship with perceived value and customer satisfaction. Manag. Glob. Transit. 2010, 8, 187. [Google Scholar]
  38. Ma, Z. Assessing serviceability and reliability to affect customer satisfaction of internet banking. J. Softw. 2012, 7, 1601–1608. [Google Scholar] [CrossRef] [Scilit]
  39. Iberahim, H.; Taufik, N.M.; Adzmir, A.M.; Saharuddin, H. Customer satisfaction on reliability and responsiveness of self-service technology for retail banking services. Procedia Econ. Financ. 2016, 37, 13–20. [Google Scholar] [CrossRef] [Scilit]
  40. Palatinus, M.; Rusnak, P.; Voisine, A.; Bowe, S. Mnemonic Code for Generating Deterministic Keys. Bitcoin Improvement Proposal 39 (BIP-39). 2013. Available online: https://github.com/bitcoin/bips/blob/master/bip-0039.mediawiki (accessed on 26 August 2026).
  41. Temoshok, D.; Proud-Madruga, D.; Choong, Y.-Y.; Galluzzo, R.; Gupta, S.; LaSalle, C.; Lefkovitz, N.; Regenscheid, A. Digital Identity Guidelines; NIST Special Publication 800-63-4; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2025. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Experimental flow diagram.
Figure 1. Experimental flow diagram.
Systems 14 01078 g001
Figure 2. Distribution of star ratings on collected review data.
Figure 2. Distribution of star ratings on collected review data.
Systems 14 01078 g002
Figure 3. Instructions given to GPT o1-mini for labeling clusters.
Figure 3. Instructions given to GPT o1-mini for labeling clusters.
Systems 14 01078 g003
Table 1. Python libraries and versions used in the experiment.
Table 1. Python libraries and versions used in the experiment.
LibraryVersion
google-play-scraper1.2.4
Nltk3.9.1
Sentence_transformers3.3.1
UMAP0.5.7
HDBSCAN0.8.40
Scikit-learn1.6.1
Bertopic0.16.4
Table 2. Hyperparameters employed for BERTopic analysis.
Table 2. Hyperparameters employed for BERTopic analysis.
ModuleHyperparameters
SentenceTransformer“distiluse-base-multilingual-cased-v2”
UMAPn_neighbors = 35min_dist = 0.04
n_components = 6Metric = “cosine”
random_state = 43
HDBSCANmin_cluster_size = 650min_samples = 15
Metric = “euclidean”cluster_selection_method = “eom”
Vectorizer Modelstop_words = “english”max_features = 18,000
max_df = 0.65min_df = 0.07
BERTopictop_n_words = 10
Table 3. ISO/IEC 25010:2011 quality characteristics with descriptions.
Table 3. ISO/IEC 25010:2011 quality characteristics with descriptions.
Quality
Characteristics
Description
Functional
Suitability
The degree to which a software product meets specified requirements and performs functions that meet user needs
Performance
Efficiency
The ability of the software to provide appropriate performance relative to resources used, such as response time and resource consumption
CompatibilityThe software’s ability to interact and function with other systems or environments without issues
UsabilityHow effectively and efficiently users can learn, operate, and interact with the software to achieve their goals
ReliabilityThe capability of the software to maintain performance and stability under specified conditions over time
SecurityThe degree to which the software protects data and prevents unauthorized access, breaches, and other security threats
MaintainabilityThe ease with which the software can be modified to correct defects, improve performance, or adapt to changing requirements
PortabilityThe ability of the software to be transferred from one environment to another with minimal effort
Table 4. Topic modeling results and ISO/IEC 25010:2011 mapping.
Table 4. Topic modeling results and ISO/IEC 25010:2011 mapping.
#CountTopic NameISO/IEC 25010:2011
Quality Characteristic
Top 5 Keywords
KeywordsScore
18911User Interface &
Experience
Usability,
Functional Suitability
Simple0.0489
User0.0333
Interface0.0331
Friendly0.0240
Navigate0.0230
25458Security & TrustworthinessSecuritySecurity0.0193
Safe0.0130
Reliable0.0120
Simple0.0478
Trusted0.0456
33813Installation &
Loading
Performance Efficiency, Maintainability,
Compatibility,
Portability
Download0.0462
Loading0.0379
Phone0.0349
Installation0.0319
Mobile0.0280
43421Transaction Fees
& Gas Costs
Usability,
Functional Suitability
Gas0.0485
Fee0.0435
Transaction0.0392
Convert0.0209
Swap0.0177
53105Application
Responsiveness
Performance Efficiency,
Reliability,
Maintainability
Open0.1227
Opening0.0711
Time0.0669
Slow0.0654
Long0.0472
62981Fraud &
Scam Concerns
Security,
Maintainability
Fraud0.1724
Money0.1127
Scam0.0466
Transaction0.0386
Fund0.0373
72600Update &
But Fix
Maintainability,
Reliability
Update0.0856
Version0.0483
Latest0.0448
Open0.0429
Updating0.0416
82571Loading &
Bugs Issues
Maintainability,
Reliability,
Performance Efficiency
Lagging0.0689
Loading0.0387
Time0.0220
Slow0.0218
Fix0.0217
92533Device
Compatibility
Compatibility, Functional Suitability
Portability
Android0.5321
Phone0.0471
iPhone0.0455
Samsung0.0386
Desktop0.0350
101213Network &
Connectivity Issues
Performance Efficiency, ReliabilityConnect0.1724
Network0.1127
Internet0.0466
Abnormal0.0386
Link0.0373
111170Account Security & RecoverySecurity,
Usability
Password0.1261
Phrase0.0709
Account0.0537
Recovery0.0487
Phone0.0475
121164Trading &
Transaction Experience
Functional Suitability,
Usability
Exchange0.0632
Trading0.0375
Transaction0.0301
Smooth0.0296
Trade0.0282
13788Functionality& PerformanceFunctional Suitability,
Performance Efficiency
Working0.2727
Smoothly0.0790
Properly0.0562
Job0.0494
Smooth0.0441
14750Performance & SpeedPerformance Efficiency,
Reliability
Speed0.0867
Fast0.0705
Time0.0647
Crash0.0570
Working0.0459
15725Token &
Balance Updates
Functional Suitability, MaintainabilityUpdate0.1110
Token0.0884
Coin0.0801
Custom0.0709
Add0.0603
Table 5. Review distribution by ISO/IEC 25010:2011 quality characteristics after fractional allocation.
Table 5. Review distribution by ISO/IEC 25010:2011 quality characteristics after fractional allocation.
Quality CharacteristicTopicFractionally Allocated
Review Count
Usability1, 4, 11, 127333.00
Security2, 6, 117533.50
Performance Efficiency3, 5, 8, 10, 13, 144220.75
Functional Suitability1, 4, 9, 12, 13, 158348.83
Maintainability3, 5, 6, 7, 8, 155998.25
Reliability5, 7, 8, 10, 144173.50
Compatibility3, 91797.58
Portability3, 91797.58
Table 6. Average user ratings by topic.
Table 6. Average user ratings by topic.
TopicAverage Rating
14.637310
22.445350
32.175084
44.093307
54.369715
63.971408
73.020917
84.601473
92.223176
102.798759
112.475302
123.912301
134.012563
142.509881
154.215068
Table 7. Comparison of weighted values and priority rankings under fractional and full allocation.
Table 7. Comparison of weighted values and priority rankings under fractional and full allocation.
Quality CharacteristicFractional AllocationFull Allocation
Weighted
Value
RankWeighted
Value
Rank
Security35.0166148.05055
Performance Efficiency27.9102270.20061
Reliability26.5742355.88103
Maintainability23.0377460.45582
Functional Suitability22.3119551.87834
Usability18.0978636.19568
Compatibility13.0784745.05906
Portability13.0784745.05906
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Jung, H.S.; Lee, H. Evaluating Software Quality in Cryptocurrency Wallet Applications: A User-Centered Approach Using ISO/IEC 25010:2011 Quality Model. Systems 2026, 14, 1078. https://doi.org/10.3390/systems14091078

AMA Style

Jung HS, Lee H. Evaluating Software Quality in Cryptocurrency Wallet Applications: A User-Centered Approach Using ISO/IEC 25010:2011 Quality Model. Systems. 2026; 14(9):1078. https://doi.org/10.3390/systems14091078

Chicago/Turabian Style

Jung, Hae Sun, and Haein Lee. 2026. "Evaluating Software Quality in Cryptocurrency Wallet Applications: A User-Centered Approach Using ISO/IEC 25010:2011 Quality Model" Systems 14, no. 9: 1078. https://doi.org/10.3390/systems14091078

APA Style

Jung, H. S., & Lee, H. (2026). Evaluating Software Quality in Cryptocurrency Wallet Applications: A User-Centered Approach Using ISO/IEC 25010:2011 Quality Model. Systems, 14(9), 1078. https://doi.org/10.3390/systems14091078

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop