GemSP: An Ensemble Model for User Story Point Estimation Using Gemini Embeddings
Abstract
1. Introduction
- 1.
- As far as we are aware, this study is one of the earliest to apply Gemini embeddings to story point detection.
- 2.
- Our proposed method demonstrates superior performance compared to selected state-of-the-art baselines, including GPT-2, Deep-SE, and GPT2SP, under cross-project evaluation on JIRA datasets
- 3.
- We also present the open research questions that guide our study and provide a detailed analysis of the proposed algorithm. The representations generated by Gemini embeddings can be precomputed and reused across multiple downstream tasks, including regression, classification, similarity search, clustering, ranking, and retrieval. This property enables efficient deployment of the proposed framework, as the embedding extraction step does not need to be repeated during model retraining or inference. Such general-purpose embeddings have been shown to transfer effectively across tasks and domains [20].
2. Literature Survey
3. Proposed Framework
3.1. Step 1: Extracting Gemini Embeddings from User Stories
3.2. Step 2: Predictive Ensemble Learning Model
3.3. Final Prediction
4. Experimental Setup
4.1. Studied Datasets
4.2. Model Implementation
| Algorithm 1: GemSP Algorithm |
Input: represents User Story Text, where , N is the number of user stories, and D is the dimensionality of text features Output: Predicted User Story Points Step 1: Gemini Embeddings Generation to N Gemini API to generate text embeddings for user story Step 2: Apply PCA for Dimensionality Reduction Apply PCA to the embeddings to reduce dimensionality to K. Step 3: Ensemble Random Forest and XGBoost Regressors to N Train Random Forest Regressor on :
Train XGBoost Regressor on :
Step 4: Ensembling Combine the predictions of RF and XGB regressors through voting regressor for final prediction: , the predicted user story points for each input user story text. |
4.3. Hyperparameters
4.4. Metrics
4.5. Handling the Subjectivity of Story Points
5. Experimental Results
6. Discussion
6.1. Interpretation of Cross-Project Evaluation
6.2. RQ1
6.3. RQ2
6.4. RQ3
6.5. RQ4
- 1.
- Scalability: GemSP demonstrates a high degree of scalability, making it particularly suitable for large-scale software projects of varying complexity. Its embedding generation mechanism allows for processing of relevant user stories, eliminating the need for full retraining of models. Indeed as the dataset size increases, we only extract only Gemini representations thereby avoiding the need to retrain the Gemini model. This modular architecture, coupled with efficient regression models, ensures that GemSP can manage increasing data volumes without much computational costs.
- 2.
- Gemini is LLM: The Gemini embeddings, pre-trained on a large-scale corpus, endow GemSP with exceptional contextual and semantic comprehension of textual inputs. This pre-training enables the model to capture deep meanings, subtle interrelationships, and linguistic nuances in user stories, providing a more robust and contextually aware representation compared to traditional feature extraction techniques.
- 3.
- Efficiency in Processing Large Volumes of Data Advantage: GemSP excels in efficiently processing large volumes of textual data, benefiting from the combination of Gemini embeddings and optimized regression models implemented using tools such as Scikit-learn. This automated pipeline accelerates the user story point estimation process, making it ideal for fast-paced development environments where rapid and reliable analysis is crucial.
- 4.
- Accuracy: GemSP achieves superior accuracy in user story point estimation through the integration of Gemini embeddings and regression approaches, further optimized by regression models. Empirical results indicate a 5–7% improvement in performance (0.17 score increase) over models such as GPT2SP, with validation from human evaluators showing a high degree of agreement with expert-labeled story points. The observed performance gains should therefore be interpreted as a combination of improved semantic representations and robust ensemble regression, rather than as evidence of inherent superiority of the regression architecture alone.
7. Limitations
8. Conclusions and Future Work
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Schwaber, K. Scrum development process. In Proceedings of the Business Object Design and Implementation: OOPSLA’95 Workshop Proceedings, Austin, TX, USA, 16 October 1995; Springer: Cham, Switzerland, 1997; pp. 117–134. [Google Scholar]
- Schwaber, K. Agile Project Management with Scrum; Microsoft Press: Redmond, WA, USA, 2004. [Google Scholar]
- dos Santos, C.A.; Bouchard, K.; Minetto Napoleão, B. Automatic user story generation: A comprehensive systematic literature review. Int. J. Data Sci. Anal. 2024, 20, 1–24. [Google Scholar] [CrossRef]
- Coelho, E.; Basu, A. Effort estimation in agile software development using story points. Int. J. Appl. Inf. Syst. 2012, 3, 7–10. [Google Scholar] [CrossRef]
- Thomas, D.; Hunt, A. User Stories Applied: For Agile Software Development; Addison-Wesley Professional: Boston, MA, USA, 2002; pp. 1–350. [Google Scholar]
- Rodríguez Sánchez, E.; Vázquez Santacruz, E.F.; Cervantes Maceda, H. Effort and cost estimation using decision tree techniques and story points in agile software development. Mathematics 2023, 11, 1477. [Google Scholar] [CrossRef]
- Kochbati, T.; Li, S.; Gérard, S.; Mraidha, C. From user stories to models: A machine learning empowered automation. In Proceedings of the MODELSWARD 2022—9th International Conference on Model-Driven Engineering and Software Development; SCITEPRESS-Science and Technology Publications: Setúbal, Portugal, 2021; Volume 1, pp. 28–40. [Google Scholar]
- Gultekin, M.; Kalipsiz, O. Story point-based effort estimation model with machine learning techniques. Int. J. Softw. Eng. Knowl. Eng. 2020, 30, 43–66. [Google Scholar] [CrossRef]
- Alsaadi, B.; Saeedi, K. Data-driven effort estimation techniques of agile user stories: A systematic literature review. Artif. Intell. Rev. 2022, 55, 5485–5516. [Google Scholar] [CrossRef]
- Prasada Rao, C.; Siva Kumar, P.; Rama Sree, S.; Devi, J. An agile effort estimation based on story points using machine learning techniques. In Proceedings of the Second International Conference on Computational Intelligence and Informatics: ICCII 2017; Springer: Singapore, 2018; pp. 209–219. [Google Scholar]
- Sembhoo, A.; Gobin-Rahimbux, B. A SLR on Deep Learning Models Based on Textual Information For Effort Estimation in Scrum; Research Square Platform LLC: Durham, NC, USA, 2023. [Google Scholar] [CrossRef]
- Choetkiertikul, M.; Dam, H.K.; Tran, T.; Pham, T.; Ghose, A.; Menzies, T. A deep learning model for estimating story points. IEEE Trans. Softw. Eng. 2018, 45, 637–656. [Google Scholar] [CrossRef]
- Fu, M.; Tantithamthavorn, C. GPT2SP: A transformer-based agile story point estimation approach. IEEE Trans. Softw. Eng. 2022, 49, 611–625. [Google Scholar] [CrossRef]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, 2–7 June 2019; Long and Short Papers. Volume 1, pp. 4171–4186. [Google Scholar]
- Choi, Y.; Asif, M.A.; Han, Z.; Willes, J.; Krishnan, R.G. Teaching llms how to learn with contextual fine-tuning. arXiv 2025, arXiv:2503.09032. [Google Scholar] [CrossRef]
- Yalçıner, B.; Dinçer, K.; Karaçor, A.G.; Efe, M.Ö. Enhancing Agile Story Point Estimation: Integrating Deep Learning, Machine Learning, and Natural Language Processing with SBERT and Gradient Boosted Trees. Appl. Sci. 2024, 14, 7305. [Google Scholar] [CrossRef]
- Narzary, S.; Brahma, B.; Mahilary, H.; Brahma, M.; Som, B.; Nandi, S. Comparative Study of Zero-Shot Cross-Lingual Transfer for Bodo POS and NER Tagging Using Gemini 2.0 Flash Thinking Experimental Model. arXiv 2025, arXiv:2503.04405. [Google Scholar]
- Lee, G.G.; Latif, E.; Shi, L.; Zhai, X. Gemini Pro Defeated by GPT-4V: Evidence from Education. arXiv 2023, arXiv:2401.08660. [Google Scholar] [CrossRef]
- Rahman, T.; Zhu, Y. Automated user story generation with test case specification using large language model. arXiv 2024, arXiv:2404.01558. [Google Scholar] [CrossRef]
- Lee, J.; Chen, F.; Dua, S.; Cer, D.; Shanbhogue, M.; Naim, I.; Ábrego, G.H.; Li, Z.; Chen, K.; Vera, H.S.; et al. Gemini embedding: Generalizable embeddings from gemini. arXiv 2025, arXiv:2503.07891. [Google Scholar] [CrossRef]
- Porru, S.; Murgia, A.; Demeyer, S.; Marchesi, M.; Tonelli, R. Estimating story points from issue reports. In Proceedings of the 12th International Conference on Predictive Models and Data Analytics in Software Engineering, Ciudad Real, Spain, 9 September 2016; pp. 1–10. [Google Scholar]
- Sánchez, E.R.; Maceda, H.C.; Santacruz, E.V. Software effort estimation for Agile Software Development using a strategy based on K-nearest neighbors algorithm. In Proceedings of the 2022 IEEE Mexican International Conference on Computer Science (ENC), Xalapa, Mexico, 24–26 August 2022; IEEE: Piscataway, NJ, USA, 2022; pp. 1–6. [Google Scholar]
- Zhang, S.; Xing, Z.; Guo, R.; Xu, F.; Chen, L.; Zhang, Z.; Zhang, X.; Feng, Z.; Zhuang, Z. Empowering Agile-Based Generative Software Development through Human-AI Teamwork. ACM Trans. Softw. Eng. Methodol. 2025, 34, 156. [Google Scholar] [CrossRef]
- Almalki, S.S. AI-Driven Decision Support Systems in Agile Software Project Management: Enhancing Risk Mitigation and Resource Allocation. Systems 2025, 13, 208. [Google Scholar] [CrossRef]
- Islam, M.R.; Sandborn, P. Multimodal Generative AI for Story Point Estimation in Software Development. arXiv 2025, arXiv:2505.16290. [Google Scholar] [CrossRef]
- Younas, W.; Chen, R.; Zhao, J.; Iqbal, T.; Sharaf, M.; Imran, A. SPERT: Reinforcement Learning-Enhanced Transformer Model for Agile Story Point Estimation. Int. J. Softw. Eng. Knowl. Eng. 2025, 35, 293–325. [Google Scholar] [CrossRef]
- dos Santos, C.A. Leveraging Text Generation for Enhanced User Story Quality. Ph.D. Thesis, Université du Québec à Chicoutimi, Saguenay, QC, Canada, 2025. [Google Scholar]
- Hallmann, D.; Jacob, K.; Lüttgen, G.; Schmid, U.; von der Weth, R. USeR: A Web-based User Story eReviewer for Assisted Quality Optimizations. arXiv 2025, arXiv:2503.02049. [Google Scholar]
- Marapelli, B.; Carie, A.; Islam, S.M. RNN-CNN model: A bi-directional long short-term memory deep learning network for story point estimation. In Proceedings of the 2020 5th International Conference on Innovative Technologies in Intelligent Systems and Industrial Applications (CITISIA), Sydney, Australia, 25–27 November 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 1–7. [Google Scholar]
- Phan, H.; Jannesari, A. Story point effort estimation by text level graph neural network. arXiv 2022, arXiv:2203.03062. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is All You Need. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Volume 30. [Google Scholar]




| Approach | Text Representation | Prediction Model | Cross-Project | Handling of Subjectivity | Key Limitations |
|---|---|---|---|---|---|
| TF-IDF + ML | Handcrafted features (TF-IDF, metadata) | Classical ML regressors (RF, SVM) | No | None | Limited semantic understanding; weak generalization |
| Deep-SE | Learned word embeddings per project | LSTM + RHWN | No | Implicit (project-specific) | Retraining required; high computational cost |
| BiLSTM/CNN-based | Learned sequence representations | Deep neural networks | Limited | Implicit | Limited transferability; low interpretability |
| TextLevelGNN | Graph-based word representations | GNN classifier | No | None | Classification-based formulation; coarse granularity |
| GPT2SP | GPT-2 (small) embeddings | Transformer-based regressor | Limited | Implicit | Computationally expensive; fixed foundation model |
| SBERT + LightGBM | Sentence-level embeddings (SBERT) | Gradient-boosted trees | Yes | Partial (normalization) | Embedding quality bounded by SBERT |
| GemSP (Ours) | Gemini embeddings | Ensemble regression (RF + XGBoost) | Yes | Explicit (per-project normalization) | Interpretability of ensemble models |
| Repository | Project | #Issues | ||||||
|---|---|---|---|---|---|---|---|---|
| Apache | Mesos | 1680 | 1 | 40 | 3.09 | 3 | 5.87 | 2.42 |
| Usergrid | 482 | 1 | 8 | 2.85 | 3 | 1.97 | 1.4 | |
| Appcelerator | Appcelerator Studio | 2919 | 1 | 40 | 5.64 | 5 | 11.07 | 3.33 |
| Aptana Studio | 829 | 1 | 34 | 8.02 | 8 | 25.97 | 5.1 | |
| Titanium SDK/CLI | 2251 | 1 | 34 | 6.32 | 5 | 25.97 | 5.1 | |
| Dura Space | DuraCloud | 666 | 1 | 6 | 3.12 | 3 | 4.12 | 2.03 |
| Atlassian | Bamboo | 521 | 1 | 20 | 4.2 | 2 | 4.6 | 2.14 |
| Clover | 384 | 1 | 49 | 3.59 | 1 | 42.95 | 6.55 | |
| JIRA Software | 352 | 1 | 20 | 4.43 | 3 | 12.35 | 3.51 | |
| Moodle | Moodle | 1166 | 1 | 155 | 15.34 | 5 | 468.53 | 21.66 |
| Lsstcorp | Data Management | 4667 | 1 | 100 | 9.57 | 4 | 275.71 | 16.61 |
| Mulesoft | Mule | 889 | 1 | 30 | 5.08 | 5 | 12.24 | 3.5 |
| Mule Studio | 732 | 1 | 34 | 6.4 | 5 | 19.94 | 4.46 | |
| Spring | Spring XD | 3526 | 1 | 3.7 | 3.7 | 3 | 10.42 | 3.23 |
| Talendforge | Talend Data Quality | 1381 | 1 | 40 | 5.92 | 5 | 26.96 | 5.19 |
| Talend ESB | 868 | 1 | 13 | 2.16 | 2 | 2.24 | 1.5 | |
| Total | 23,313 | - | - | - | - | - | - |
| Model | MAE ↓ | RMSE ↓ |
|---|---|---|
| GPT-2 | 5.306 | 5.83 |
| Deep-SE | 3.50 | 3.93 |
| GPT2SP | 2.14 | 2.45 |
| GemSP (Ours) | 1.98 | 2.29 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Moufidi, I.; Achour, S.; Benattou, M. GemSP: An Ensemble Model for User Story Point Estimation Using Gemini Embeddings. Information 2026, 17, 110. https://doi.org/10.3390/info17010110
Moufidi I, Achour S, Benattou M. GemSP: An Ensemble Model for User Story Point Estimation Using Gemini Embeddings. Information. 2026; 17(1):110. https://doi.org/10.3390/info17010110
Chicago/Turabian StyleMoufidi, Imad, Safaa Achour, and Mohammed Benattou. 2026. "GemSP: An Ensemble Model for User Story Point Estimation Using Gemini Embeddings" Information 17, no. 1: 110. https://doi.org/10.3390/info17010110
APA StyleMoufidi, I., Achour, S., & Benattou, M. (2026). GemSP: An Ensemble Model for User Story Point Estimation Using Gemini Embeddings. Information, 17(1), 110. https://doi.org/10.3390/info17010110

