Consolidating Access to Candidate Data for Recruitment Headhunting: Leveraging Explainable Machine Learning
Abstract
1. Introduction
- We demonstrate how the Coresignal API consolidates candidate data from LinkedIn, Indeed, and other platforms into a unified interface, eliminating redundant multi-platform searches;
- We compare three regression approaches, Ridge, Gradient Boosting, and Random Forest, for TF-IDF-based candidate–job matching with rigorous statistical validation, showing that regularized linear models achieve superior performance while maintaining interpretability;
- We implement Shapash-based explanations that translate model outputs into intuitive visualizations, making feature importance accessible to users without technical backgrounds.
2. Literature Review
2.1. Data Aggregation in Candidate Sourcing
2.2. Machine Learning in Recruitment
2.3. Explainable AI in Recruitment
2.4. Research Gap and Contributions
- Demonstrating Coresignal API integration for consolidated candidate data access across multiple job portals;
- Conducting rigorous comparative evaluation of Ridge regression, Gradient Boosting, and Random Forest for TF-IDF-based matching with full statistical validation;
- Implementing Shapash-based explainability designed for recruiter accessibility, advancing practical XAI applications in recruitment.
3. Methodology
3.1. Study Design and Data Collection
3.1.1. Selection of Data Source and API Evaluation
3.1.2. Job Description Development
- Python (100% of postings);
- SQL (100%);
- Statistical Analysis (100%);
- Machine Learning (90%);
- Data Visualization (70%);
- Cloud Platforms, specifically AWS and/or Azure (60%).
- Required Skills: Python, SQL, Machine Learning, Statistical Analysis, Cloud Platforms (AWS and/or Azure), Data Visualization;
- Required Education: Master’s degree in Computer Science, Statistics, Mathematics, Data Science, or a related quantitative field;
- Required Experience: 5+ years in a data science or analytical role, with at least 2 years at senior or lead level.
3.1.3. Data Collection
3.2. User Interface Development
3.3. Data Preprocessing and Feature Engineering
3.3.1. Text Processing
- Text extraction: Consolidated all textual fields (skills, experience descriptions, education, certifications) into unified profile texts;
- Lowercasing: Converted all text to lowercase for case-insensitive matching;
- Punctuation removal: Removed special characters while preserving spaces between words;
- Stopwords removal: Filtered common English stopwords (e.g., “the”, “and”, “is”);
- Tokenization: Split texts into individual word tokens.
- Raw text: “Experienced Senior Data Scientist with Python, Machine Learning & AWS”;
- Transformed text: “experienced senior data scientist python machine learning aws”.
3.3.2. TF-IDF Vectorization
- max_features = 230, retaining only the top 230 terms by corpus-wide frequency;
- min_df = 2, excluding terms appearing in fewer than two documents;
- max_df = 0.75, excluding terms appearing in more than 75% of documents to remove near-universal terms;
- ngram_range = (1, 3), capturing unigrams, bigrams, and trigrams to preserve multi-word technical phrases such as “senior data scientist”, “machine learning”, and “data science”.
3.3.3. Similarity Score Calculation
3.3.4. Feature Engineering
3.4. Machine Learning Model Development
- α = 1.0;
- solver = ‘auto’;
- fit_intercept = True;
- max_iter = 10,000;
- random_state = 42.
- n_estimators = 100;
- max_depth = None;
- min_samples_split = 2;
- min_samples_leaf = 1;
- max_features = None;
- random_state = 42.
- n_estimators = 100;
- learning_rate = 0.1;
- max_depth = 3;
- min_samples_split = 2;
- min_samples_leaf = 1;
- subsample = 1.0;
- random_state = 42.
3.4.1. Ridge Regression
3.4.2. Gradient Boosting Regressor
3.4.3. Random Forest Regressor
3.5. Model Evaluation and Statistical Analysis
3.5.1. Baseline Establishment: Cosine Similarity
3.5.2. Cross-Validation Procedure
3.5.3. Performance Evaluation Metrics
3.5.4. Statistical Significance Tests
- R2 scores were recorded for each of the five folds for both models under comparison, and the fold-wise differences in R2 were computed;
- A paired t-test was performed on these differences, with the test statistic defined as , where is the mean of the fold-wise differences and , with SDd denoting the standard deviation of the differences and = 5 the number of folds;
- A Bonferroni correction was applied to account for multiple comparisons across three pairwise model combinations (Ridge vs. Gradient Boosting, Ridge vs. Random Forest, and Gradient Boosting vs. Random Forest), yielding an adjusted significance threshold of ;
- Effect sizes were calculated using Cohen’s d to quantify the practical magnitude of performance differences between models: , where .
3.5.5. Confidence Intervals
- Test predictions were resampled with replacement over 1000 bootstrap iterations;
- Corresponding true similarity scores were retained for each resampled set;
- R2 was calculated for each bootstrap sample;
- The 2.5th and 97.5th percentiles of the resulting bootstrap distribution were taken as the lower and upper confidence interval bounds respectively.
3.6. Explainability Analysis with Shapash
3.6.1. Shapash Implementation
3.6.2. Feature Importance Interpretation
- “machine learning”: +0.076 contribution (increases similarity);
- “data science”: +0.027 contribution (increases similarity);
- Absence of “python”: −0.009 contribution (decreases similarity).
3.6.3. Validation of Feature–Job Alignment
- The top 10 features by global importance were extracted using Shapash;
- These features were compared against the explicit job requirements specified in the standardized job description, including Python, Machine Learning, and master’s degree, among others;
- A feature–job alignment percentage was computed as: Alignment (%) = (Number of matched features ÷ Total required features) × 100.
3.7. Implementation and Software
- Requests 2.32.5 for API integration;
- pandas 1.3.5 and nltk 3.6.7 for data preprocessing;
- scikit-learn 1.0.2 for TF-IDF vectorization and machine learning modelling;
- scipy 1.7.3 for statistical significance testing;
- shapash 2.0.0 for model explainability;
- matplotlib 3.5.1 and seaborn 0.11.2 for visualization.
3.8. Ethical Considerations
4. Results
4.1. Exploratory Data Analysis
4.2. Model Performance Results
4.2.1. Baseline Performance: Cosine Similarity
4.2.2. Model Performance on Held-Out Test Data
4.2.3. Model Cross-Validation Performance
4.2.4. Statistical Significance Testing
4.3. Model Explainability and Feature Importance
Global Feature Importance
4.4. Individual Candidate Explanations
4.5. Model Interpretability: Ridge Coefficients vs. Shapash Visualization
- TF-IDF scale: Coefficients are expressed on the TF-IDF scale, which is not intuitive; it is unclear whether β = 0.169 represents a large or small effect without understanding the underlying value distribution;
- Feature context: Coefficients alone do not indicate how a feature compares to others in terms of its actual impact on individual predictions;
- Individual predictions: Coefficients reflect global average effects and cannot explain why a specific candidate received a particular score.
- Expresses contributions in the predicted similarity score scale: Rather than β values, Shapash reports candidate-specific contributions. For example, “machine learning” contributed +0.076 to Candidate 195’s predicted similarity score;
- Provides global and local views simultaneously: The dashboard’s feature importance panel (top left) mirrors the ranking in Table 6, while the local explanation panel (bottom right) drills into any individual candidate, as shown for Candidate 195;
- Supports non-technical decision-making: The clustering view groups candidates by explainability patterns, enabling recruiters to identify candidate segments without interpreting raw coefficients or SHAP mathematics.
4.6. Summary of Key Findings
- Stratified 5-fold cross-validation and bootstrap confidence intervals confirmed robust, generalizable performance, with Ridge maintaining stable R2 values ranging from 0.914 to 0.954 across all folds and a held-out test RMSE of 0.026, consistent with the cross-validated mean RMSE of 0.025;
- Ridge Regression significantly outperformed both ensemble methods for TF-IDF-based candidate–job matching (cross-validated R2 = 0.935, bootstrap R2 = 0.954, 95% CI [0.939, 0.965] vs. Gradient Boosting R2 = 0.840 and Random Forest R2 = 0.733), with all pairwise differences statistically significant after Bonferroni correction (all ps ≤ 0.001, Cohen’s d ≥ 1.992);
- Global feature importance analysis revealed strong alignment between the model’s top-ranked features and job description requirements (14/15 features, 93%), validating that the model learned meaningful job-relevant relationships rather than spurious correlations;
- Shapash enhances the practical usability of Ridge regression’s inherent interpretability by translating global coefficients into candidate-specific contribution scores, enabling non-technical recruiters to understand not just candidate rankings but the specific skills, qualifications, and experience gaps driving each score.
5. Discussion
5.1. Principal Findings and Theoretical Implications
5.2. Efficiency Gains Through Data Consolidation
5.2.1. Aggregation of Candidate Data
5.2.2. Transparent Algorithmic Decision Support
- Verification: recruiters can confirm that rankings reflect actual job requirements rather than accepting opaque scores;
- Gap identification: for promising candidates with moderate scores, specific skill gaps can be identified and communicated directly;
- Bias detection: if irrelevant features such as geographic tokens or institution names rank highly, this signals problems requiring correction before deployment.
5.2.3. Interpretability Design Rationale
5.3. Comparison with Prior Work
5.4. Limitations
5.4.1. Dataset Scope and Generalizability
5.4.2. Single Job Description
5.4.3. Absence of Recruiter User Validation
- Do recruiters find Shapash visualizations more comprehensible than raw regression coefficients?
- Does the consolidated interface reduce actual time-to-hire in real recruitment workflows?
- How do recruiters integrate algorithmic rankings into holistic hiring decisions?
- Does transparency increase recruiter trust and willingness to use AI recommendations?
5.4.4. Fairness and Bias Considerations
5.4.5. Regulatory Compliance and Data Governance
5.4.6. API Dependency and Cost Considerations
5.4.7. Positioning Relative to Commercial Recruitment Tools
6. Conclusions
- Multi-role and multi-country validation to establish generalizability boundaries beyond the South African Data Scientist context;
- Formal recruiter usability trials measuring time-to-hire reduction, trust calibration, and decision quality when using Shapash-based explanations versus conventional screening;
- Fairness auditing with demographic parity and equal opportunity metrics to identify and mitigate potential bias in candidate rankings.
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A. Synthesized Job Description
References
- Köchling, A.; Wehner, M.C. Discriminated by an algorithm: A systematic review of discrimination and fairness by algorithmic decision-making in the context of HR recruitment and HR development. Bus. Res. 2020, 13, 795–848. [Google Scholar] [CrossRef] [Scilit]
- Burrell, J. How the machine ‘thinks’: Understanding opacity in machine learning algorithms. Big Data Soc. 2016, 3, 2053951715622512. [Google Scholar] [CrossRef] [Scilit]
- Zhaldak, A.; Krasovska, M. Modeling the phased implementation of headhunting as a way to fill vacancies. Technol. Audit. Prod. Reserves 2021, 6, 6–11. [Google Scholar] [CrossRef] [Scilit]
- Bogers, T.; Kaya, M. An Exploration of the Information Seeking Behavior of Recruiters. In Proceedings of the HR@ RecSys, Amsterdam, The Netherlands, 27 September–1 October 2021. [Google Scholar]
- Peicheva, M. Data analysis from the applicant tracking system. Choveshki Resur. Tehnol. HR Technol. Creat. Space Assoc. 2022, 2, 6–15. [Google Scholar]
- Tambe, P.; Cappelli, P.; Yakubovich, V. Artificial intelligence in human resources management: Challenges and a path forward. Calif. Manag. Rev. 2019, 61, 15–42. [Google Scholar] [CrossRef] [Scilit]
- Jirjees, A.K.; Ahmed, A.M.; Abdulla, A.A.; Lu, J.; Noori, E.M.; Kareem, R.N.; Hassan, B.A.; Veisi, H.; Rashid, T.A. Machine Learning for Recruitment: Analysing Job-Matching Algorithm. Doctoral Dissertation, University of Kurdistan Hewlêr, Erbil, Iraq, 2024. [Google Scholar]
- Madanchian, M. From recruitment to retention: AI tools for human resource decision-making. Appl. Sci. 2024, 14, 11750. [Google Scholar] [CrossRef] [Scilit]
- Kostopoulos, G.; Davrazos, G.; Kotsiantis, S. Explainable artificial intelligence-based decision support systems: A recent review. Electronics 2024, 13, 2842. [Google Scholar] [CrossRef] [Scilit]
- Hunkenschroer, A.L.; Kriebitz, A. Is AI recruiting (un) ethical? A human rights perspective on the use of AI for hiring. AI Ethics 2023, 3, 199–213. [Google Scholar] [CrossRef] [Scilit]
- Beretta, A.; Ercoli, G.; Ferraro, A.; Guidotti, R.; Iommi, A.; Mastropietro, A.; Monreale, A.; Rotelli, D.; Ruggieri, S. Requirements of eXplainable AI in Algorithmic Hiring. In Proceedings of the AIMMES 2024 Workshop on AI Bias: Measurements, Mitigation, Explanation Strategies|Co-Located with EU Fairness Cluster Conference 2024, Amsterdam, The Netherlands, 20 March 2024. [Google Scholar]
- Hofeditz, L.; Clausen, S.; Rieß, A.; Mirbabaie, M.; Stieglitz, S. Applying XAI to an AI-based system for candidate management to mitigate bias and discrimination in hiring. Electron. Mark. 2022, 32, 2207–2233. [Google Scholar] [CrossRef] [Scilit]
- Glenny Jocelyn, G.; Judy Grace Nitta, J. The Effectiveness and Challenges of Online Platforms for Talent Sourcing: A Perception Study of Recruiters in the IT Sector. Int. J. Sci. Res. Technol. 2025, 2, IJSRT/250307017. [Google Scholar] [CrossRef]
- Mydyti, H.; Ware, A. Integrating Intelligent Web Scraping Techniques in Internship Management Systems: Enhancing Internship Matching. Ann. Emerg. Technol. Comput. (AETiC) 2025, 9, 1–23. [Google Scholar]
- Kumar, N.; Gupta, M.; Sharma, D.; Ofori, I. Technical job recommendation system using APIs and web crawling. Comput. Intell. Neurosci. 2022, 2022, 7797548. [Google Scholar] [CrossRef] [Scilit]
- Frazzetto, P.; Haq, M.U.U.; Fabris, F.; Sperduti, A. From Text to Talent: A Pipeline for Extracting Insights from Candidate Profiles. arXiv 2025, arXiv:2503.17438. [Google Scholar] [CrossRef] [Scilit]
- Smelyakov, K.; Hurova, Y.; Osiievskyi, S. Analysis of the effectiveness of using machine learning algorithms to make hiring decisions. In Proceedings of the International Conference on Computational Linguistics and Intelligent Systems, Kharkiv, Ukraine, 20–21 April 2023; Volume 1613, p. 73. [Google Scholar]
- Mori, M.; Sassetti, S.; Cavaliere, V.; Bonti, M. A systematic literature review on artificial intelligence in recruiting and selection: A matter of ethics. Pers. Rev. 2025, 54, 854–878. [Google Scholar] [CrossRef] [Scilit]
- Mashayekhi, Y.; Li, N.; Kang, B.; Lijffijt, J.; De Bie, T. A challenge-based survey of e-recruitment recommendation systems. ACM Comput. Surv. 2024, 56, 1–33. [Google Scholar]
- Campion, E.D.; Campion, M.A. Impact of machine learning on personnel selection. Organ. Dyn. 2024, 53, 101035. [Google Scholar] [CrossRef] [Scilit]
- Ajjam, M.-H.; Al-Raweshidy, H.S. AI-driven semantic similarity-based job matching framework for recruitment systems. Inf. Sci. 2025, 724, 122728. [Google Scholar]
- Datto, S.; Ahmed, M.; Emran, A.; Redwan, K.; Nadim, R.; Mahmood, S. Enhancing Candidate Selection with NLP-Driven Resume Analysis for Industry 4.0 Recruitment Systems. In International Conference on Data Science, AI and Applications; Springer: Cham, Switzerland, 2025; pp. 46–60. [Google Scholar]
- Deshmukh, A.; Raut, A. Applying bert-based nlp for automated resume screening and candidate ranking. Ann. Data Sci. 2025, 12, 591–603. [Google Scholar] [CrossRef] [Scilit]
- Pessach, D.; Singer, G.; Avrahami, D.; Ben-Gal, H.C.; Shmueli, E.; Ben-Gal, I. Employees recruitment: A prescriptive analytics approach via machine learning and mathematical programming. Decis. Support Syst. 2020, 134, 113290. [Google Scholar] [CrossRef] [Scilit]
- Larsson, J.; Wallin, J. The Choice of Normalization Influences Shrinkage in Regularized Regression. arXiv 2025, arXiv:2501.03821. [Google Scholar] [CrossRef] [Scilit]
- Friedman, J.H. Greedy function approximation: A gradient boosting machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef] [Scilit]
- Liaw, A.; Wiener, M. Classification and regression by randomForest. R News 2002, 2, 18–22. [Google Scholar]
- Albaroudi, E.; Mansouri, T.; Alameer, A. A comprehensive review of AI techniques for addressing algorithmic bias in job hiring. AI 2024, 5, 383–404. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z. Ethics and discrimination in artificial intelligence-enabled recruitment practices. Humanit. Soc. Sci. Commun. 2023, 10, 567. [Google Scholar] [CrossRef] [Scilit]
- Smuha, N.A. Regulation 2024/1689 of the Eur. Parl. & Council of June 13, 2024 (eu artificial intelligence act). Int. Leg. Mater. 2025, 64, 1234–1381. [Google Scholar] [CrossRef] [Scilit]
- Buçinca, Z.; Swaroop, S.; Paluch, A.E.; Doshi-Velez, F.; Gajos, K.Z. Contrastive explanations that anticipate human misconceptions can improve human decision-making skills. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, 2025; pp. 1–25. [Google Scholar]
- Gómez-Talal, I.; Azizsoltani, M.; Bote-Curiel, L.; Rojo-Álvarez, J.L.; Singh, A. Towards Explainable Artificial Intelligence in Machine Learning: A study on efficient Perturbation-Based Explanations. Eng. Appl. Artif. Intell. 2025, 155, 110664. [Google Scholar] [CrossRef] [Scilit]
- Ribeiro, M.T.; Singh, S.; Guestrin, C. “Why should i trust you?” Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; Association for Computing Machinery: New York, NY, USA, 2016; pp. 1135–1144. [Google Scholar]
- Qin, L.; Zhu, Y.; Liu, S.; Zhang, X.; Zhao, Y. The Shapley Value in Data Science: Advances in Computation, Extensions, and Applications. Mathematics 2025, 13, 1581. [Google Scholar] [CrossRef] [Scilit]
- Shapley, L.S. A Value for n-Person Games; Princeton University Press: Princeton, NJ, USA, 1953. [Google Scholar]
- Kaur, H.; Nori, H.; Jenkins, S.; Caruana, R.; Wallach, H.; Wortman Vaughan, J. Interpreting interpretability: Understanding data scientists’ use of interpretability tools for machine learning. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, 2020; pp. 1–14. [Google Scholar]
- Miller, T. Explanation in artificial intelligence: Insights from the social sciences. Artif. Intell. 2019, 267, 1–38. [Google Scholar] [CrossRef] [Scilit]
- MAIF. Shapash: Making Machine Learning Interpretable. Available online: https://github.com/MAIF/shapash (accessed on 20 August 2025).
- Spinner, T.; Schlegel, U.; Schäfer, H.; El-Assady, M. explAIner: A visual analytics framework for interactive and explainable machine learning. IEEE Trans. Vis. Comput. Graph. 2019, 26, 1064–1074. [Google Scholar] [CrossRef] [Scilit]
- Liao, Q.V.; Gruen, D.; Miller, S. Questioning the AI: Informing design practices for explainable AI user experiences. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, 2020; pp. 1–15. [Google Scholar]
- Ravi, K.M. Mitigating bias in AI-driven recruitment: The role of explainable machine learning (XAI). Int. J. 2024, 10, 461–469. [Google Scholar]
- Delecraz, S.; Eltarr, L.; Oullier, O. Transparency and explainability of a machine learning model in the context of human resource management. In Proceedings of the Workshop on Ethical and Legal Issues in Human Language Technologies and Multilingual De-Identification of Sensitive Data in Language Resources Within the 13th Language Resources and Evaluation Conference; European Language Resources Association: Marseille, France, 2022; pp. 38–43. [Google Scholar]
- Fabris, A.; Baranowska, N.; Dennis, M.J.; Graus, D.; Hacker, P.; Saldivar, J.; Zuiderveen Borgesius, F.; Biega, A.J. Fairness and bias in algorithmic hiring: A multidisciplinary survey. ACM Trans. Intell. Syst. Technol. 2025, 16, 1–54. [Google Scholar] [CrossRef] [Scilit]
- Alsubaie, N.; Aleisa, N. Mitigating bias in AI model using explainable AI in terms of hiring process in the industry. IEEE Access 2025, 13, 147218–147241. [Google Scholar] [CrossRef] [Scilit]
- Salton, G.; Buckley, C. Approaches to Text Retreival for Structured Documents; Cornell University: Ithaca, NY, USA, 1990. [Google Scholar]
- Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
- Ngwenya, B.; Paepae, T.; Bokoro, P.N. Advancing SDG 6.3. 2 with machine learning-based virtual sensors for high-frequency nutrient monitoring. J. Water Process Eng. 2025, 79, 108831. [Google Scholar] [CrossRef] [Scilit]
- Ngwenya, B.; Paepae, T.; Bokoro, P.N. Monitoring ambient water quality using machine learning and IoT: A review and recommendations for advancing SDG indicator 6.3.2. J. Water Process Eng. 2025, 73, 107664. [Google Scholar] [CrossRef] [Scilit]
- Paepae, T.; Bokoro, P.N.; Kyamakya, K. From fully physical to virtual sensing for water quality assessment: A comprehensive review of the relevant state-of-the-art. Sensors 2021, 21, 6971. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Paepae, T.; Bokoro, P.N.; Kyamakya, K. A virtual sensing concept for Nitrogen and Phosphorus monitoring using machine learning techniques. Sensors 2022, 22, 7338. [Google Scholar] [CrossRef] [Scilit]
- Paepae, T.; Bokoro, P.N.; Kyamakya, K. Data augmentation for a virtual-sensor-based nitrogen and phosphorus monitoring. Sensors 2023, 23, 1061. [Google Scholar] [CrossRef] [Scilit]
- Mokgwatjane, K.; Paepae, T. An explainable ensemble machine learning approach for multi-domain, multiclass sentiment analysis in Amazon product reviews. Mach. Learn. Appl. 2026, 23, 100825. [Google Scholar] [CrossRef] [Scilit]
- Wilimitis, D.; Walsh, C.G. Practical considerations and applied examples of cross-validation for model development and evaluation in health care: Tutorial. JMIR AI 2023, 2, e49023. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lumumba, V.W.; Kiprotich, D.; Lemasulani Mpaine, M.; Grace Makena, N.; Daniel Kavita, M. Comparative analysis of cross-validation techniques: LOOCV, K-folds cross-validation, and repeated K-folds cross-validation in machine learning models. Am. J. Theor. Appl. Stat. 2024, 13, 127–137. [Google Scholar]
- Hodson, T.O. Root mean square error (RMSE) or mean absolute error (MAE): When to use them or not. Geosci. Model Dev. Discuss. 2022, 15, 5481–5487. [Google Scholar] [CrossRef] [Scilit]
- Demšar, J. Statistical comparisons of classifiers over multiple data sets. J. Mach. Learn. Res. 2006, 7, 1–30. [Google Scholar]
- Cohen, J. Set correlation and contingency tables. Appl. Psychol. Meas. 1988, 12, 425–434. [Google Scholar] [CrossRef] [Scilit]
- Rudin, C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat. Mach. Intell. 2019, 1, 206–215. [Google Scholar] [CrossRef] [Scilit]
- Republic of South Africa. Protection of Personal Information Act, No. 4 of 2013. Gov. Gaz. 2013, 581, 1–148.

















| Study | Year | Data Source | ML Method | XAI Technique | Non-Technical User Explainability | Multi-Portal Aggregation | Key Limitation |
|---|---|---|---|---|---|---|---|
| Ajjam & Al-Raweshidy [21] | 2025 | Simulated & real-world datasets | TF-IDF & Cosine Similarity | None | No | No | No multi-platform sourcing; no explainability |
| Alsubaie & Aleisa [44] | 2025 | Single Platform | RF, XGBoost, LightGBM | SHAP | Partial | No | No multi-platform sourcing; no recruiter-facing explainability |
| Delecraz et al. [42] | 2022 | Single proprietary database | XG Boost | SHAP | Partial | No | Single proprietary database; no recruiter-facing explainability |
| Deshmukh & Raut [23] | 2025 | Single Platform | BERT-based NLP | None | No | No | Opacity of deep learning; no XAI |
| Datto et al. [22] | 2025 | Single Platform | TF-IDF & Cosine Similarity | None | No | No | No explainability; no multi-platform sourcing |
| Frazzetto et al. [16] | 2025 | Single-source CVs | LLM & Graph Neural Networks | None | No | No | Single-source input; no multi-portal aggregation |
| Hofeditz et al. [12] | 2022 | Controlled experiment | Wizard-of-Oz AI recommendation system | Text explanation, not a formal XAI method | Yes (user study) | No | Simulated AI; no real ML model |
| Jirjees et al. [7] | 2024 | Synthetic dataset from Kaggle | SVM, LSTM, MLP, BERT Classification | None | No | No | Synthetic dataset; high accuracy but no explainability |
| Kumar et al. [15] | 2022 | Multi-source job postings | Content-based & Collaborative Filtering | None | No | No | Web scraping ethical & data quality concerns; no candidate data aggregation |
| Magham [41] | 2024 | No dataset used | None | SHAP & LIME | No | No | No empirical validation; no recruiter-facing explainability |
| Smelyakov et al. [17] | 2023 | Single dataset from Kaggle | Decision Tree, Random Forest, Gradient Boosting | None | No | No | No explainability; single-source data |
| This Study | 2026 | Coresignal API (multi-platform) | Ridge Regression, Gradient Boosting, Random Forest | Shapash | Yes | Yes | Single occupation and geographic market; no formal recruiter user validation |
| Feature | Description |
|---|---|
| Headline/Title | Professional title or headline of a candidate. |
| Industry | Industry or field the candidate works in. |
| Roles | Job positions held by a candidate in their professional experience. |
| Companies worked at | Companies where the candidate has worked. |
| Total years’ experience | Total duration of professional experience in years. |
| Educational institutions | Institutions where a candidate studied. |
| Qualifications | Candidate’s academic qualifications. |
| Skills | Candidate’s technical or professional skills. |
| Certifications | Professional certifications held by a candidate. |
| Candidate Profile | Total Years of Experience | Similarity Score |
|---|---|---|
| data scientist standard bank group data analyst automation specialist university venda bachelor science computer information python programming microsoft certified data analyst associate | 3 | 0.606362 |
| senior specialist data science mtn clover south africa master statistics microsoft certified azure fundamental intro python data science | 6 | 0.754634 |
| Model | R2 Mean (SD) | Bootstrap R2 95% CI | RMSE Mean (SD) |
|---|---|---|---|
| Ridge Regression | 0.935 (0.018) | [0.939, 0.965] | 0.025 (0.003) |
| Gradient Boosting | 0.840 (0.038) | [0.846, 0.916] | 0.039 (0.003) |
| Random Forest | 0.733 (0.065) | [0.746, 0.845] | 0.050 (0.004) |
| Comparison | t(4) | p-Value | Significant | Cohen’s d | Effect Size |
|---|---|---|---|---|---|
| Ridge vs. Gradient Boosting | 9.599 | 0.0007 | Yes | 3.165 | Large |
| Ridge vs. Random Forest | 9.073 | 0.0008 | Yes | 4.205 | Large |
| Gradient Boosting vs. Random Forest | 8.569 | 0.001 | Yes | 1.992 | Large |
| Rank | Feature | Importance Score | Job Description Alignment |
|---|---|---|---|
| 1 | engineering | 0.057 | Yes (Required Field) |
| 2 | data science | 0.057 | Yes (Required Field) |
| 3 | machine learning | 0.036 | Yes (Required Skill) |
| 4 | python | 0.030 | Yes (Required Skill) |
| 5 | statistical analysis | 0.024 | Yes (Required Skill) |
| 6 | computer science | 0.022 | Yes (Required Field) |
| 7 | mathematics | 0.017 | Yes (Required Field) |
| 8 | lead | 0.015 | Yes (Experience Level) |
| 9 | senior data | 0.014 | Yes (Experience Level) |
| 10 | aws | 0.014 | Yes (Experience Skill) |
| 11 | cloud | 0.014 | Yes (Required Skill) |
| 12 | business | 0.014 | Partial |
| 13 | master degree | 0.013 | Yes (Required Education) |
| 14 | bachelor degree | 0.011 | Yes (Minimum Education Level) |
| 15 | engineer | 0.011 | Yes (Required Field) |
| Candidate | Score | Score Range | Top Positive Contributors | Top Negative Contributors |
|---|---|---|---|---|
| 148 | 0.420 | Very High (0.35+) | master (+0.0170), data science (+0.0161), aws (+0.0135), cloud (+0.0103), analysis (+0.0092) | engineering (−0.0088), management (−0.0082), learning (−0.0079), machine (−0.0059), machine learning (−0.0059) |
| 91 | 0.260 | High (0.25–0.35) | statistic (+0.0368), statistical (+0.0357), mathematical (+0.0066), design (+0.0062), senior (+0.0047) | data science (−0.0117), engineering (−0.0088), learning (−0.0079), machine (−0.0059), machine learning (−0.0059) |
| 98 | 0.234 | Medium-High (0.20–0.25) | intern (+0.0029), academy (+0.0012), management (+0.0009), information (+0.0008), data analyst (+0.0008) | data science (−0.0117), exploreai (−0.0099), engineering (−0.0088), learning (−0.0079), senior (−0.0073) |
| 78 | 0.109 | Medium (0.10–0.20) | statistical (+0.0336), analysis (+0.0291), quantitative (+0.0207), senior (+0.0154), statistic (+0.0115) | engineering (−0.0088), learning (−0.0079), machine (−0.0059), machine learning (−0.0059), degree (−0.0051) |
| 83 | 0.092 | Low (0.00–0.10) | engineering (+0.0360), mathematics (+0.0061), industrial (+0.0048), trainee (+0.0033), tutor (+0.0016) | data science (−0.0117), learning (−0.0079), senior (−0.0073), machine (−0.0059), machine learning (−0.0059) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Mncwabe, M.; Paepae, T. Consolidating Access to Candidate Data for Recruitment Headhunting: Leveraging Explainable Machine Learning. Informatics 2026, 13, 94. https://doi.org/10.3390/informatics13060094
Mncwabe M, Paepae T. Consolidating Access to Candidate Data for Recruitment Headhunting: Leveraging Explainable Machine Learning. Informatics. 2026; 13(6):94. https://doi.org/10.3390/informatics13060094
Chicago/Turabian StyleMncwabe, Mncedisi, and Thulane Paepae. 2026. "Consolidating Access to Candidate Data for Recruitment Headhunting: Leveraging Explainable Machine Learning" Informatics 13, no. 6: 94. https://doi.org/10.3390/informatics13060094
APA StyleMncwabe, M., & Paepae, T. (2026). Consolidating Access to Candidate Data for Recruitment Headhunting: Leveraging Explainable Machine Learning. Informatics, 13(6), 94. https://doi.org/10.3390/informatics13060094

