Figure 1.
Methodological workflow for knowledge base construction and evaluation.
Figure 2.
Workflow of the chunking module: section-aware chunking with character-based fallback. Red boxes indicate chunks mined.
Figure 3.
Workflow of the field-extraction module: title, skills, and experience detection through keyword matching, evidence selection, and rule-based validation. Red boxes indicate chunks mined.
Figure 4.
Evaluation workflow including system predictions, dual-annotator labeling, ground-truth (GT) consolidation, and system evaluation. Sample sizes were n = 76 CVs for Queries 1–2 and n = 350 CVs for Query 3. Green cells denote valid CVs, and red cells denote invalid CVs.
Figure 5.
Example of correct requirement matching (title/degree, skills, and experience) and agreement with human labeling in Query 3. Red boxes indicate sections where relevant information appears. Highlighted text shows matches identified by both the system and the ground truth (GT). Checkmarks denote correct extraction.
Figure 6.
Example of mismatched requirement detection in Query 1. Red boxes indicate sections where relevant information typically appears. Highlighted text shows matches identified by both the system and the ground truth (GT). Checkmarks denote correct extraction, and crosses denote incorrect extraction.
Figure 7.
Performance metrics across the seven evaluation scenarios for Query 3.
Figure 8.
Performance metrics across the seven evaluation scenarios for Query 2.
Figure 9.
Performance metrics across the seven evaluation scenarios for Query 1.
Table 1.
Illustrative confusion matrix for binary classification.
| | Predicted: 1 | Predicted: 0 |
|---|
| Actual: 1 | True Positives | False Negatives |
| Actual: 0 | False Positives | True Negatives |
Table 2.
Definition of evaluation queries applied.
| Query | Degree | Skills * | Experience (Years) |
|---|
| 1 | “Software Developer” | Java, Python, Leadership | 5 |
| 2 | “Electrical Engineer” | Leadership, Communication, AutoCAD | 4 |
| 3 | “Mechanical Engineer” | Communication, Office, Creativity | 3 |
Table 3.
Binary representation of evaluation scenarios for candidate selection.
| Scenario | Requirement 1 Degree | Requirement 2 Skills | Requirement 3 Experience |
|---|
| 1 (D) | 1 | 0 | 0 |
| 2 (S) | 0 | 1 | 0 |
| 3 (E) | 0 | 0 | 1 |
| 4 (D + S) | 1 | 1 | 0 |
| 5 (D + E) | 1 | 0 | 1 |
| 6 (S + E) | 0 | 1 | 1 |
| 7 (D + S + E) | 1 | 1 | 1 |
Table 4.
Confusion-matrix counts for Query 3 across evaluation scenarios.
| Scenario (D + S + E) * | True Positives | True Negatives | False Positives | False Negatives |
|---|
| 1 (D) | 23 | 312 | 0 | 15 |
| 2 (S) | 224 | 107 | 7 | 12 |
| 3 (E) | 157 | 146 | 18 | 29 |
| 4 (D + S) | 17 | 323 | 1 | 9 |
| 5 (D + E) | 4 | 333 | 6 | 7 |
| 6 (S + E) | 105 | 210 | 17 | 18 |
| 7 (D + S + E) | 4 | 339 | 4 | 3 |
Table 5.
Accuracy (%), specificity (%), sensitivity/recall (%), precision (%) and error rate (%) for Query 3 derived from the confusion matrix.
| Scenario (D + S + E) * | Accuracy | Specificity | Sensitivity | Precision | Error Rate |
|---|
| 1 (D) | 95.71% | 100.00% | 60.53% | 100.00% | 4.29% |
| 2 (S) | 94.57% | 93.86% | 94.92% | 96.97% | 5.43% |
| 3 (E) | 86.57% | 89.02% | 84.41% | 89.71% | 13.43% |
| 4 (D + S) | 97.14% | 99.69% | 65.38% | 94.44% | 2.86% |
| 5 (D + E) | 96.29% | 98.23% | 36.36% | 40.00% | 3.71% |
| 6 (S + E) | 90.00% | 92.51% | 85.37% | 86.07% | 10.00% |
| 7 (D + S + E) | 98.00% | 98.83% | 57.14% | 50.00% | 2.00% |
Table 6.
Confusion-matrix counts for Query 2 across evaluation scenarios.
| Scenario (D + S + E) * | True Positives | True Negatives | False Positives | False Negatives |
|---|
| 1 (D) | 5 | 59 | 0 | 12 |
| 2 (S) | 27 | 27 | 1 | 21 |
| 3 (E) | 25 | 26 | 4 | 21 |
| 4 (D + S) | 3 | 63 | 0 | 10 |
| 5 (D + E) | 2 | 63 | 0 | 11 |
| 6 (S + E) | 9 | 46 | 1 | 20 |
| 7 (D + S + E) | 1 | 66 | 0 | 9 |
Table 7.
Accuracy (%), specificity (%), sensitivity/recall (%), precision (%) and error rate (%) for Query 2 derived from the confusion matrix.
| Scenario (D + S + E) * | Accuracy | Specificity | Sensitivity | Precision | Error Rate |
|---|
| 1 (D) | 84.21% | 100.00% | 29.41% | 100.00% | 15.79% |
| 2 (S) | 71.05% | 96.43% | 56.25% | 96.43% | 28.95% |
| 3 (E) | 67.11% | 86.67% | 54.35% | 86.21% | 32.89% |
| 4 (D + S) | 86.84% | 100.00% | 23.08% | 100.00% | 13.16% |
| 5 (D + E) | 85.53% | 100.00% | 15.38% | 100.00% | 14.47% |
| 6 (S + E) | 72.37% | 97.87% | 31.03% | 90.00% | 27.63% |
| 7 (D + S + E) | 88.16% | 100.00% | 10.00% | 100.00% | 11.84% |
Table 8.
Confusion-matrix counts for Query 1 across evaluation scenarios.
| Scenario (D + S + E) * | True Positives | True Negatives | False Positives | False Negatives |
|---|
| 1 (D) | 2 | 63 | 2 | 9 |
| 2 (S) | 15 | 48 | 3 | 10 |
| 3 (E) | 21 | 29 | 3 | 23 |
| 4 (D + S) | 2 | 67 | 0 | 7 |
| 5 (D + E) | 0 | 68 | 1 | 7 |
| 6 (S + E) | 2 | 59 | 1 | 14 |
| 7 (D + S + E) | 1 | 71 | 0 | 4 |
Table 9.
Accuracy (%), specificity (%), sensitivity/recall (%), precision (%) and error rate (%) for Query 1 derived from the confusion matrix.
| Scenario (D + S + E) * | Accuracy | Specificity | Sensitivity | Precision | Error Rate |
|---|
| 1 (D) | 85.53% | 96.92% | 18.18% | 50.00% | 14.47% |
| 2 (S) | 82.89% | 94.12% | 60.00% | 83.33% | 17.11% |
| 3 (E) | 65.79% | 90.63% | 47.73% | 87.50% | 34.21% |
| 4 (D + S) | 90.79% | 100.00% | 22.22% | 100.00% | 9.21% |
| 5 (D + E) | 89.47% | 98.55% | 0.00% | 0.00% | 10.53% |
| 6 (S + E) | 80.26% | 98.33% | 12.50% | 66.67% | 19.74% |
| 7 (D + S + E) | 94.74% | 100.00% | 20.00% | 100.00% | 5.26% |
Table 10.
Accuracy (%) per query across the seven evaluation scenarios.
| Scenario (D + S + E) * | Query 1 | Query 2 | Query 3 | Average |
|---|
| 1 (D) | 85.53% | 84.21% | 95.71% | 88.48% |
| 2 (S) | 82.89% | 71.05% | 94.57% | 82.84% |
| 3 (E) | 65.79% | 67.11% | 86.57% | 73.16% |
| 4 (D + S) | 90.79% | 86.64% | 97.14% | 91.59% |
| 5 (D + E) | 89.47% | 85.53% | 96.29% | 90.43% |
| 6 (S + E) | 80.26% | 72.37% | 90.00% | 80.88% |
| 7 (D + S + E) | 94.74% | 88.16% | 98.00% | 93.63% |
Table 11.
Sensitivity (%) per query across the seven evaluation scenarios.
| Scenario (D + S + E) * | Query 1 | Query 2 | Query 3 | Average |
|---|
| 1 (D) | 18.18% | 29.41% | 60.53% | 36.04% |
| 2 (S) | 60.00% | 56.25% | 94.92% | 70.39% |
| 3 (E) | 47.73% | 54.35% | 84.41% | 62.16% |
| 4 (D + S) | 22.22% | 23.08% | 65.38% | 36.89% |
| 5 (D + E) | 0.00% | 15.28% | 36.36% | 17.25% |
| 6 (S + E) | 12.59% | 31.03% | 85.37% | 42.97% |
| 7 (D + S + E) | 20.00% | 10.00% | 57.14% | 29.05% |
Table 12.
Specificity (%) per query across the seven evaluation scenarios.
| Scenario (D + S + E) * | Query 1 | Query 2 | Query 3 | Average |
|---|
| 1 (D) | 96.92% | 100.00% | 100.00% | 98.97% |
| 2 (S) | 94.12% | 96.43% | 93.86% | 94.80% |
| 3 (E) | 90.63% | 86.67% | 89.02% | 88.77% |
| 4 (D + S) | 100.00% | 100.00% | 99.69% | 99.90% |
| 5 (D + E) | 98.55% | 100.00% | 98.23% | 98.93% |
| 6 (S + E) | 98.33% | 97.87% | 92.51% | 96.24% |
| 7 (D + S + E) | 100.00% | 100.00% | 98.83% | 99.61% |
Table 13.
Confusion matrix for Query 3 under Scenario 3.
| | Predicted: 1 | Predicted: 0 |
|---|
| Actual: 1 | 157 | 29 |
| Actual: 0 | 18 | 146 |
Table 14.
Derived metrics from confusion matrix for Query 3 under Scenario 3.
| Metric | Value |
|---|
| Accuracy | 86.57% |
| Specificity | 89.02% |
| Sensitivity | 84.41% |
| Precision | 89.71% |
| Error rate | 13.43% |
Table 15.
Confusion matrix for Query 3 under Scenario 7.
| | Predicted: 1 | Predicted: 0 |
|---|
| Actual: 1 | 4 | 3 |
| Actual: 0 | 4 | 339 |
Table 16.
Derived metrics from confusion matrix for Query 3 under Scenario 7.
| Metric | Value |
|---|
| Accuracy | 98.00% |
| Specificity | 98.83% |
| Sensitivity | 57.14% |
| Precision | 50.00% |
| Error rate | 2.00% |
Table 17.
Confusion matrix for Query 2 under Scenario 3.
| | Predicted: 1 | Predicted: 0 |
|---|
| Actual: 1 | 25 | 21 |
| Actual: 0 | 4 | 26 |
Table 18.
Derived metrics from confusion matrix for Query 2 under Scenario 3.
| Metric | Value |
|---|
| Accuracy | 67.11% |
| Specificity | 86.67% |
| Sensitivity | 54.35% |
| Precision | 86.21% |
| Error rate | 32.89% |
Table 19.
Confusion matrix for Query 2 under Scenario 7.
| | Predicted: 1 | Predicted: 0 |
|---|
| Actual: 1 | 1 | 9 |
| Actual: 0 | 0 | 66 |
Table 20.
Derived metrics from confusion matrix for Query 2 under Scenario 7.
| Metric | Value |
|---|
| Accuracy | 88.16% |
| Specificity | 100.00% |
| Sensitivity | 10.00% |
| Precision | 100.00% |
| Error rate | 11.84% |
Table 21.
Confusion matrix for Query 1 under Scenario 3.
| | Predicted: 1 | Predicted: 0 |
|---|
| Actual: 1 | 21 | 23 |
| Actual: 0 | 3 | 29 |
Table 22.
Derived metrics from confusion matrix for Query 1 under Scenario 3.
| Metric | Value |
|---|
| Accuracy | 65.79% |
| Specificity | 90.63% |
| Sensitivity | 47.73% |
| Precision | 87.50% |
| Error rate | 34.21% |
Table 23.
Confusion matrix for Query 1 under Scenario 7.
| | Predicted: 1 | Predicted: 0 |
|---|
| Actual: 1 | 1 | 4 |
| Actual: 0 | 0 | 71 |
Table 24.
Derived metrics from confusion matrix for Query 1 under Scenario 7.
| Metric | Value |
|---|
| Accuracy | 94.74% |
| Specificity | 100.00% |
| Sensitivity | 20.00% |
| Precision | 100.00% |
| Error rate | 5.26% |
Table 25.
Query 1 Friedman test.
| Statistic | Value |
|---|
| N | 76 |
| χ2 (df = 6) | 42.27 |
Table 26.
Query 1 mean ranks (Friedman test).
| Scenario (D + S + E) * | Scenario | Mean Rank |
|---|
| 7 (D + S + E) | 7 | 4.37 |
| 4 (D + S) | 4 | 4.23 |
| 5 (D + E) | 5 | 4.18 |
| 1 (D) | 1 | 4.05 |
| 2 (S) | 2 | 3.95 |
| 6 (S + E) | 6 | 3.86 |
| 3 (E) | 3 | 3.36 |
Table 27.
Query 1 Cochran’s Q.
| Statistic | Value |
|---|
| N | 76 |
| Q (df = 6) | 42.271 |
| Asymptotic Sig. | <0.001 |
Table 28.
Query 1 Cochran’s Q frequencies.
| Scenario (D + S + E) * | Success (1) | Failure (0) |
|---|
| 1 (D) | 65 | 11 |
| 2 (S) | 63 | 13 |
| 3 (E) | 50 | 26 |
| 4 (D + S) | 69 | 7 |
| 5 (D + E) | 68 | 8 |
| 6 (S + E) | 61 | 15 |
| 7 (D + S + E) | 72 | 4 |
Table 29.
Query 1 Wilcoxon comparisons (vs. Scenario 7).
Scenarios Comparison | Z | p-Value |
|---|
| 1 vs. 7 | −2.65 | 0.008 |
| 2 vs. 7 | −2.50 | 0.013 |
| 3 vs. 7 | −4.31 | <0.001 |
| 4 vs. 7 | −1.73 | 0.083 |
| 5 vs. 7 | −2.00 | 0.046 |
| 6 vs. 7 | −3.32 | 0.001 |
Table 30.
Query 2 Friedman test.
| Statistic | Value |
|---|
| N | 76 |
| χ2 (df = 6) | 28.95 |
Table 31.
Query 2 mean ranks (Friedman test).
| Scenario (D + S + E) * | Scenario | Mean Rank |
|---|
| 7 (D + S + E) | 7 | 4.31 |
| 4 (D + S) | 4 | 4.26 |
| 5 (D + E) | 5 | 4.22 |
| 1 (D) | 1 | 4.17 |
| 6 (S + E) | 2 | 3.76 |
| 2 (S) | 6 | 3.71 |
| 3 (E) | 3 | 3.57 |
Table 32.
Query 2 Cochran’s Q.
| Statistic | Value |
|---|
| N | 76 |
| Q (df = 6) | 28.948 |
| Asymptotic Sig. | <0.001 |
Table 33.
Query 2 Cochran’s Q frequencies.
| Scenario (D + S + E) * | Scenario | Success (1) | Failure (0) |
|---|
| 1 (D) | 1 | 64 | 12 |
| 2 (S) | 2 | 54 | 22 |
| 3 (E) | 3 | 51 | 25 |
| 4 (D + S) | 4 | 66 | 10 |
| 5 (D + E) | 5 | 65 | 11 |
| 6 (S + E) | 6 | 55 | 21 |
| 7 (D + S + E) | 7 | 67 | 9 |
Table 34.
Query 2 Wilcoxon comparisons (vs. Scenario 7).
Scenarios Comparison | Z | p-Value |
|---|
| 1 vs. 7 | −1.13 | 0.257 |
| 2 vs. 7 | −2.84 | 0.005 |
| 3 vs. 7 | −3.02 | 0.002 |
| 4 vs. 7 | −0.58 | 0.564 |
| 5 vs. 7 | −1.41 | 0.157 |
| 6 vs. 7 | −2.83 | 0.005 |
Table 35.
Query 3 Friedman test.
| Statistic | Value |
|---|
| N | 350 |
| χ2 (df = 6) | 83.28 |
Table 36.
Query 3 mean ranks (Friedman test).
| Scenario (D + S + E) * | Mean Rank |
|---|
| 7 (D + S + E) | 4.14 |
| 4 (D + S) | 4.11 |
| 5 (D + E) | 4.08 |
| 1 (D) | 4.06 |
| 2 (S) | 4.02 |
| 6 (S + E) | 3.86 |
| 3 (E) | 3.74 |
Table 37.
Query 3 Cochran’s Q.
| Statistic | Value |
|---|
| N | 350 |
| Q (df = 6) | 83.282 |
| Asymptotic Sig. | <0.001 |
Table 38.
Query 3 Cochran’s Q frequencies.
| Scenario (D + S + E) * | Success (1) | Failure (0) |
|---|
| 1 (D) | 335 | 15 |
| 2 (S) | 331 | 19 |
| 3 (E) | 303 | 47 |
| 4 (D + S) | 340 | 10 |
| 5 (D + E) | 337 | 13 |
| 6 (S + E) | 315 | 35 |
| 7 (D + S + E) | 343 | 7 |
Table 39.
Query 3 Wilcoxon comparisons (vs. Scenario 7).
Scenarios Comparison (D + S + E) * | Z | p-Value |
|---|
| 1 vs. 7 | −2.00 | 0.046 |
| 2 vs. 7 | −2.45 | 0.014 |
| 3 vs. 7 | −5.90 | <0.001 |
| 4 vs. 7 | −1.00 | 0.317 |
| 5 vs. 7 | −2.45 | 0.014 |
| 6 vs. 7 | −4.80 | <0.001 |
Table 40.
Summary of statistical results of the Friedman test and the Cochran Q test across queries.
| Query | n | Friedman χ2 (df = 6) | p-Value | Cochran Q (df = 6) | p-Value | Highest Scenario | Lowest Scenario |
|---|
| Query 1 | 76 | 42.27 | <0.001 | 42.271 | <0.001 | S7 | S3 |
| Query 2 | 76 | 28.95 | <0.001 | 28.948 | <0.001 | S7 | S3 |
| Query 3 | 350 | 83.28 | <0.001 | 83.282 | <0.001 | S7 | S3 |
Table 41.
Summary of statistical results of Wilcoxon test across queries.
Comparison vs. Scenario 7 (D + S + E) * | p-Value |
|---|
| Query 1 | Query 2 | Query 3 |
|---|
| Scenario 1 (D) | 0.008 | 0.257 | 0.046 |
| Scenario 2 (S) | 0.013 | 0.005 | 0.014 |
| Scenario 3 (E) | <0.001 | 0.002 | <0.001 |
| Scenario 4 (D + S) | 0.083 | 0.564 | 0.317 |
| Scenario 5 (D + E) | 0.046 | 0.157 | 0.014 |
| Scenario 6 (S + E) | 0.001 | 0.005 | <0.001 |