Next Article in Journal
Objective Audiovestibular Assessment After Traumatic Brain Injury in Medico-Legal Contexts: A Narrative Expert Review and Practical Cross-Check Framework
Next Article in Special Issue
Repeated Blunt-Force Trauma in an Elderly Male: An Atypical Intimate Partner Homicide Case Involving a Female Partner
Previous Article in Journal
Chained Lives: Veterinary Perceptions of Dog Tethering and Their Implications for Regulatory and Criminal Frameworks in Portugal
Previous Article in Special Issue
The Spectrum of Choice: A Review of European Abortion Legal Frameworks from a Medicolegal Perspective
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

AI-Based Quantitative Handwriting and Signature Feature Analysis: Development and Validation of a Mobile Application for Forensic Document Examination—A Preliminary Study

1
Department of Forensic Medicine, Faculty of Medicine, Balikesir University, 10145 Balikesir, Turkey
2
Department of Forensic Medicine, Balikesir Ataturk City Hospital, 10100 Balikesir, Turkey
*
Authors to whom correspondence should be addressed.
Forensic Sci. 2026, 6(2), 41; https://doi.org/10.3390/forensicsci6020041
Submission received: 13 April 2026 / Revised: 15 May 2026 / Accepted: 16 May 2026 / Published: 20 May 2026
(This article belongs to the Special Issue Feature Papers in Forensic Sciences)

Abstract

Background/Objective: Forensic document examination (FDE) traditionally relies on subjective expert opinion. This preliminary study was designed to develop and validate a hybrid deep learning model (ResNet-50 + bidirectional long short-term memory [BiLSTM]) for quantitative handwriting and signature feature analysis, and to compare its performance, under standardized experimental conditions, with that of three certified forensic document examiners. Methods: Handwriting and signature samples were collected from 225 individuals in a standardized setting. Fifteen quantitative handwriting features were extracted, the dataset was split into training (70%, n = 158) and testing (30%, n = 67) subsets using stratified random sampling, and ground truth for analytic categories was defined by majority consensus among the three examiners (with adjudicated review for disagreements). A hybrid architecture combining a ResNet-50 backbone and a bidirectional LSTM encoder was used. Results: The model demonstrated 93.4% accuracy, an F1-score of 0.926, and an AUC-ROC of 0.968 on the held-out test set. Under our task-specific experimental conditions, the model performed better than examiners on slant analysis (96.8% vs. 93.2%, p = 0.002), pressure profiling (94.1% vs. 91.7%, p = 0.019), and age estimation (87.4% vs. 82.1%, p = 0.011); examiners performed better on forgery detection (95.8% vs. 91.2%, p = 0.008) and signature verification (96.1% vs. 92.3%, p < 0.012). Mean processing time was reduced by 99.6% (0.8 s vs. 197 s per case). Conclusions: Within the limits of this preliminary single-centre study, the system showed performance comparable to certified examiners on several quantitative tasks and complementary strengths overall, supporting its feasibility as an adjunctive tool in a hybrid human–AI workflow. Broader, multi-centre validation and explainability work are required before any forensic deployment can be considered.

1. Introduction

Forensic document examination (FDE) is one of the most critical fields in forensic sciences and is a pillar of evidence analysis utilised to ascertain validity during civil trials, criminal investigations, will disputes, financial fraud cases, or identity verification (ID) pleadings [1,2]. Handwriting analysis and signature verification are still the most sought-after FDE services, accounting for 60–75% of all document examination casework [3].
Traditionally, forensic handwriting and signature examination has depended on the trained visual judgement and experience of certified document examiners. These specialists analyse a constellation of features, such as letter formations, slant angle, pen pressure dynamics, baseline alignment, and spacing patterns, in order to provide opinions on authorship attribution or document authenticity. Although expert examiners possess unique and invaluable expertise, established limitations of subjective visual approaches persist. In particular, between-examiner reproducibility and the existence of known error rates have become central concerns for legal admissibility under contemporary forensic science standards [3,4]; published Cohen’s kappa values for inter-examiner agreement vary between 0.72 and 0.91 [3,4].
It is also important to clarify the terminology used in this work. The term “graphology” has historically been used to denote the inference of personality traits from handwriting, an interpretive practice that lacks empirical validation and is not part of forensic document examination. In the present study, we therefore restrict our terminology to quantitative handwriting and signature feature analysis, that is, the objective measurement of geometric, kinematic, and morphological features of handwritten material, in order to align with current forensic practice and avoid ambiguity.
The introduction of deep learning architectures has begun a paradigm shift in many scientific fields. Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have achieved human-level or super-human performance in medical imaging, natural language processing, and biometric identification tasks previously thought to be the exclusive domain of expert judgement [5,6]. The forensic sciences, however, have lagged these advances. Recent reviews have highlighted both the promise and the cautions of integrating AI into forensic decision-making, emphasising that AI tools should be viewed as decision support rather than as replacements for human experts, and stressing the need for transparency, reproducibility, and known error rates as prerequisites for legal admissibility [7,8].
Several pioneering works have applied machine learning to handwriting and signature analysis. Hafemann et al. proposed SigNet, a CNN-based architecture for offline signature verification [4]; Lai and Jin introduced path-signature features for offline writer identification [9]; Diaz et al. provided a comprehensive perspective on handwritten signature technology [10]; Tolosana et al. introduced DeepSign for online signature verification [11]; and Zois et al. presented sequential motif profiles and topological plots for offline signature verification [12]. Complementary contributions include the writer-independent offline signature verification framework of Zois et al. [13], the ensemble-based forgery-reduction approach of Bertolini et al. [14], and the convolutional Siamese network of Dey et al. [15]. However, these prior studies share several recurrent shortcomings: most evaluate a single task (typically signature verification) rather than a multi-domain feature set; very few include a head-to-head comparison with certified forensic document examiners under matched experimental conditions; and reporting of formal ethics-committee oversight and standardised acquisition protocols is limited [16].
To address these shortcomings, the specific objectives of the present preliminary study were:
(1) to develop a hybrid deep learning architecture combining a ResNet-50 convolutional backbone [5] and a bidirectional LSTM temporal encoder [6] for the simultaneous quantitative analysis of handwriting and signature features;
(2) to validate the system on a prospectively collected, ethically approved cohort of 225 participants;
(3) to benchmark its diagnostic performance against three certified forensic document examiners on ten predefined analytic categories under matched experimental conditions; and
(4) to characterise the relative strengths, weaknesses, and complementarity of human and AI performance for incorporation into a hybrid human–AI workflow.
The ten analytic categories used for the comparison were: (i) font/style recognition; (ii) forgery detection; (iii) age estimation (within ±5 years); (iv) gender prediction; (v) slant analysis; (vi) pressure profile assessment; (vii) signature verification; (viii) forged signature detection; (ix) letter connectivity assessment; and (x) overall handwriting consistency.

2. Materials and Methods

2.1. Study Design and Ethical Approval

This cross-sectional, single-centre, diagnostic accuracy study was carried out at the Department of Forensic Medicine, Faculty of Medicine, Balikesir University, from 1 December 2025 to 1 March 2026. Ethical approval. The study protocol was approved by the Balikesir University Health Research Ethics Committee (Decision No: 2025/9-4, Date: 4 November 2025) and conducted in accordance with the Declaration of Helsinki (2013 revision) [17]. Written informed consent was obtained from all participants at the time of enrolment. Participants were each given a unique identifier (K-001 to K-225) which kept their identities anonymous.

2.2. Participants

In total, 225 participants were recruited via convenience sampling from the staff and visitors of Balikesir University Faculty of Medicine, which is acknowledged as a potential source of selection bias and a limitation of the present preliminary work (see Section 4). Inclusion criteria were: (1) age ≥ 18 years; (2) ability to write fluently in Turkish script; and (3) willingness to provide samples under standardised conditions. Exclusion criteria were musculoskeletal disorders of the dominant hand, acute pain conditions, or neurological states affecting fine motor control. The cohort comprised 120 males (53.3%) and 105 females (46.7%), with a mean age of 48.0 years (SD = 17.2, range 18–75). Educational status was reported as: primary school 4.4%, secondary school 11.1%, high school 26.7%, university 44.0%, and postgraduate 13.8%. 88.4% of participants reported being right-handed. The dataset was split into training (70%, n = 158) and testing (30%, n = 67) sets by stratified random sampling. Detailed demographic characteristics, including the breakdown by training and test sets, are provided in Supplementary Table S3.

2.3. Data Collection Protocol

Handwriting and signature samples were collected under controlled environmental conditions (500 lux illumination, 22 ± 2 °C, fixed-height desk surface). Standardised materials were unlined A4 white paper (80 g/m2) and a medium-point ballpoint pen (0.7 mm tip, blue ink). Participants completed: (1) cursive transcription of a standardised pangram repeated three times; (2) block-letter transcription of a 50-word passage; (3) numerical sequence writing (0–9, repeated five times); and (4) signature production (five repetitions at natural pace and five at accelerated pace). Forty-five participants additionally produced simulated forgeries of three randomly assigned target signatures. These forgeries were produced by non-specialist volunteer participants under unrestricted simulation conditions and therefore represent low-skill simulated forgeries; results should not be generalised to skilled or adversarial forgery scenarios encountered in real casework. All samples were digitised at 600 DPI using a Canon CanoScan LiDE 400 scanner (Canon Inc., Tokyo, Japan), with Otsu’s adaptive thresholding [18], morphological noise removal, and Hough-transform-based skew correction.

2.4. Quantitative Handwriting Feature Extraction

Python 3.10 with OpenCV 4.8 and scikit-image 0.21 was used to compute fifteen quantitative features across five domains: (1) Spatial features—slant angle, letter height, line spacing, word spacing, and margin width; (2) Pressure and dynamics features—pressure score (0–100), writing speed, and pen-lift counts; (3) Connectivity features—letter connectivity ratio and line slant; (4) Regularity features—consistency score (0–100) and regularity score (0–100); (5) Morphological complexity features—contour complexity [18], fractal dimension [19], and a composite morphological index. To improve reproducibility, the operational definitions and computational formulas of all custom-defined features (pressure score, consistency score, regularity score, composite morphological index, and the global quality index of Section 2.5) are provided in Supplementary Table S1.

2.5. Signature Morphological Analysis

Signature specimens were used to extract eleven morphometric parameters: total length, maximum height, bounding-box area, complexity index, pressure variance, speed-profile score (SPS), texture analysis output, continuity ratio (CR), repeatability index (RI), uniqueness score (US), forgery-resistance score, and global quality index. Definitions and equations for non-standard parameters are listed in Supplementary Table S1.

2.6. Deep Learning Architecture

The model used a hybrid architecture combining a ResNet-50 backbone [5] pretrained on ImageNet, which produced a 2048-dimensional visual feature vector, with a two-layer bidirectional LSTM [6] (256 hidden units per direction) for sequential representation. LSTM outputs were pooled through Bahdanau attention [20] and concatenated with the CNN features, yielding a 2560-dimensional embedding. This embedding was fed through two fully connected layers (1024 → 256 units) with batch normalisation, GELU activation, and dropout (0.35). The network has multiple task-specific output heads. Specifically, a Siamese branch with weight-sharing and contrastive loss is used for writer identification and signature-pair verification (categories vii–viii of Section 1) [21]; a softmax cross-entropy classifier is used for the discrete classification categories (i, ii, iv, ix, x); and a regression head with smooth L1 loss is used for the quantitative measurement categories (iii, v, vi). This separation ensures that each analytic category is mapped to a clearly defined output and loss function, supporting downstream auditability.

2.7. Training Protocol

Training was performed for 50 epochs with the AdamW optimiser [22] (initial learning rate 1 × 10−3, weight decay 1 × 10−4, cosine annealing, batch size 32). Data augmentation comprised random rotation (±15°), translation (±10%), elastic deformation, Gaussian noise (σ = 0.02), and brightness/contrast adjustment (±20%). Early stopping was applied with a patience of 10 epochs. Mixed-precision (FP16) training was performed on a single NVIDIA A100 GPU and required approximately 4.5 h.

2.8. Expert Examiner Assessment

The reference standard was provided by three independent certified forensic document examiners with 22, 15, and 10 years of casework experience, respectively. None of the examiners were aware of the AI model’s predictions or each other’s assessments. Each examiner independently assessed all 67 cases in the test set across the ten predefined analytic categories using a standardised evaluation form. Examiners were explicitly permitted to record an “inconclusive” judgement; for the diagnostic-accuracy analyses, “inconclusive” was treated as a separate category and reported in Supplementary Table S2 rather than forced to a binary outcome. Disagreement among the three examiners was resolved by majority consensus (2/3 or 3/3), and ties or fully discordant cases were flagged and adjudicated in a structured discussion meeting; the reference label assigned at this meeting was used as the ground truth for training and evaluation.

2.9. Statistical Analysis

Diagnostic performance metrics (accuracy, sensitivity, specificity, precision, F1-score, AUC-ROC, Cohen’s kappa [23], Matthews correlation coefficient [24]) were computed. We performed five-fold stratified cross-validation and an ablation analysis. Paired comparisons used McNemar’s test [25]. Intraclass correlation coefficients (ICC, two-way random effects) [26] and Bland–Altman analysis [27] were used for agreement; ICC interpretation followed Cicchetti’s guidelines [28]. Statistical significance was set at α = 0.05 with Bonferroni correction. Analyses were performed in Python 3.10 with SciPy 1.11 and scikit-learn 1.3.

3. Results

3.1. Overall Model Performance

On the training set, the AI model achieved 96.2% accuracy, F1-score 0.951, and AUC-ROC 0.987. On the held-out test set: 93.4% accuracy, 92.1% sensitivity, 94.6% specificity, 93.1% precision, F1-score 0.926, AUC-ROC 0.968, Cohen’s κ = 0.867, and MCC = 0.866 (Table 1). The 2.8-percentage-point training-test accuracy gap indicates acceptable generalisation.
Overall accuracy did not differ significantly between the AI (93.4%) and the expert mean (94.4%) (McNemar’s χ2 = 2.33, p = 0.127). Specificity was higher for the AI (94.6% vs. 93.3%, p = 0.018), while sensitivity was higher for examiners (95.4% vs. 92.1%, p = 0.043). These differences are reported descriptively and should not be interpreted as evidence of generalised superiority of either approach.

3.2. Category-Specific Performance

Across the ten predefined categories (Table 2), the AI was numerically higher than experts in six categories, with statistically significant advantages in slant analysis (+3.6%, p = 0.002), age estimation (+5.3%, p = 0.011), pressure profile (+2.4%, p = 0.019), letter connectivity (+2.8%, p = 0.015), and font recognition (+0.8%, p = 0.034). Examiners performed significantly better than the AI on forgery detection (−4.6%, p = 0.008), signature verification (−3.8%, p = 0.012), and forged-signature detection (−4.5%, p = 0.006). The differences in gender prediction (+2.1%) and overall consistency (+0.2%) did not reach statistical significance (p = 0.285 and p = 0.412, respectively) and should be interpreted as comparable performance, not as a meaningful advantage. Operational definitions of each analytic category, including the criteria used for binary scoring, are provided in Supplementary Table S2.

3.3. Cross-Validation Results

Five-fold cross-validation showed consistent performance: accuracy 92.8–94.2% (mean 93.46%, SD 0.53%), F1-score 0.921–0.935 (mean 0.928, SD 0.005), AUC-ROC 0.964–0.971 (mean 0.968, SD 0.003) (Table 3).

3.4. Ablation Study

The most discriminative features were slant angle (Δ = −2.9%, p = 0.003) and pressure profile (Δ = −2.2%, p = 0.008). Removing the bidirectional LSTM (CNN-only) and removing the CNN (BiLSTM-only) reduced accuracy by 4.3 and 6.7 percentage points, respectively (both p < 0.001), indicating that the spatial and temporal streams are complementary (Table 4).

3.5. Reliability and Agreement Analysis

The intra-rater reliability of the AI system (ICC = 0.978, 95% CI 0.962–0.988) was higher than inter-expert agreement (ICC = 0.891, 95% CI 0.843–0.928). AI–expert agreement was highest with Expert 1 (ICC = 0.912) and lowest with Expert 3 (ICC = 0.884) (Table 5). Bland–Altman analysis showed minimal systematic bias for slant angle (mean difference +0.32°, limits of agreement [LoA] −3.31 to +3.95°), pressure score (−1.45, LoA −9.70 to +6.80), and letter size (+0.08 mm, LoA −0.74 to +0.90 mm) (Table 6).

3.6. Comparison with Published Literature

The accuracy of the proposed system (93.4%) was numerically higher than that of most prior comparators except SynSig2Vec (93.1%) (Table 7). These comparisons should be interpreted cautiously: prior studies differ from ours in dataset composition, task definitions, evaluation metrics, and the presence or absence of expert benchmarking, and therefore numerical differences in headline accuracy do not in themselves establish methodological superiority.

4. Discussion

This preliminary study describes the development and validation of a hybrid AI system for quantitative handwriting and signature feature analysis, with an AUC-ROC of 0.968 on the held-out test set. Our central finding is that the AI system and certified examiners showed task-dependent, complementary performance under matched experimental conditions, rather than generalised superiority of one approach over the other [7,8].
The AI model performed better than examiners in six of ten categories, with the largest gains in slant analysis (+3.6%), age estimation (+5.3%), and pressure profiling (+2.4%). These categories share a common characteristic: they rely heavily on precise quantitative measurements of geometric or pressure-related descriptors, which are intrinsically well suited to computational extraction and which suffer from known measurement variability when performed by human raters [3,4]. Crawford et al., for instance, have shown that statistical decomposition of handwritten material into graphical components and a Bayesian writership analysis can outperform traditional intuitive comparison on similar geometric tasks [16]. Our results are consistent with this trend and extend it by showing that an end-to-end deep learning system, trained jointly on multiple feature domains, can match or exceed examiner accuracy on these specific quantitative subtasks. By contrast, examiners were significantly better than the AI on forgery detection (−4.6%), signature verification (−3.8%), and forged-signature detection (−4.5%). These tasks require integrative pattern recognition, contextual judgement, and the cumulative casework experience that the published literature has long associated with expert performance [3,4,10]. The pattern observed here mirrors recent results in other forensic AI evaluations, in which AI tools provide useful decision support but do not replace expert reasoning, particularly when adversarial or skilled forgeries are involved [7,8].
The reliability findings deserve particular emphasis. The intra-rater reliability of the AI system (ICC = 0.978) exceeded inter-expert agreement (ICC = 0.891). In forensic practice, reproducibility is one of the explicit Daubert/Kumho admissibility criteria [7] and is repeatedly cited in cross-examination [3]. A deterministic system that returns identical outputs on identical inputs therefore offers a measurable advantage on this specific dimension; however, this advantage does not translate directly into evidentiary weight, because reproducibility without explainability is insufficient for legal admissibility [8]. The fact that AI–expert agreement was highest with the most experienced examiner (Expert 1, ICC = 0.912) is suggestive of convergence in outputs, but we are careful to note that convergence in measurements does not imply that the model has learned the same internal reasoning processes that experts use; this would require explicit explainability analyses (see below).
The ablation analysis identified slant angle and pressure profile as the two most discriminative features, in agreement with the long-standing graphological literature [10] and with the more recent quantitative writership work of Crawford et al. [16]. The 4.3- and 6.7-percentage-point drops associated with removing the BiLSTM and the CNN respectively confirm that spatial and temporal streams contribute non-redundant information, in line with the architectural rationale of Tolosana et al.’s DeepSign [11] and Lai et al.’s SynSig2Vec [31], and provide a methodological argument for hybrid rather than single-stream architectures in this domain. The reduction of mean processing time from 197 s to 0.8 s per case, while not in itself a measure of forensic validity, is consistent with the workflow benefits described by recent reviews of forensic AI [7,8] and supports the feasibility of using the system for triage and quality-control purposes within a hybrid human–AI workflow.
Several limitations should be acknowledged in addition to those previously mentioned. First, the sample of 225 participants from a single Turkish-speaking centre limits generalisability and is one of the principal reasons we frame this work as a preliminary study; multi-centre validation across multiple scripts is a necessary next step. Second, the simulated forgeries were produced by non-specialist volunteers, and our results therefore characterise low-skill simulated forgeries rather than skilled or adversarial casework; performance under skilled-forger conditions remains an open question. Third, only three certified examiners were available as the human reference; while this is comparable to several recent published studies, expanding the panel would tighten the inter-expert ICC estimate. Fourth, and most importantly for forensic deployment, the present model does not yet provide case-level explanations of its decisions. Lack of explainability is not merely a feature gap to be added later but a major barrier to forensic and legal use, because admissibility frameworks explicitly require that experts (whether human or algorithmic) can describe the basis for their conclusions [7,8]. Future work will therefore prioritise integrating Grad-CAM [32] and SHAP-based [33] explanations, evaluating adversarial robustness against skilled forgeries, and conducting prospective multi-centre validation under casework conditions.

5. Conclusions

In this preliminary single-centre study, a hybrid ResNet-50/BiLSTM system achieved test-set accuracy of 93.4% (F1-score 0.926, AUC-ROC 0.968) on a multi-domain quantitative handwriting and signature feature analysis task. Under matched experimental conditions, the system performed comparably to three certified forensic document examiners overall, with task-specific advantages on quantitative measurement categories and task-specific disadvantages on forgery- and signature-recognition categories that depend on integrative casework experience. The system demonstrated higher intra-rater reproducibility than the inter-examiner reference and substantially shorter processing time, which together support its feasibility as an adjunctive decision-support tool within a hybrid human–AI workflow rather than as a replacement for certified examiners. Multi-centre validation across multiple scripts, evaluation against skilled forgeries, and the integration of explainability tools are required before any forensic or legal deployment can be considered.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/forensicsci6020041/s1, Table S1: Operational definitions, computational formulas, and observed value ranges of the quantitative handwriting and signature features (n = 225); Table S2: Operational definitions and scoring rules for the ten analytic categories used in the AI–examiner comparison (test set, n = 67), including the distribution of inconclusive judgements; Table S3: Demographic characteristics of the study participants (n = 225), by training and test sets.

Author Contributions

Conceptualisation, M.C. (Muhammet Can); methodology, M.C. (Muhammet Can); software, M.C. (Muhammet Can); validation, M.C. (Muhammet Can), M.C. (Meksel Cengiz) and C.I.; formal analysis, M.C. (Muhammet Can); investigation, M.C. (Muhammet Can) and M.C. (Meksel Cengiz); resources, M.C. (Muhammet Can); data curation, M.C. (Meksel Cengiz) and C.I.; writing—original draft preparation, M.C. (Muhammet Can); writing—review and editing, M.C. (Meksel Cengiz) and C.I.; visualisation, M.C. (Muhammet Can); supervision, M.C. (Muhammet Can); project administration, M.C. (Muhammet Can). All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and approved by the Balikesir University Health Research Ethics Committee (Decision No: 2025/9-4, Date: 4 November 2025).

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

The datasets generated during the current study are available from the corresponding author upon reasonable request, subject to institutional data protection policies and participant privacy requirements.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Huber, R.A.; Headrick, A.M. Handwriting Identification: Facts and Fundamentals; CRC Press: Boca Raton, FL, USA, 1999. [Google Scholar]
  2. Srihari, S.N.; Cha, S.H.; Arora, H.; Lee, S. Individuality of Handwriting. J. Forensic Sci. 2002, 47, 856–872. [Google Scholar] [CrossRef] [Scilit]
  3. Impedovo, D.; Pirlo, G. Automatic Signature Verification: The State of the Art. IEEE Trans. Syst. Man Cybern. C 2008, 38, 609–635. [Google Scholar] [CrossRef] [Scilit]
  4. Hafemann, L.G.; Sabourin, R.; Oliveira, L.S. Learning Features for Offline Handwritten Signature Verification Using Deep Convolutional Neural Networks. Pattern Recognit. 2017, 70, 163–176. [Google Scholar] [CrossRef] [Scilit]
  5. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
  6. Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit]
  7. Sessa, F.; Esposito, M.; Cocimano, G.; Sablone, S.; Karaboue, M.A.A.; Chisari, M.; Albano, D.G.; Salerno, M. Artificial Intelligence and Forensic Genetics: Current Applications and Future Perspectives. Appl. Sci. 2024, 14, 2113. [Google Scholar] [CrossRef] [Scilit]
  8. Hicklin, R.A.; Eisenhart, L.; Richetelli, N.; Miller, M.D.; Belcastro, P.; Burkes, T.M.; Parks, C.L.; Smith, M.A.; Buscaglia, J.; Peters, E.M.; et al. Accuracy and Reliability of Forensic Handwriting Comparisons. Proc. Natl. Acad. Sci. USA 2022, 119, e2119944119. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Lai, S.; Jin, L. Offline Writer Identification Based on the Path Signature Feature. In Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR), Sydney, Australia, 20–25 September 2019; pp. 1137–1142. [Google Scholar]
  10. Diaz, M.; Ferrer, M.A.; Impedovo, D.; Malik, M.I.; Pirlo, G.; Plamondon, R. A Perspective Analysis of Handwritten Signature Technology. ACM Comput. Surv. 2019, 51, 1–39. [Google Scholar] [CrossRef] [Scilit]
  11. Tolosana, R.; Vera-Rodriguez, R.; Fierrez, J.; Ortega-Garcia, J. DeepSign: Deep On-Line Signature Verification. IEEE Trans. Biom. Behav. Identity Sci. 2021, 3, 229–239. [Google Scholar] [CrossRef] [Scilit]
  12. Zois, E.N.; Zervas, E.; Tsourounis, D.; Economou, G. Sequential Motif Profiles and Topological Plots for Offline Signature Verification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 13245–13255. [Google Scholar]
  13. Zois, E.N.; Alexandridis, A.; Economou, G. Writer Independent Offline Signature Verification Based on Asymmetric Pixel Relations and Unrelated Training-Testing Datasets. Expert. Syst. Appl. 2019, 125, 14–32. [Google Scholar] [CrossRef] [Scilit]
  14. Bertolini, D.; Oliveira, L.S.; Justino, E.; Sabourin, R. Reducing Forgeries in Writer-Independent Off-Line Signature Verification Through Ensemble of Classifiers. Pattern Recognit. 2010, 43, 387–396. [Google Scholar] [CrossRef] [Scilit]
  15. Dey, S.; Dutta, A.; Toledo, J.I.; Ghosh, S.K.; Lladós, J.; Pal, U. SigNet: Convolutional Siamese Network for Writer Independent Offline Signature Verification. arXiv 2017, arXiv:1707.02131. [Google Scholar] [CrossRef] [Scilit]
  16. Crawford, A.M.; Ommen, D.M.; Carriquiry, A.L. A Statistical Approach to Aid Examiners in the Forensic Analysis of Handwriting. J. Forensic Sci. 2023, 68, 1768–1779. [Google Scholar] [CrossRef] [Scilit]
  17. World Medical Association. World Medical Association Declaration of Helsinki: Ethical Principles for Medical Research Involving Human Subjects. JAMA 2013, 310, 2191–2194. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Otsu, N. A Threshold Selection Method from Gray-Level Histograms. IEEE Trans. Syst. Man Cybern. 1979, 9, 62–66. [Google Scholar] [CrossRef] [Scilit]
  19. Mandelbrot, B.B. The Fractal Geometry of Nature; W.H. Freeman: New York, NY, USA, 1982. [Google Scholar]
  20. Bahdanau, D.; Cho, K.; Bengio, Y. Neural Machine Translation by Jointly Learning to Align and Translate. In Proceedings of the 3rd International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
  21. Bromley, J.; Guyon, I.; LeCun, Y.; Säckinger, E.; Shah, R. Signature Verification Using a “Siamese” Time Delay Neural Network. In Proceedings of the Advances in Neural Information Processing Systems 6 (NIPS 1993), Denver, CO, USA, 29 November–2 December 1993; DBLP: Trier, Germany, 1993; pp. 737–744. [Google Scholar]
  22. Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. In Proceedings of the 7th International Conference on Learning Representations (ICLR), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
  23. Cohen, J. A Coefficient of Agreement for Nominal Scales. Educ. Psychol. Meas. 1960, 20, 37–46. [Google Scholar] [CrossRef] [Scilit]
  24. Matthews, B.W. Comparison of the Predicted and Observed Secondary Structure of T4 Phage Lysozyme. Biochim. Biophys. Acta 1975, 405, 442–451. [Google Scholar] [CrossRef] [Scilit]
  25. McNemar, Q. Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages. Psychometrika 1947, 12, 153–157. [Google Scholar] [CrossRef] [Scilit]
  26. Shrout, P.E.; Fleiss, J.L. Intraclass Correlations: Uses in Assessing Rater Reliability. Psychol. Bull. 1979, 86, 420–428. [Google Scholar] [CrossRef]
  27. Bland, J.M.; Altman, D.G. Statistical Methods for Assessing Agreement Between Two Methods of Clinical Measurement. Lancet 1986, 327, 307–310. [Google Scholar] [CrossRef] [Scilit]
  28. Cicchetti, D.V. Guidelines, Criteria, and Rules of Thumb for Evaluating Normed and Standardized Assessment Instruments in Psychology. Psychol. Assess. 1994, 6, 284–290. [Google Scholar] [CrossRef]
  29. Tsourounis, D.; Theodorakopoulos, I.; Zois, E.N.; Economou, G. From Text to Signatures: Knowledge Transfer for Efficient Deep Feature Learning in Offline Signature Verification. Expert. Syst. Appl. 2022, 189, 116136. [Google Scholar] [CrossRef] [Scilit]
  30. Lai, S.; Jin, L. Recurrent Adaptation Networks for Online Signature Verification. IEEE Trans. Inf. Forensics Secur. 2019, 14, 1624–1637. [Google Scholar] [CrossRef] [Scilit]
  31. Lai, S.; Jin, L.; Zhu, Y.; Li, Z.; Lin, L. SynSig2Vec: Forgery-Free Learning of Dynamic Signature Representations by Sigma Lognormal-Based Synthesis and 1D CNN. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 6472–6485. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  32. Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 618–626. [Google Scholar]
  33. Lundberg, S.M.; Lee, S.I. A Unified Approach to Interpreting Model Predictions. In Proceedings of the Advances in Neural Information Processing Systems 30 (NeurIPS), Long Beach, CA, USA, 4–9 December 2017; DBLP: Trier, Germany, 2017; pp. 4766–4777. [Google Scholar]
Table 1. Overall performance comparison between the AI model and expert examiners.
Table 1. Overall performance comparison between the AI model and expert examiners.
MetricAI (Train)AI (Test)Expert 1Expert 2Expert 3Expert Meanp-Value
Accuracy0.9620.9340.9510.9430.9380.9440.127
Sensitivity0.9480.9210.9620.9540.9470.9540.043
Specificity0.9710.9460.9390.9310.9280.9330.018
Precision0.9550.9310.9420.9360.9300.9360.089
F1-Score0.9510.9260.9520.9450.9380.9450.068
AUC-ROC0.9870.968
Cohen’s κ0.9230.8670.9010.8850.8740.8870.094
MCC0.9210.8660.8990.8840.8730.8850.082
Time (s)0.80.8185210195197<0.001
Table 2. Category-specific performance comparison (test set, n = 67).
Table 2. Category-specific performance comparison (test set, n = 67).
CategoryAI AccExpert AccAI F1Expert F1AI SuperiorΔAccp-Value
Font Recognition0.9560.9480.9490.941Yes+0.0080.034
Forgery Detection0.9120.9580.9050.951No−0.0460.008
Age Estimation (±5 y)0.8740.8210.8670.814Yes+0.0530.011
Gender Prediction0.7830.7620.7710.755Yes+0.0210.285
Slant Analysis0.9680.9320.9610.928Yes+0.0360.002
Pressure Profile0.9410.9170.9340.912Yes+0.0240.019
Signature Verification0.9230.9610.9180.955No−0.0380.012
Forged Sig. Detection0.8970.9420.8890.935No−0.0450.006
Letter Connectivity0.9520.9240.9450.919Yes+0.0280.015
Overall Consistency0.9380.9360.9310.929Yes+0.0020.412
Table 3. Five-fold cross-validation results.
Table 3. Five-fold cross-validation results.
FoldTrain Acc (%)Test Acc (%)F1-ScoreAUC-ROCPrecisionRecall
196.493.10.9240.9670.9310.917
295.894.20.9350.9710.9420.928
396.192.80.9210.9640.9280.914
495.993.70.9300.9690.9370.923
596.393.50.9280.9680.9350.921
Mean ± SD96.1 ± 0.293.5 ± 0.50.928 ± 0.0050.968 ± 0.0030.935 ± 0.0050.921 ± 0.005
Table 4. Ablation study results.
Table 4. Ablation study results.
ConfigurationAccuracy (%)F1-ScoreAUCΔAcc (%)Impactp-Value
Full Model93.40.9260.968
(–) Fractal Dim.92.80.9200.963−0.6Low0.312
(–) Pressure91.20.9040.951−2.2High0.008
(–) Slant Angle90.50.8970.945−2.9High0.003
(–) Letter Connect.92.10.9130.958−1.3Medium0.045
(–) Contour Compl.92.50.9170.961−0.9Low0.187
(–) Writing Speed91.80.9100.955−1.6Medium0.032
CNN Only89.10.8820.938−4.3<0.001
BiLSTM Only86.70.8580.921−6.7<0.001
Table 5. Intraclass correlation coefficients for reliability assessment.
Table 5. Intraclass correlation coefficients for reliability assessment.
MeasurementICC (2,1)95% CI Lower95% CI UpperSEMReliability
AI Repeatability0.9780.9620.9880.012Excellent
Inter-Expert0.8910.8430.9280.034Good
AI vs. Expert 10.9120.8710.9420.028Excellent
AI vs. Expert 20.8970.8510.9310.031Good
AI vs. Expert 30.8840.8350.9210.035Good
Table 6. Bland–Altman agreement analysis between AI system and expert mean.
Table 6. Bland–Altman agreement analysis between AI system and expert mean.
ParameterMean BiasSDLoA LowerLoA UpperAgreement
Slant Angle (°)+0.321.85−3.31+3.95Acceptable
Pressure Score−1.454.21−9.70+6.80Acceptable
Letter Size (mm)+0.080.42−0.74+0.90Good
Complexity Index−0.020.06−0.14+0.10Good
Consistency Score+1.825.34−8.65+12.29Moderate
Table 7. Performance comparison with published literature.
Table 7. Performance comparison with published literature.
StudyYearnMethodAccuracy (%)TaskΔ vs. Ours
Present Study2026225ResNet-50 + BiLSTM93.4HW + Sig
Hafemann et al. [4]2017170CNN (SigNet)87.2Signature+6.2
Lai and Jin [9]2019200Path Signature91.5Writer ID+1.9
Diaz et al. [10]2019150DTW + SVM84.3Signature+9.1
Tolosana et al. [11]2021310DeepSign (RNN)92.8Sig Verif.+0.6
Zois et al. [12]2020185Motif + Topo90.1Signature+3.3
Tsourounis et al. [29]2022250KD + CNN91.9Signature+1.5
Lai and Jin [30]2019180RAN92.3Sig Verif.+1.1
Lai et al. [31]2022300SynSig2Vec93.1Sig Verif.+0.3
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Can, M.; Işık, C.; Cengiz, M. AI-Based Quantitative Handwriting and Signature Feature Analysis: Development and Validation of a Mobile Application for Forensic Document Examination—A Preliminary Study. Forensic Sci. 2026, 6, 41. https://doi.org/10.3390/forensicsci6020041

AMA Style

Can M, Işık C, Cengiz M. AI-Based Quantitative Handwriting and Signature Feature Analysis: Development and Validation of a Mobile Application for Forensic Document Examination—A Preliminary Study. Forensic Sciences. 2026; 6(2):41. https://doi.org/10.3390/forensicsci6020041

Chicago/Turabian Style

Can, Muhammet, Cihangir Işık, and Meksel Cengiz. 2026. "AI-Based Quantitative Handwriting and Signature Feature Analysis: Development and Validation of a Mobile Application for Forensic Document Examination—A Preliminary Study" Forensic Sciences 6, no. 2: 41. https://doi.org/10.3390/forensicsci6020041

APA Style

Can, M., Işık, C., & Cengiz, M. (2026). AI-Based Quantitative Handwriting and Signature Feature Analysis: Development and Validation of a Mobile Application for Forensic Document Examination—A Preliminary Study. Forensic Sciences, 6(2), 41. https://doi.org/10.3390/forensicsci6020041

Article Metrics

Back to TopTop