1. Introduction
Forensic document examination (FDE) is one of the most critical fields in forensic sciences and is a pillar of evidence analysis utilised to ascertain validity during civil trials, criminal investigations, will disputes, financial fraud cases, or identity verification (ID) pleadings [
1,
2]. Handwriting analysis and signature verification are still the most sought-after FDE services, accounting for 60–75% of all document examination casework [
3].
Traditionally, forensic handwriting and signature examination has depended on the trained visual judgement and experience of certified document examiners. These specialists analyse a constellation of features, such as letter formations, slant angle, pen pressure dynamics, baseline alignment, and spacing patterns, in order to provide opinions on authorship attribution or document authenticity. Although expert examiners possess unique and invaluable expertise, established limitations of subjective visual approaches persist. In particular, between-examiner reproducibility and the existence of known error rates have become central concerns for legal admissibility under contemporary forensic science standards [
3,
4]; published Cohen’s kappa values for inter-examiner agreement vary between 0.72 and 0.91 [
3,
4].
It is also important to clarify the terminology used in this work. The term “graphology” has historically been used to denote the inference of personality traits from handwriting, an interpretive practice that lacks empirical validation and is not part of forensic document examination. In the present study, we therefore restrict our terminology to quantitative handwriting and signature feature analysis, that is, the objective measurement of geometric, kinematic, and morphological features of handwritten material, in order to align with current forensic practice and avoid ambiguity.
The introduction of deep learning architectures has begun a paradigm shift in many scientific fields. Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have achieved human-level or super-human performance in medical imaging, natural language processing, and biometric identification tasks previously thought to be the exclusive domain of expert judgement [
5,
6]. The forensic sciences, however, have lagged these advances. Recent reviews have highlighted both the promise and the cautions of integrating AI into forensic decision-making, emphasising that AI tools should be viewed as decision support rather than as replacements for human experts, and stressing the need for transparency, reproducibility, and known error rates as prerequisites for legal admissibility [
7,
8].
Several pioneering works have applied machine learning to handwriting and signature analysis. Hafemann et al. proposed SigNet, a CNN-based architecture for offline signature verification [
4]; Lai and Jin introduced path-signature features for offline writer identification [
9]; Diaz et al. provided a comprehensive perspective on handwritten signature technology [
10]; Tolosana et al. introduced DeepSign for online signature verification [
11]; and Zois et al. presented sequential motif profiles and topological plots for offline signature verification [
12]. Complementary contributions include the writer-independent offline signature verification framework of Zois et al. [
13], the ensemble-based forgery-reduction approach of Bertolini et al. [
14], and the convolutional Siamese network of Dey et al. [
15]. However, these prior studies share several recurrent shortcomings: most evaluate a single task (typically signature verification) rather than a multi-domain feature set; very few include a head-to-head comparison with certified forensic document examiners under matched experimental conditions; and reporting of formal ethics-committee oversight and standardised acquisition protocols is limited [
16].
To address these shortcomings, the specific objectives of the present preliminary study were:
(1) to develop a hybrid deep learning architecture combining a ResNet-50 convolutional backbone [
5] and a bidirectional LSTM temporal encoder [
6] for the simultaneous quantitative analysis of handwriting and signature features;
(2) to validate the system on a prospectively collected, ethically approved cohort of 225 participants;
(3) to benchmark its diagnostic performance against three certified forensic document examiners on ten predefined analytic categories under matched experimental conditions; and
(4) to characterise the relative strengths, weaknesses, and complementarity of human and AI performance for incorporation into a hybrid human–AI workflow.
The ten analytic categories used for the comparison were: (i) font/style recognition; (ii) forgery detection; (iii) age estimation (within ±5 years); (iv) gender prediction; (v) slant analysis; (vi) pressure profile assessment; (vii) signature verification; (viii) forged signature detection; (ix) letter connectivity assessment; and (x) overall handwriting consistency.
4. Discussion
This preliminary study describes the development and validation of a hybrid AI system for quantitative handwriting and signature feature analysis, with an AUC-ROC of 0.968 on the held-out test set. Our central finding is that the AI system and certified examiners showed task-dependent, complementary performance under matched experimental conditions, rather than generalised superiority of one approach over the other [
7,
8].
The AI model performed better than examiners in six of ten categories, with the largest gains in slant analysis (+3.6%), age estimation (+5.3%), and pressure profiling (+2.4%). These categories share a common characteristic: they rely heavily on precise quantitative measurements of geometric or pressure-related descriptors, which are intrinsically well suited to computational extraction and which suffer from known measurement variability when performed by human raters [
3,
4]. Crawford et al., for instance, have shown that statistical decomposition of handwritten material into graphical components and a Bayesian writership analysis can outperform traditional intuitive comparison on similar geometric tasks [
16]. Our results are consistent with this trend and extend it by showing that an end-to-end deep learning system, trained jointly on multiple feature domains, can match or exceed examiner accuracy on these specific quantitative subtasks. By contrast, examiners were significantly better than the AI on forgery detection (−4.6%), signature verification (−3.8%), and forged-signature detection (−4.5%). These tasks require integrative pattern recognition, contextual judgement, and the cumulative casework experience that the published literature has long associated with expert performance [
3,
4,
10]. The pattern observed here mirrors recent results in other forensic AI evaluations, in which AI tools provide useful decision support but do not replace expert reasoning, particularly when adversarial or skilled forgeries are involved [
7,
8].
The reliability findings deserve particular emphasis. The intra-rater reliability of the AI system (ICC = 0.978) exceeded inter-expert agreement (ICC = 0.891). In forensic practice, reproducibility is one of the explicit Daubert/Kumho admissibility criteria [
7] and is repeatedly cited in cross-examination [
3]. A deterministic system that returns identical outputs on identical inputs therefore offers a measurable advantage on this specific dimension; however, this advantage does not translate directly into evidentiary weight, because reproducibility without explainability is insufficient for legal admissibility [
8]. The fact that AI–expert agreement was highest with the most experienced examiner (Expert 1, ICC = 0.912) is suggestive of convergence in outputs, but we are careful to note that convergence in measurements does not imply that the model has learned the same internal reasoning processes that experts use; this would require explicit explainability analyses (see below).
The ablation analysis identified slant angle and pressure profile as the two most discriminative features, in agreement with the long-standing graphological literature [
10] and with the more recent quantitative writership work of Crawford et al. [
16]. The 4.3- and 6.7-percentage-point drops associated with removing the BiLSTM and the CNN respectively confirm that spatial and temporal streams contribute non-redundant information, in line with the architectural rationale of Tolosana et al.’s DeepSign [
11] and Lai et al.’s SynSig2Vec [
31], and provide a methodological argument for hybrid rather than single-stream architectures in this domain. The reduction of mean processing time from 197 s to 0.8 s per case, while not in itself a measure of forensic validity, is consistent with the workflow benefits described by recent reviews of forensic AI [
7,
8] and supports the feasibility of using the system for triage and quality-control purposes within a hybrid human–AI workflow.
Several limitations should be acknowledged in addition to those previously mentioned. First, the sample of 225 participants from a single Turkish-speaking centre limits generalisability and is one of the principal reasons we frame this work as a preliminary study; multi-centre validation across multiple scripts is a necessary next step. Second, the simulated forgeries were produced by non-specialist volunteers, and our results therefore characterise low-skill simulated forgeries rather than skilled or adversarial casework; performance under skilled-forger conditions remains an open question. Third, only three certified examiners were available as the human reference; while this is comparable to several recent published studies, expanding the panel would tighten the inter-expert ICC estimate. Fourth, and most importantly for forensic deployment, the present model does not yet provide case-level explanations of its decisions. Lack of explainability is not merely a feature gap to be added later but a major barrier to forensic and legal use, because admissibility frameworks explicitly require that experts (whether human or algorithmic) can describe the basis for their conclusions [
7,
8]. Future work will therefore prioritise integrating Grad-CAM [
32] and SHAP-based [
33] explanations, evaluating adversarial robustness against skilled forgeries, and conducting prospective multi-centre validation under casework conditions.