1. Introduction
Digital forensics (DF) has evolved from a largely tool-assisted investigative practice into a data-intensive scientific discipline. The modern digital environment, characterized by cloud infrastructures, distributed storage, and widespread Internet of Things (IoT) devices, has generated evidence volumes that are both massive and heterogeneous. Traditional approaches, which relied on exhaustive disk duplication and manual artifact review, are increasingly inadequate for meeting contemporary investigative timelines and evidential integrity requirements. In recent years, the integration of intelligence-led automation, adaptive reasoning, and machine learning has begun to reshape parts of the digital forensics process. Intelligent systems now assist investigators by dynamically identifying relevant evidence, optimizing imaging parameters, and prioritizing analytical tasks based on contextual relevance. These developments reflect an emerging shift from static data collection toward responsive, knowledge-driven discovery in certain areas of the field [
1,
2,
3]. As digital storage grows and evidence originates from more sources, there is a stronger need for systems that balance speed, thoroughness, and legal trust. Ref. [
4] showed that features such as TRIM and garbage collection in SSDs can erase data too quickly, making forensic tests difficult to reproduce. To address this, researchers have proposed flexible imaging, faster data collection, and the selective retention of the most relevant evidence [
5]. In parallel, privacy-oriented techniques such as encrypted deduplication and selective imaging help meet legal and ethical standards [
6].
While prior works have examined discrete phases of intelligent digital forensics, the field still lacks a systematic account of how farintelligent methods have actually penetrated its various domains. Earlier studies such as [
7] demonstrated the application of clustering and classification algorithms to automate reasoning in the analysis stage of investigations; however, such efforts were largely confined to analytical intelligence and offer little sense of the broader landscape.
This study addresses that gap through a systematic mapping review of six key domains of digital forensics research: Imaging & Acquisition, Data Reduction & Optimization, Memory Forensics, Mobile & IoT Forensics, Multimedia Forensics, and Frameworks, Standards & Privacy. Rather than assuming that intelligence is uniformly present, the review treats the adoption of intelligent methods as a variable to be measured, mapping where such methods have taken hold and where the field remains reliant on classical, deterministic techniques.
Accordingly, this work contributes by:
Providing a systematic map of six digital forensics domains that positions the adoption of intelligent methods as a measured dimension rather than an assumed feature.
Analyzing each domain through a unified lens (overview, challenges, intelligent approaches, and future outlook) and quantifying the extent of intelligent-method adoption within it.
Revealing a pronounced and uneven penetration of intelligence across the field, concentrated in multimedia and mobile/IoT forensics while imaging, data reduction, memory, and governance remain predominantly classical, and discussing the cross-domain challenges (interoperability, explainability, and ethical accountability) that this asymmetry raises.
Table 1 situates this review against representative prior work, clarifying its distinct contribution.
The remainder of this paper is organized as follows.
Section 2 reviews related studies and situates this work within the evolution of intelligent digital forensics.
Section 3 details the review methodology.
Section 4 analyzes the six research streams and the distribution of intelligent methods across them.
Section 5 presents a cross-domain discussion, and
Section 6 concludes with implications and future directions.
2. Background
Digital forensics has undergone a paradigm shift from largely manual, tool-assisted examination toward intelligence-supported and automated investigation. Early conceptual work anticipated this transition: Irons and Lallie argued that traditional digital forensics workflows would increasingly require intelligent support to remain scalable and effective [
12]. The growing heterogeneity of evidence, from solid-state drives and distributed cloud environments to mobile and IoT ecosystems, has since exposed the practical limitations of linear, exhaustive workflows. As Horsman noted, full bit-stream duplication once ensured evidential completeness, but modern multi-terabyte storage systems have rendered exhaustive imaging inefficient and often impractical [
1].
Researchers have long recognized these pressures. Rughani argued that conventional, manpower-intensive methods cannot keep pace with the growth of cybercrime, proposing one of the earliest AI-driven frameworks to automate routine digital forensics processes [
13]. Jarrett and Choo further emphasized that automation and AI are reshaping digital forensics practice by accelerating triage, reducing analyst workload, and uncovering patterns that manual review often misses [
14]. Dubey et al. similarly highlighted that the increasing complexity of digital investigations demands more adaptive, intelligence-enabled models to maintain investigative efficiency and reliability [
15].
2.1. Origins of Intelligent Digital Forensics
The formal articulation of intelligence within digital forensics workflows was consolidated by Adam and Varol, who demonstrated how clustering and classification algorithms could support analytical reasoning by identifying relationships within crime datasets [
7]. Their work focused primarily on the analysis phase, that is, the post-acquisition inference stage, showing how computational models could enhance human decision making once evidence had already been collected.
Building on this conceptual foundation, Kuforiji expanded the role of intelligence beyond analytical reasoning and into operational automation [
16]. His AI-driven DFIR framework integrated machine learning, natural language processing, and robotic process automation to accelerate triage and correlation while still maintaining traceability and digital forensics soundness.
Hall et al. added an important caution: as algorithmic reasoning becomes more prevalent, explainable-AI mechanisms are essential to ensure transparency, interpretability, and courtroom defensibility [
17].
2.2. Advances in Intelligent Processes
Recent studies show that intelligence can strengthen several phases of the digital forensics process, although, as this review makes clear, its adoption remains uneven across them. Within acquisition, selective and integrity-verified methods have emerged to reduce overhead in time, storage, and analyst effort while still maintaining evidential reliability [
2,
3]. In storage-aware contexts, Freiling et al. demonstrated that volatile SSD behaviors such as TRIM and garbage collection undermine repeatability, making adaptive and learning-capable acquisition strategies increasingly necessary [
4]. Karadsheh et al. further addressed efficiency and legal compliance through encrypted deduplication and semantically guided selective imaging [
6].
Beyond data collection, researchers have also been developing broader, process-focused frameworks. Song and Li showed that big-data systems can support scalable, standardized analysis when investigators face massive volumes of evidence [
18]. Bhavsar and colleagues demonstrated that distributed forensic workflows can manage large data variety without sacrificing analytical quality [
19]. Dimitriadis et al. proposed D4I, a structured, attack-agnostic investigation model that organizes artifacts and links them to Cyber Kill Chain stages, making examinations clearer, more repeatable, and more rigorous [
20].
Collectively, these efforts show that “intelligence” in digital forensics now emerges across multiple layers: identifying useful data early, adapting imaging and reduction strategies dynamically, improving transparency and explainability, and supporting governance that protects privacy and aligns with standards. Even so, the field remains fragmented, with most intelligent components solving narrow problems, and many areas continuing to depend on traditional methods rather than end-to-end intelligent processes.
2.3. Motivation for a Multi-Domain Analysis of Intelligent Digital Forensics
Digital forensics research has made real progress across many areas, including imaging, data reduction, memory analysis, mobile and IoT ecosystems, multimedia evidence, and governance-focused frameworks. Yet, as Dubey et al. and Jarrett and Choo point out, these advances often appear in isolation. Intelligent techniques are typically developed and evaluated within narrow problem spaces rather than as components of a connected forensic architecture [
14,
15].
For that reason, this study treats intelligence not as a single stage in the forensic process but as a dimension that can be assessed across the field. The six domains examined are Imaging & Acquisition, Data Reduction & Optimization, Memory Forensics, Mobile & IoT Forensics, Multimedia Forensics, and Frameworks, Standards & Privacy, which reflect the major research streams in contemporary practice. Using the methodology outlined in
Section 3, the review evaluates the extent to which each domain has adopted intelligent capabilities. As the analysis will show, adoption is uneven: strong in multimedia and mobile/IoT forensics, and far more limited in imaging, data reduction, memory, and governance.
These domains are not intended to represent a step-by-step investigative sequence. Instead, they capture parallel research tracks where intelligent methods have been introduced at different levels to address challenges such as scale, volatility, interpretability, and governance. By bringing these streams together under a single analytical lens, this work identifies where machine learning, automation, adaptive reasoning, and policy-aware design have taken hold, and where classical methods still dominate, with implications for scalability, evidential soundness, and ethical accountability in modern digital forensics.
3. Review Methodology
This study was conducted as a systematic mapping review and is reported in accordance with the PRISMA Extension for Scoping Reviews (PRISMA-ScR) [
21]. Unlike a meta-analytic review, a mapping review charts the distribution, maturity, and methodological character of a research field rather than pooling quantitative outcomes. Accordingly, the objective here is not only to identify primary studies across the principal digital forensics (DF) domains, but to characterize the extent to which intelligent methods such as machine learning, deep learning, and related learning or reasoning-based techniques have been adopted within each domain.
3.1. Research Questions
The review is guided by three questions:
- RQ1.
Across which primary DF domains is contemporary research distributed, and in what proportions?
- RQ2.
To what extent have intelligent methods been adopted within each domain, and which domains remain predominantly classical?
- RQ3.
How has the adoption of intelligent methods evolved over time?
Intelligence is therefore treated as an analyzed variable of the corpus, not as an eligibility filter (see
Section 3.4).
3.2. Information Sources
Searches were executed across the following bibliographic databases: IEEE Xplore, the ACM Digital Library, Science Direct (Elsevier), Wiley Online Library, Scopus, Web of Science, SpringerLink, and Google Scholar. No trial registers were searched, as none are applicable to this field. The searches were executed between October 2025 and January 2026. Records were restricted to peer-reviewed, English language publications within 2006–2026. Articles published online ahead of print were eligible if their online publication date fell within the search window, even where the formal volume or issue assignment carried a later (2026) date.
3.3. Search Strategy
Because the DF sub-communities use divergent terminology, a single umbrella query under-represents several domains. Each domain was therefore searched with a domain-specific term block, conjoined (AND) with a shared intelligence/automation term block. The intelligence block functioned as a thematic scoping device, orienting the search toward automation-relevant work, rather than as a strict inclusion gate:
(“machine learning” OR “deep learning” OR “neural network*” OR “artificial intelligence” OR “explainable AI” OR XAI OR “knowledge graph*” OR “natural language processing” OR “automat*” OR “intelligent” OR “selective”)
The domain-specific term blocks are listed in
Table 2. Strings were adapted to each database’s syntax, and backward/forward citation searching was applied to included studies.
3.4. Eligibility Criteria
Consistent with the mapping objective, a study was included if it met all of the following:
- IC1.
It is a primary study or a substantive technical framework (i.e., it presents an implemented, evaluated, or otherwise original contribution).
- IC2.
It addresses a digital forensic investigative task within one of the six domains defined in
Section 3.6.
- IC3.
It engages, at least thematically, with automation or intelligent methods.
- IC4.
It is peer-reviewed and within the stated window.
Crucially, the use of an intelligent method was
not an inclusion requirement. Classical, deterministic, and rule-based studies were retained and are reported; whether a study employs an intelligent method was instead recorded as a coded variable (
Section 3.7), enabling the adoption analysis in RQ2–RQ3.
Studies were excluded (or re-tiered) if:
- EC1.
They are secondary studies (systematic reviews, surveys, or bibliometric analyses). These were not counted among the primary corpus but were retained as a separate related-reviews tier for context and comparison.
- EC2.
They concern cybersecurity or intrusion detection without direct forensic-investigative applicability.
- EC3.
They present a computer-vision or machine-learning method with no evidential or forensic framing.
- EC4.
They are non-peer-reviewed, editorial, or otherwise non-primary.
3.5. Study Selection
Searches across the eight databases yielded a large initial return, which was relevance-screened at the database interface against the domain scope during search execution; the 96 records retained from this process were exported for formal screening and constitute the records identified in the PRISMA flow. Seven duplicate records were removed, leaving 89 records for screening. All records were screened by the first author against the criteria in
Section 3.4, with the screening decision and its rationale recorded per record in a structured screening worksheet; uncertain or borderline cases were discussed with the second author and resolved by consensus. Following title/abstract and full-text assessment, seven secondary studies were reassigned to the related-reviews tier, yielding a final corpus of 82 primary studies. The complete selection process is summarized in the PRISMA flow diagram (
Figure 1 and
Supplementary Materials). To prevent a single study from inflating multiple categories, each included study was assigned to exactly one primary domain on the basis of its principal contribution; secondary domain relevance was recorded but not double-counted.
The original identification search was conducted across eight databases between October 2025 and January 2026 and, as is common for domain-scoped literature searches, relevance filtering at the database interface was applied during execution; per-database raw yields were not separately retained at that stage. To transparently characterize the breadth of the underlying literature, a supplementary domain-only search was subsequently conducted in Scopus on 10 August 2026, running each domain term block within a forensic context and restricted to the Computer Science and Engineering subject areas over the review window. This supplementary search returned 8756 records across the six domains (
Table 3), of which the 82 studies included in this review represent approximately 0.9%. Because this search postdates the original identification window and covers a single database, its yields characterize domain magnitude rather than reconstruct the original multi-database funnel. The distribution is consistent with the review’s central finding: the domains with the largest candidate literatures do not correspond to the smallest included sets, and the uneven adoption of intelligent methods reported below is therefore not an artifact of search breadth. To ensure the uneven adoption we observed is not just a quirk of how different subfields appear in search results, we compared each domain’s share of the broader Scopus literature with its share of studies that ultimately met our criteria. The two distributions pull apart in a way that directly contradicts a recall artifact. Multimedia forensics makes up 56.2% of the baseline literature but only 17.1% of the included corpus, while imaging and acquisition accounts for just 2.8% of the baseline yet rises to 22.0% of the included studies. If high adoption was simply the result of easier retrieval, the domain with the strongest intelligent uptake would appear more frequently in the included set. Instead, it is substantially under-represented relative to its literature size. This pattern shows that the adoption differences we report reflect real variation in how intelligence is being developed across domains—not a bias introduced by search recall.
3.6. Domain Taxonomy
Included studies were mapped to six primary DF domains: (i) Imaging & Acquisition; (ii) Data Reduction & Optimization; (iii) Memory Forensics; (iv) Mobile & IoT Forensics; (v) Multimedia Forensics; and (vi) Frameworks, Standards & Privacy. Mobile and IoT forensics were treated as a single domain, reflecting their shared evidentiary characteristics and the frequency with which reviewed studies address both ecosystems jointly.
3.7. Data Extraction and Coding
For each included study, the following were extracted: bibliographic metadata, publication year, assigned primary domain, and, central to this review, the intelligence method type. The latter was coded as one of: none (classical/deterministic), machine learning, deep learning, natural language processing, explainable AI, knowledge-graph reasoning, or automation. This coding underpins the domain×method analysis that constitutes the review’s principal contribution. To enable a finer analysis than a binary intelligent/classical split, each study’s coded method type was mapped onto a five-level intelligence maturity ladder, ordered by the degree of learning and inspectable reasoning involved: Level 0, classical or deterministic methods with no learning component; Level 1, rule-based automation that accelerates workflow without learning; Level 2, classical machine learning using engineered features; Level 3, deep learning and representation-learning methods, including natural-language processing and neural approaches; and Level 4, explainable or reasoning-based methods, including explainable-AI mechanisms and knowledge-graph reasoning. Studies employing more than one method were assigned to the highest level reached. This ladder distinguishes not merely whether a domain has adopted intelligent methods, but what kind and maturity of intelligence it has attained.
3.8. Synthesis and Mapping
Findings were synthesized descriptively and visualized as: a per-domain distribution of included studies; a longitudinal bubble map of studies by domain and year; a domain-level breakdown of intelligent versus classical methods; and a temporal trend of intelligent-method adoption. No quantitative meta-analysis was performed, consistent with the mapping design.
4. Intelligence in Digital Forensics: A Multi-Domain Perspective
This study adopts a structured and transparent literature mapping approach to examine how intelligent techniques manifest across digital forensics. Rather than treating intelligence as a single processing stage, the review conceptualizes it as a unifying analytical dimension that cuts across six distinct research domains: Imaging & Acquisition, Data Reduction & Optimization, Memory Forensics, Mobile & IoT Forensics, Multimedia Forensics, and Frameworks, Standards & Privacy. These domains represent parallel research streams within the broader digital forensics discipline, rather than sequential steps of a standardized investigative workflow.
4.1. Mapping Overview and Corpus Distribution
The review protocol, information sources, search strategy, eligibility criteria, and the domain-classification procedure that produced the analytical corpus are reported in
Section 3. Following classification, each included study was assigned to a single dominant domain to prevent the same study from inflating multiple categories. The distribution of the 82 included studies across the six domains is shown in
Figure 2: imaging and acquisition is the largest domain (18 studies), followed by frameworks, standards and privacy (15), data reduction and optimization (14), multimedia forensics (14), mobile and IoT forensics (13), and memory forensics (8).
Their dispersion and longitudinal spread appear in the literature bubble map of
Figure 3, which reveals a pronounced concentration of recent activity: the 2024–2025 cluster is dominated by multimedia forensics and by frameworks, standards and privacy, while the earliest intelligent contribution in the corpus dates to 2007.
Most importantly for this review’s central question, intelligence adoption is markedly uneven. Of the 82 included studies, 38 employ an intelligent or automation component while 44 rely on classical techniques; as
Figure 4 shows, this near-even aggregate conceals a sharp divide, from 93% intelligence adoption in multimedia forensics and 77% in mobile and IoT forensics, down to 33% in frameworks, standards and privacy, 28% in imaging and acquisition, 25% in memory forensics, and 21% in data reduction and optimization. Put together, these establish the contextual foundation for the domain-specific analyses that follow.
4.2. Intelligence Maturity Across Domains
Resolving intelligence adoption onto the maturity ladder (
Figure 5) reveals structure that a binary classification obscures. The domains are not simply more or less intelligent; they occupy distinct maturity signatures.
Table 4 summarises each domain’s adoption profile, maturity ceiling, and characteristic gap.
Imaging and acquisition, despite five non-classical studies, never advances beyond classical machine learning: its intelligent share is four automation studies and a single machine-learning study, with no deep-learning or reasoning-based work at all. Multimedia forensics shows the opposite profile, bypassing the lower rungs almost entirely to concentrate at Levels 3 and 4, with ten deep-learning and three explainable studies against a single classical one. Mobile and IoT forensics clusters at Level 2, its six classical-machine-learning triage studies marking a mature but not deep adoption. Frameworks, standards and privacy displays a barbell: automation at Level 1 and reasoning-based knowledge-graph and explainable methods at Level 4, with the machine-learning and deep-learning middle almost empty, consistent with a governance domain that favors inspectable reasoning over statistical pattern recognition. Memory forensics remains overwhelmingly classical, its only two intelligent studies sitting at Levels 3 and 4 and separated by nearly two decades. The uneven penetration reported above is therefore not only a matter of degree but of kind: domains differ in the maturity of intelligence they have reached and, in several cases, in the rungs they have skipped.
4.3. Imaging & Acquisition
4.3.1. Overview
Imaging and acquisition remain foundational to digital forensics, defining how evidence is captured, preserved, and validated for later analysis. Historically, imaging relied on exhaustive bit-stream duplication to guarantee completeness and reproducibility, a practice clearly articulated by Horsman [
1]. As storage capacities and system complexity have grown, this once reliable approach has become increasingly inefficient, prompting substantial work on selective and targeted acquisition methods that determine what to capture, how quickly, and to what extent based on evidential relevance.
It is important to characterize this literature precisely. Although often described as “intelligent,” the selective imaging tradition is overwhelmingly rule-based and heuristic rather than learning-driven. Indeed, among the six domains examined in this review, imaging and acquisition is both the largest (18 included studies) and the least transformed by intelligent methods: only five studies employ any non-classical technique, and four of those focus on automation rather than machine learning. Imaging, therefore, stands as the clearest example of a mature, high-volume forensic domain in which classical, deterministic methods continue to dominate.
4.3.2. Challenges
Modern acquisition environments introduce technical, procedural, and ethical constraints. Freiling et al. demonstrated that solid state drives complicate repeatable imaging because TRIM, wear-leveling, and garbage-collection mechanisms can unpredictably erase residual data [
4]. Cloud and virtualized infrastructures further fragment evidence across multi-tenant environments, limiting an examiner’s control over source data, as shown by Alghamdi et al. [
22]. Remote access platforms add another layer of complexity by dispersing artifacts across synchronized devices, a challenge highlighted by Soni et al. [
23].
Investigators also face bottlenecks created by massive storage capacities, constrained acquisition bandwidth, and privacy regulations that increasingly restrict indiscriminate collection. These pressures explain why the domain has invested so heavily in selective acquisition, but, as the next section shows, that investment has largely taken the form of engineering and automation rather than learning-driven approaches.
4.3.3. Selective and Automated Approaches
The dominant response to acquisition-scale pressure has been selective imaging: capturing only evidentially relevant regions. This is a rich but fundamentally classical lineage. Turner introduced selective and targeted imaging through Digital Evidence Bags [
24]; Grier and Richard accelerated imaging of large disks using sifting collectors [
25]; Halboob et al. formalized a selective-imaging model [
26]; and Faust et al. extended selective imaging to live systems [
27]. More recent engineering continues the same deterministic tradition: Nugroho and Amiruddin built a low-cost, hash-verified selective-imaging device [
2], and Ozcan et al. used Spark-based parallelism to accelerate artifact extraction [
3]. Supporting tooling and formats, including the AFF image format [
28], forensic XML and metadata representations [
29,
30], differential analysis [
31], and NTFS-level analysis [
32,
33], likewise operate on deterministic rules.
Genuinely intelligent contributions are the exception rather than the rule. Four studies introduce automation without learning: automated selective capture [
24,
25], automated pipeline processing [
34], and automated evidence-set generation [
35]. Only one study in the entire domain applies a learning method: Alqahtany et al. used cluster analysis to guide forensic acquisition of non-volatile memory in cloud (IaaS) environments [
36]. Privacy and standards-oriented acquisition work also exists, for example, encrypted selective imaging by Karadsheh [
6], but its contribution is governance rather than learning, and it is discussed within the frameworks domain. Finally, Voigt et al. cautioned that even these efficiency gains must be evaluated under realistic conditions, since synthetic benchmarks can distort performance claims [
37].
4.3.4. Integration Challenges and Future Outlook
The near absence of learning in imaging is best understood not as a deficiency in the existing work, which is methodologically strong, but as a clear research opportunity. Selective acquisition currently decides what to capture using fixed rules and predefined profiles; a natural next step is learning-based relevance prediction that adapts to case type and prior investigations, supported by explainable AI modules that justify why particular regions were captured or excluded. Realizing this vision will require standardized benchmark datasets and evaluation metrics that the domain still lacks, a gap highlighted by Voigt et al. [
37]. It will also require privacy-aware design so that adaptive selection remains legally defensible.
In short, imaging and acquisition is the domain where intelligent methods have the furthest still to travel, and given its foundational role in digital forensics, it is also the domain with the most to contribute once those methods mature.
4.4. Data Reduction & Optimization
4.4.1. Overview
Data reduction and optimization have become essential as investigators confront unprecedented volumes of evidence. With expanding storage capacities and increasingly interconnected systems, forensic teams are routinely buried under terabytes of repetitive or irrelevant data, slowing analysis and stretching infrastructure to its limits. The dominant response has been to preserve evidential completeness while minimizing processing overhead through deduplication, similarity-based filtering, and triage.
It is worth stating plainly, however, that most of this machinery is algorithmic and deterministic rather than learning-driven. Among the 14 included studies, only three employ a genuinely intelligent method. Data reduction is therefore, like imaging, a domain whose core techniques remain classical even as the volume pressures that motivate them intensify. And because reduction decisions reshape the evidentiary record, they carry epistemic risk that must be managed explicitly rather than assumed away, this is a concern amplified in cross-jurisdictional settings where data minimization and privacy preservation apply, as Karadsheh et al. emphasize [
6].
4.4.2. Challenges
The foremost challenge is maintaining evidential integrity when data are filtered, compressed, or deduplicated. Lanterna and Barili showed that file-system level deduplication can fragment evidence into chunks whose reconstruction depends on intact hash mappings; when those mappings are lost, investigators risk reconstructing files that never existed [
38]. Du et al. demonstrated that deduplicated acquisition can yield substantial storage and speed gains, but only under high duplication ratios and centralized infrastructure [
39]. Shayau et al. observed that incomplete baseline libraries may cause reduction engines to exclude modified or hidden artifacts [
40].
These are challenges of algorithm design and coverage, not of model training, which is an instructive reminder of how classical the domain’s foundations remain.
4.4.3. Reduction Techniques and the Role of Intelligence
The bulk of the domain comprises deterministic reduction methods. Deduplication-based acquisition, as demonstrated by Du et al. [
39], and indexed-baseline reduction approaches such as DIFReM by Shayau et al. [
40], shrink evidence sets using fixed rules and reference libraries. Similarity-based content triage, context-triggered piecewise hashing, and similarity digests, introduced and refined by Roussev and Quates [
41], likewise reduce volume through deterministic comparison rather than learning. This body of work is effective and widely used, but it does not learn.
Against this classical backbone, three studies introduce genuine intelligence, and all three do so through machine-learning-based triage rather than through reduction algorithms themselves. Del Mar-Raave et al. used a machine-learning image classifier to identify evidentially relevant media within large seized collections, reducing analyst review burden [
42]. Du and Scanlon [
43] automated metadata-based classification of incriminating artifacts, and Roussev et al. advanced real-time triage that prioritizes material for examination as it is acquired [
44]. Notably, all three operate at the triage/prioritization layer, deciding what to look at first, rather than within the reduction mechanics themselves, which remain classical.
Broader machine learning-driven orchestration of triage and correlation has also been proposed within incident-response frameworks, most prominently by Kuforiji [
16], though that contribution is discussed within the frameworks domain.
4.4.4. Integration Challenges and Future Outlook
The narrow footprint of learning in this domain is itself a finding: reduction decisions, that is, what to keep, are beginning to be learned, but reduction mechanisms, in the form of how to compress or deduplicate them, remain almost entirely rule-based. The clearest opportunity is therefore to bring learning into the reduction process itself, enabling models that adapt to case type and prior investigations, paired with explainable AI modules that allow investigators to justify why specific data were excluded or compressed.
Realizing this vision depends on standardized benchmark datasets and shared fidelity metrics that the domain currently lacks is a gap underscored by Voigt et al. [
37]. It also requires auditable mechanisms that keep learned reduction defensible in court, ensuring that adaptive methods do not compromise evidential soundness or legal acceptability.
The long-term goal is a reduction ecosystem that learns from prior investigations while producing lean, relevant, and demonstrably sound evidence sets, in an evolution that would shift data reduction from a purely deterministic practice to an adaptive, accountable, and intelligence-supported component of modern digital forensics.
4.5. Memory Forensics
4.5.1. Overview
Memory forensics addresses the acquisition and analysis of volatile system state, such as running processes, loaded modules, memory-mapped files, and application traces that exist only in RAM and vanish at power-off. As adversaries increasingly rely on fileless malware, in-memory payloads, and anti-forensic techniques that leave minimal disk residue, volatile memory has become one of the few places where such activity remains observable at all. Within our corpus, however, memory forensics is also the domain where intelligent methods have penetrated the least: of the eight included studies, only two employ a learning-based method, giving the lowest intelligence share (25%) of any of the six domains. The domain’s core machinery remains, by both necessity and convention, a deterministic one.
4.5.2. Challenges
Memory forensics operates under constraints that are unusually unforgiving. Volatile evidence offers a single acquisition opportunity: memory state changes continuously on a live system, and cannot be re-captured for verification once the machine powers down, so any error or omission at acquisition time is permanent. Dangi et al. highlighted this pressure in the context of live systems, where investigators must decide in real time what to preserve [
5]. Interpretation is equally fragile. Reliable analysis depends on reconstructing operating system data structures whose layouts shift across versions and builds, making tooling perpetually vulnerable to obsolescence, a dependence that Prem et al. made visible through comparative assessment of the framework tooling on which analyses rest [
45]. Adversaries exploit precisely this layer: Block showed that memory-mapped image files can be maliciously modified in ways that mimic legitimate state [
46], while modern applications scatter their traces across live memory, disk, and synchronized services, forcing examiners to correlate across sources [
23]. Finally, placing recovered in-memory traces on a defensible timeline remains difficult, as volatile artifacts rarely carry the temporal anchors that disk artifacts provide [
47]. Together, these pressures explain the domain’s conservatism: when evidence is unrepeatable and structures are exact, approximate methods carry a cost that other domains do not face.
4.5.3. Classical Approaches: Structured State Analysis
Memory forensics is “classical” for a simple reason: the work demands exactness. Investigators must rebuild operating system structures exactly as they appear in RAM, down to the byte and the specific OS version. Anything approximate or probabilistic risks misinterpreting a volatile state, and that level of uncertainty is unacceptable in a forensic setting. The six classical studies in our corpus all follow this strict, precision-driven model because it is the only approach that reliably preserves evidential integrity. At the core sits a structured analysis of Windows volatile state. Lapso et al. developed whitelisting approaches that model known good system state so anomalous entries stand out in forensic memory visualizations [
48]. Block identified malicious modifications of memory-mapped image files through systematic structural comparison [
46]. Around this core, the remaining studies extend memory analysis toward adjacent investigative needs: Soni et al. recovered application-level artifacts of remote-access tooling by combining live, disk, and memory examination [
23]; Dangi et al. used memory forensics to drive selective imaging of live systems [
5]; Prem et al. conducted comparative assessments of memory-forensics framework tooling [
45]; and Weyermann et al. situated recovered digital traces within a common temporal frame for event reconstruction [
47]. What ties this body of work together is its commitment to determinism, and for good forensic reasons. An examiner must be able to show, step by step, how an in-memory artifact was recovered and interpreted, and a court must be able to trust that explanation. Methods that rely on approximation or probabilistic inference do not offer that level of transparency, which is why they fit poorly at this layer of the forensic stack.
4.5.4. Intelligent Approaches: An Early Start, a Long Stall
What makes the memory domain analytically distinctive is not merely that intelligent methods are rare, but that its two intelligent studies sit at the chronological extremes of our entire corpus, with nothing between them. The earliest intelligent study across all six domains belongs here: Khan et al. applied neural networks to post-event timeline reconstruction as far back as 2007 [
49], demonstrating that learned models could infer event sequences from system state well before the current wave of deep learning enthusiasm. The thread then goes silent for nearly two decades before resurfacing in recent work applying unsupervised learning, specifically clustering and anomaly detection with SHAP-based explanations to memory analysis [
50]. This resurgence is explicitly motivated by two long-standing constraints: the scarcity of labeled forensic memory data and the need for explainability strong enough to satisfy courtroom admissibility. The eighteen-year gap between these studies is itself a finding. It suggests that the barrier has never been awareness but fit: both intelligent contributions operate at the interpretive layer, inferring event sequences and surfacing anomalies, while the extraction machinery beneath them remains, and arguably must remain exact. This division of labor mirrors the decisions-versus-mechanisms distinction observed in data reduction: intelligence can guide what an analyst pays attention to, but the evidentiary mechanisms themselves stay classical.
4.5.5. Integration Challenges and Future Outlook
The obstacles to broader intelligent adoption here are concrete. Memory captures are large, unlabeled, and highly version-dependent, which means training data that generalizes across operating system builds is scarce; ground truth requires expert annotation of raw structures; and the court-facing requirement of explainability weighs even more heavily on volatile evidence, which cannot be re-acquired for verification. Yet these same constraints also define the opportunity. The corpus’s most recent intelligent study points toward a viable path: unsupervised and explainable methods that operate on features extracted by mature deterministic tooling [
50]. Such approaches can detect injected code, fileless payloads, or anomalous system state without displacing the sound acquisition and parsing layer beneath them.
Memory forensics is therefore the domain where the gap between intelligent potential and intelligent practice is widest, and where the 2007 precedent [
49] shows that the direction was recognized long before the field was ready to pursue it.
4.6. Mobile & IoT Forensics
4.6.1. Overview
Mobile and Internet-of-Things (IoT) forensics deal with an ecosystem of interconnected devices that continuously generate heterogeneous, distributed, and often short-lived data. Evidence is frequently dispersed across mobile file systems, app sandboxes, volatile memory, network interactions, and cloud-synchronized repositories. This dispersion fundamentally limits post hoc reconstruction, as short-lived artifacts, selective logging, and cloud-mediated interactions often eliminate the possibility of complete retrospective acquisition [
51,
52]. Surveys of the field have documented how increasing device diversity, encryption, and rapid OS evolution complicate acquisition and analysis, motivating more adaptive and intelligent tooling [
9]. Among the six domains examined, mobile and IoT forensics is the second most transformed by intelligent methods (10 of its 13 included studies). Unlike the deep learning saturation seen in multimedia forensics, however, intelligence here is dominated by machine learning-based triage and automation, a lineage that extends back more than a decade.
4.6.2. Challenges
Mobile and IoT investigations face multi-layered challenges. Traditional extraction struggles with device heterogeneity, proprietary app structures, and inconsistent logging formats; fragmentation across Android and iOS ecosystems creates acquisition gaps when tools cannot keep pace with rapid OS updates [
9]. Boztas et al. showed that app-level evidence is difficult to reconstruct because artifacts are scattered across internal databases, logs, caches, and configuration files protected by sand-boxing [
53]. In IoT environments, Quick et al. demonstrated that datasets are frequently incomplete due to selective logging, limited onboard storage, and volatile network-driven interactions [
51].
The combined effect is a forensic landscape defined by high volume, high heterogeneity, and low persistence, precisely the conditions under which manual analysis breaks down and learning-based triage becomes attractive.
4.6.3. Classical Approaches: Deterministic Reduction and Volatile Evidence
Before turning to the intelligent majority, the domain’s classical stream deserves its own account, both because it remains operationally useful and because it provides the baseline against which the intelligent shift is measurable. Quick and Choo developed selective IoT data reduction procedures that shrink heterogeneous device datasets into manageable evidential subsets through deterministic filtering rather than learned relevance [
51]. Zhang et al. analysed volatile evidence from IoT botnet infrastructure, reconstructing attack activity from short-lived network and device traces through structured, rule-driven examination [
52].
These studies share the same character seen in classical imaging and memory forensics work: sound, repeatable, and explainable procedures whose limitations emerge only at scale. In a domain defined by high volume, high heterogeneity, and low persistence, deterministic pipelines cannot adapt to unseen app structures or meaningfully prioritize among thousands of candidate artifacts. That inflexibility is precisely the gap the intelligent stream is designed to address.
4.6.4. Intelligent Approaches and Emerging Opportunities
The included studies concentrate on three intelligence-driven sub-streams.
Machine learning triage and evidence prioritization: This is the domain’s most established intelligent line, focused on automating the categorization and prioritization of relevant evidence. Marturana and Tacconi first introduced a machine-learning triage methodology for automated media categorization [
54], and later, Marturana et al. applied a quantitative, feature-based approach specifically to mobile forensics [
55]. Extending this to modern casework, Serhal and Le-Khac used machine learning over file metadata to triage smartphone evidence at scale in real investigations [
56]. Verma et al. generalized the idea into the Df 2.0 framework, coupling machine learning evidence prediction with privacy evaluation and linking triage with governance concerns [
57].
Automated extraction and specialized inference: Researchers increasingly rely on automation and machine learning to handle the messy, fast-moving nature of mobile evidence. Boztas et al. showed that automated, interaction-driven app analysis can surface files that never appear in a standard static acquisition [
53]. Bhattarai et al. combined deep learning and NLP to triage cryptocurrency wallet artifacts across mobile platforms [
58], while Peng et al. used large-scale clustering to impose structure on huge, heterogeneous evidence sets [
59]. In a different direction, Freire-Obregon et al. demonstrated that deep learning can link captured media back to the specific handset that produced it, a direct attribution capability with clear investigative value [
60].
Edge and IoT intelligence: Similar shifts are emerging in edge and IoT environments. Kshirsagar et al. integrated generative AI and deep-learning modules into a portable forensic device capable of reconstructing fingerprints, detecting weapons, and interpreting activity on-scene, illustrating how intelligent pre-processing can accelerate field investigations [
61].
The emerging trajectory shows a shift toward context-aware extraction, machine learning-driven triage, and cross-platform correlation, while classical approaches remain appropriate in areas where deterministic, easily explainable procedures are still essential.
4.6.5. Integration Challenges and Future Outlook
Integrating intelligent analytics across heterogeneous mobile and IoT environments remains difficult. Models trained on one device family or OS variant often fail to generalize due to proprietary firmware, inconsistent logging, and rapidly evolving app ecosystems. Approaches such as on-device learning and federated learning introduce further complications: keeping models synchronized, validating updates, and maintaining an auditable evidential trail are all non-trivial. Standardization also lags; there is still no widely accepted schema for representing mobile or IoT forensic artifacts in a consistent, interoperable format.
Looking ahead, progress depends on shared ontologies, cross-device benchmarking datasets, and explainable AI layers that allow investigators to justify automated findings in a legally defensible manner. The longer-term vision is an ecosystem in which intelligent edge devices collaborate with central analytical engines to deliver real-time, privacy-aware investigative insight without compromising forensic rigor.
4.7. Multimedia Forensics
4.7.1. Overview
Multimedia forensics focuses on the authentication, attribution, and interpretation of visual and audio evidence drawn from cameras, mobile devices, sensors, and synthetic media generators. As investigations increasingly rely on digital photographs, video, and audio, and as generative models make convincing manipulation trivial, establishing the integrity and provenance of these traces has become essential. Berube et al. [
62] showed that even technically accurate multimedia evidence can be challenged when reconstruction steps or enhancement procedures are poorly documented. They further demonstrated that in a real trial setting, an opaque tooling can undermine the interpretation of visual evidence. Reflecting the pace of change, Singh et al. [
10] and Klasen et al. [
63] emphasized that modern AI-driven manipulation now requires detection and attribution methods of comparable sophistication. Of the six domains examined here, multimedia forensics is the most thoroughly transformed by intelligent methods, with 13 of its 14 included studies employing a substantive learning-based technique.
4.7.2. Challenges
Multimedia evidence presents challenges of data quality, deliberate manipulation, and contextual ambiguity. Visual and audio data are often degraded by compression, filtering, or platform-specific transformations that obscure the forensic signatures on which analysis depends. Modern deep-learning manipulation, including face-swapping, splicing, and fully synthetic generation, further complicates source identification and authenticity verification, as shown by Singh and Sehgal [
10]. Anti-forensic operations add another layer of difficulty. Bhatia et al. [
64] demonstrated how quantization footprints can be intentionally altered to hide tampering, and this classical anti-forensic line of work directly motivates the learning-based detectors discussed below. The central tension is that the same learning-based techniques used to create and conceal manipulation now require learning-based methods for credible detection.
4.7.3. Intelligent Approaches and Emerging Opportunities
The included studies cluster into four intelligence-driven sub-streams.
Forgery detection and localization: Recent methods move beyond binary classification to localize tampered regions and produce output that investigators can use directly. Yang et al. [
65] proposed a dual-encoder network that preserves non-semantic forensic fingerprints to detect and localize image splicing while resisting anti-forensic attack. Diwan and Roy [
66] combined keypoint features with a CNN to localize copy-move forgeries under post-processing. Sharma et al. [
67] paired a CNN with a transformer to generate tampered-region masks. Munawar and Oussalah [
68] introduced an explainable dual-stream attention mechanism, and Bamigbade et al. [
69] framed manipulation detection for digital-forensic use through vision-attention anomaly scoring.
Deepfake detection: Ciftci et al. [
70] exploited physiological signals that synthetic video fails to reproduce. Nguyen et al. [
71] applied capsule networks to forgery detection. For evidential use, Bharati et al. [
72] coupled deepfake detection with explainable-AI reasoning oriented toward legal investigation, addressing the admissibility gap that pure detection accuracy leaves open.
Source attribution: Determining which device produced an image is inherently evidential. Klier and Baier [
73] critically evaluated deep-learning source camera identification against forensic expectations. Bennabhaktula et al. [
74] identified camera models from forensic traces in homogeneous image regions. Manisha et al. [
75] learned device-specific fingerprints that remain robust under counter-forensic manipulation.
Audio and reconstruction: Extending multimedia forensics beyond the visual, Shaaban and Yildirim [
76] applied deep learning to detect synthetic and manipulated speech, an increasingly important capability as voice-based fraud and fabricated audio evidence proliferate. Complementing detection, Sharma et al. [
77] proposed an end-to-end reconstruction pipeline that integrates super-resolution, denoising, and de-blurring with OCR and contextual analysis to recover intelligible detail from degraded imagery.
The emerging trajectory shows a move from static inspection to learning-based authentication, attribution, and reconstruction. Throughout this evolution, explainability remains the decisive criterion for evidential use, distinguishing forensic-grade methods from conventional computer-vision models.
4.7.4. Integration Challenges and Future Outlook
Despite this maturity, operational integration remains difficult. Deep models require large, ethically sourced datasets that remain scarce in investigative contexts. As the attribution and deepfake studies repeatedly show, explainability is essential because courts expect transparent reasoning to accompany technical findings, especially when AI-generated enhancements or classifications influence outcomes, as demonstrated by Berube et al. [
62] and Bharati et al. [
72]. The domain’s only non-intelligent study is the anti-forensic technique introduced by Bhatia et al. [
64], underscoring how fully the field has pivoted toward learning-based methods. The remaining open problems now concern standardization, benchmarking, and evidential defensibility rather than whether intelligence should be adopted.
Future work should establish interoperable benchmarks and domain-specific evaluation metrics, and develop explainable AI layers that link automated outputs to defensible evidential narratives. The goal is to ensure that computational gains translate into court-ready reliability, allowing intelligent multimedia-forensic systems to operate with both technical sophistication and legal robustness.
4.8. Frameworks, Standards & Privacy
4.8.1. Overview
This domain encompasses the procedural, governance, and standards layer of digital forensics: investigation process models, the application and comparison of formal standards, admissibility and quality assurance, and the growing tension between investigative reach and privacy rights. It is the connective tissue that turns technical capability into a legally defensible practice. It is also, on our evidence, a predominantly classical domain: of the fifteen included studies, only five involve an intelligent or automation component, and just three of those employ a genuine learning or reasoning method. In other words, the governance layer of digital forensics has begun to talk about intelligence far more than it has begun to use it.
4.8.2. Challenges
The pressures on this domain are procedural rather than computational, but they are no less real. Formal standards struggle to keep up with the environments they are supposed to govern. Comparative reviews show that even the two main instruments, NIST SP800-101 [
78] and ISO/IEC 27037 [
79], leave gaps and overlaps [
80], and every new investigative setting, from public clouds to building automation systems to schema-less databases, exposes assumptions that those standards never anticipated [
81,
82,
83]. Quality assurance is under similar strain. Missed opportunities in real investigations often come down to process failures rather than technical limits [
84], and courtroom experience shows that even technically sound digital evidence can falter when procedures or tooling are poorly documented [
62]. At the same time, privacy obligations increasingly restrict what investigators can collect, retain, or process at all, putting investigative reach and individual rights in direct tension [
6].
The result is a domain that has to keep reinterpreting its own rule book while the evidence landscape, and now the analytical tools themselves, shift beneath it.
4.8.3. Classical Approaches: Standards in Practice and Governance Under Pressure
The classical core of this domain falls into two streams, each responding to the pressures outlined above. The first deals with standards in practice: instead of proposing new formal standards, these studies test, compare, and adapt existing instruments to new environments and offense types. The comparative review of NIST SP800-101 and ISO/IEC 27037 turns gap finding into practical guidance [
80], while the ACPO principles are stress-tested in public-cloud investigations [
81] and building-automation systems [
82]. In the same spirit, an ISO/IEC 27037-aligned six-phase methodology operationalizes standardized hard disk acquisition [
85]; offence-specific standardization is developed for Windows 10 examinations in child sexual abuse cases [
1]; and a dedicated forensic process is constructed for wide column NoSQL databases [
83]. Across this stream, the pattern is consistent: general-purpose standards need substantial, environment-specific interpretation before they become usable, and doing that interpretive work is itself a meaningful research contribution.
The second stream addresses governance under pressure: what investigation quality, legality, and legitimacy require as both evidence and analytical tooling grow more complex. Case study analysis of a criminal trial shows how digital traces succeed or fail once they reach court [
62]; an examination of missed opportunities in real investigations traces quality failures to process rather than technology [
84]; and privacy-preserving investigation frameworks confront the growing tension between investigative needs and user rights [
6]. Notably, this stream now includes governance of artificial intelligence itself: a responsible AI framework for forensic science [
86] sets out principles for adopting AI defensibly, while employing no intelligent method of its own. That inversion captures the domain’s current stance toward intelligence, which is more regulatory and anticipatory than operational.
4.8.4. Intelligent and Automated Approaches
The five studies that incorporate intelligence or automation fall into the same pattern we saw in imaging: automation comes first, learning comes second. On the automation side, the field is still working out what “automation” in digital forensics should actually mean, with early formal definitions now emerging [
87]. We also see practical engineering work, like real-time forensic collection for containerized orchestration environments using eBPF instrumentation [
88]. These are valuable operational contributions, but they are still rule-driven rather than learned.
The genuinely reasoning-based work is newer, and it is the part that feels most distinctive. We now have knowledge-graph approaches that make evidential reasoning explicit and navigable [
89], a schema for representing forensic artifacts and their relationships [
90], and an NLP-driven pipeline that automates artifact extraction and reporting for Electron-based social applications [
91]. It is telling that when “intelligence” enters this domain, it comes in the form of structured reasoning and representation, with knowledge graphs and language processing taking the place of the statistical pattern recognition that dominates multimedia and mobile forensics. Governance work demands traceable inference, and the intelligent methods gaining traction here are exactly the ones whose reasoning can be inspected.
4.8.5. Integration Challenges and Future Outlook
For intelligent methods specifically, the barriers described above compound: a domain whose standards interpretation is still manual and jurisdiction-specific, and whose admissibility doctrine for machine-generated findings remains unsettled, has little institutional room to adopt methods it cannot yet govern. The small “intelligent” corner of this literature nonetheless hints at a realistic path forward. Knowledge graphs and other explainable reasoning structures [
89,
90] could grow into the connective tissue that makes intelligent outputs from other domains auditable and fit for courtroom use. At the same time, responsible AI governance [
86] is beginning to move from broad principles toward enforceable procedure. If intelligence is going to spread beyond multimedia and mobile forensics, this domain will not just document that shift, but will rather have to authorize it.
5. Discussion
The six domains reviewed in
Section 4 show how intelligence is reshaping digital forensics, while also revealing several cross-cutting issues that reach beyond any single technique or research stream. Across imaging, reduction, memory analysis, mobile and IoT investigation, multimedia interpretation, and governance frameworks, common themes emerge around interoperability, evidential completeness, explainability, and privacy. This section considers these themes at a conceptual level, emphasizing broad expectations rather than prescriptive formulas or domain-specific measurement schemes.
5.1. Fragmentation and Lack of Interoperability
A consistent pattern across the literature is fragmentation. Intelligent tools tend to evolve as solutions to specific problems, which are faster acquisition, smarter triage, and improved visual reconstruction, without aligning to shared formats, schemas, or validation protocols. This leads to a collection of promising but isolated components rather than an integrated ecosystem. Frameworks such as [
86], privacy-preserving workflows [
6], big-data architectures [
18,
19], and structured investigative models like D4I [
20] demonstrate that more unified approaches are possible. However, these efforts remain unevenly applied, and no single architectural model can suit all investigative or jurisdictional contexts. A general expectation moving forward is that intelligent digital forensics systems should be designed with interoperability and cross-domain compatibility in mind, even if full unification is neither realistic nor desirable.
5.2. Efficiency Versus Evidential Completeness
Intelligent methods also reveal a recurring tension between efficiency and completeness. Techniques such as selective imaging, adaptive data reduction, or edge-based IoT triage can reduce volume, speed up processing, and alleviate analyst workload. However, these gains may carry the risk of omitting contextual or residual artifacts, particularly in volatile environments such as memory or distributed IoT ecosystems. Given the absence of universal rules for how much filtering is “acceptable,” the expectation is not to enforce a specific quantitative balance. Instead, practitioners should remain mindful of what may be gained or lost when adopting automated or selective approaches and ensure that efficiency does not inadvertently compromise evidential coverage or legal defensibility.
5.3. Explainability, Transparency, and Trust
Explainability is becoming central to intelligent digital forensics, yet it remains difficult to apply in a uniform way. Ref. [
17] emphasized that investigators must be able to interpret and justify algorithmic outputs, especially when models infer relevance, detect anomalies, or reconstruct media. Ref. [
14] further highlighted concerns about algorithmic bias, opaque decision pathways, and limited transparency in automated workflows. There is no single explainability model that fits all digital forensics contexts, and expectations differ across legal systems and investigative environments. As a result, intelligent tools should aim to provide interpretable and reviewable outputs where feasible, while recognizing that explainability requirements will vary by jurisdiction, tool purpose, and case-specific constraints.
5.4. Privacy, Legal Constraints, and Ethical Accountability
Privacy concerns permeate mobile, IoT, multimedia, and cloud investigations. Ref. [
6] demonstrated that privacy-aware approaches such as selective imaging and encrypted deduplication can reduce unnecessary exposure of personal data. However, privacy regulation remains heterogeneous: different jurisdictions enforce different rules on what can be collected, how long it may be retained, and what forms of automated reasoning are permissible. Frameworks like [
86] offer structured guidance for responsible AI adoption, but implementation depends heavily on institutional policy, resource capacity, and local law. Thus, rather than advocating for a single privacy mechanism, the broader expectation is that intelligent digital forensics systems should incorporate mechanisms that align with ethical norms and legal constraints relevant to their operational context.
5.5. Standardization and Validation Expectations
In a consistent practice with PRISMA-ScR reporting guidance, this review addresses the assessment of study quality and potential sources of bias through a structured qualitative evaluation, rather than the application of a formal risk-of-bias scoring instrument. Given the heterogeneous nature of intelligent digital forensics research, spanning conceptual frameworks, tool implementations, experimental evaluations, and domain-specific case studies, it will not be appropriate to use a single quantitative bias assessment tool. Across the reviewed literature, a recurring challenge is the absence of shared validation practices and commonly accepted quality baselines for intelligent digital forensics workflows. Evaluation methodologies are often developed independently within specific subdomains or for individual tools, which limit comparability across studies and complicate the assessment of methodological robustness and evidentiary reliability. To support a transparent and defensible assessment of methodological quality, this review adopts a qualitative appraisal strategy grounded in widely recognized data-quality principles. Specifically, the Dimensions of Data Quality (DDQ) framework proposed by Black et al. [
92] was used as a conceptual reference to guide evaluation. The DDQ framework outlines qualitative dimensions including accuracy, completeness, consistency, uniqueness, timeliness, interpretability, and integrity that are commonly used to assess the fitness and trustworthiness of data across analytical domains. Although DDQ was not developed specifically for digital forensics, these dimensions align closely with established forensic expectations.
Each study included in the review was assessed using the six DDQ-aligned criteria shown in
Table 5. For every criterion, studies were judged as having strong, moderate, or limited alignment based on how clearly they described their methods, how comprehensive their evaluations were, and how transparently they reported their validation procedures.
Table 5 serves as a structured guide for assessing quality and potential bias, not as a scoring tool. Its purpose is to support transparent reviewer reasoning and highlight the most common methodological weaknesses found across the studies.
The six quality dimensions (accuracy, completeness, consistency, interpretability, uniqueness, and integrity) are adapted from the data-quality literature and chosen because a forensic method is a pipeline whose output must be fit for evidential use: accurate in extraction, complete in evidential coverage, consistent across tools, interpretable to analysts and courts, sound in artifact differentiation, and integrity-preserving under automated workflows. Unlike risk-of-bias tools for intervention studies, they accommodate the corpus’s methodological heterogeneity while surfacing the concerns that bear on admissibility. Alignment with each dimension was recorded during full-text appraisal as strong, moderate, or limited (
Table 6); because appraisal was performed by a single reviewer, these levels are interpretive and descriptive, informing the cross-domain discussion of weaknesses rather than ranking or excluding studies.
Two limitations of the review process itself should be noted. Screening was performed by a single reviewer, which introduces a risk of selection bias; this was mitigated by recording criterion-based decisions for every record and by consensus resolution of uncertain cases. We did not compute formal inter-coder reliability statistics (such as Cohen’s kappa on an independently double-coded subset). The second author’s role was to collaboratively review the criteria and coding decisions rather than to perform a fully independent duplicate coding pass. We note this as a limitation of the single-coder design, though it is partly mitigated by the fully specified and reproducible coding scheme described in
Section 3.7.
Because eligibility required at least some thematic engagement with automation or intelligent methods (IC3), the adoption shares we report reflect patterns within this automation-oriented corpus rather than each domain’s entire research output. All per-study screening records and quality-appraisal notes are available from the authors upon request.
5.6. Temporal and Methodological Shifts Across Domains
Beyond cross-cutting themes, the mapped literature also suggests domain-level differences in publication span and validation style. Imaging and acquisition studies appear across nearly the full review window (imaging studies span 2006–2025, within the 2006–2026 review window), reflecting a long-standing research focus driven by changing storage technologies and evolving acquisition constraints. This domain also includes substantial empirical work evaluating acquisition behavior, integrity controls, and practical tooling under realistic conditions.
This pattern is visible in
Figure 6: intelligent contributions, sporadic before 2019, account for the majority of the corpus’s most recent additions, while classical contributions persist at a steady rate throughout the review window.
In contrast, memory forensics shows a stronger concentration of contributions in recent years (particularly from around 2020 onward), consistent with the rising relevance of volatile evidence in modern threat models, including memory-resident and file-less techniques, as well as increased containerization and cloud-hosted execution. Framework and governance research, including privacy and accountability-oriented proposals, more frequently emphasizes implementation-based validation, reflecting the difficulty of evaluating policy and compliance mechanisms through purely empirical datasets. These temporal and methodological differences matter because they shape how ready a domain’s intelligent techniques are for operational deployment. They also motivate a forward-looking research direction: bridging empirically grounded digital forensics techniques with deployable, auditable governance structures so that intelligent automation can scale without eroding transparency or legal defensibility.
5.7. Toward an Integrated Intelligent Digital Forensics Outlook
Across all six domains, the long-term opportunity lies not in defining a rigid, universal architecture, but in encouraging coherence between intelligent tools, digital forensics standards, and legal expectations. A more integrated outlook would link acquisition strategies with analytical reasoning, incorporate privacy safeguards into evidence handling, and embed transparency into automated workflows, while still allowing for jurisdictional and operational flexibility. Such alignment will require ongoing collaboration across digital forensics science, computer science, law, and policy. In this sense, intelligence in digital forensics should be understood as an evolving collection of capabilities shaped by shared expectations of transparency, reliability, and ethical responsibility, rather than as a fixed or prescriptive pipeline.
Beyond the machine-learning and deep-learning methods that dominate the intelligent portion of this corpus, a newer wave of foundation-model techniques is beginning to reach digital forensics and warrants attention even though it is not yet represented among the included primary studies. Large language models and vision-language models offer plausible support for automated report generation, cross-modal evidence correlation, and natural-language querying of heterogeneous artifacts; retrieval-augmented generation could ground such systems in case-specific evidence rather than parametric memory; and autonomous or agentic pipelines could coordinate the multi-tool investigations this review found largely absent. These capabilities carry risks that bear directly on evidential use: hallucination and non-determinism threaten reproducibility, opaque reasoning complicates courtroom admissibility, and the provenance of model-generated inferences is difficult to establish. The same explainability and integrity requirements that constrain current intelligent methods apply with greater force to generative systems, making their responsible adoption contingent on the governance and validation structures this review found to be underdeveloped.
The review highlights several research gaps that need focused attention. One of the most pressing issues is the lack of standardized, publicly available benchmark datasets for intelligent digital forensics, especially in areas like volatile memory, mobile and IoT telemetry, and degraded multimedia evidence. Without these resources, it remains difficult to compare studies or generalize machine-learning results. Second, most explainability work remains centered on technical mechanisms rather than on producing explanations that investigators or courts can actually use. This gap is particularly important in memory and multimedia forensics, where intelligent reconstruction can directly shape how evidence is interpreted. Third, there is very little research on intelligent chain-of-custody management, automated tool validation, or cross-tool provenance tracing, even though these needs grow more urgent as algorithmic pipelines replace manual processes. Fourth, privacy-preserving and governance-aware intelligence, such as federated learning across jurisdictions, selective or purpose-bound acquisition, and auditable decision logging, remains underdeveloped despite its clear operational importance. Addressing these gaps will require coordinated progress on shared evaluation datasets, interoperable ontologies, explainability layers aligned with evidentiary standards, and governance frameworks that support end-to-end auditing. These priorities should guide the next generation of research in intelligent digital forensics.
Table 7 illustrates this unevenness: benchmark availability broadly tracks intelligence adoption, with multimedia forensics both the richest in standardised datasets and the highest in adoption, while memory and data reduction remain both data-poor and largely classical.
The maturity roadmap in
Figure 7 draws these directions together, situating each domain at its current effective maturity and indicating the next step most likely to advance it.
6. Conclusions and Future Work
This paper examined how intelligence is influencing six major domains of digital forensics research: imaging and acquisition, data reduction and optimization, memory forensics, mobile and IoT investigations, multimedia analysis, and governance frameworks. Treating intelligence adoption as a measured variable rather than an assumption, the review yields four principal conclusions.
First, adoption is markedly uneven and now quantified: intelligent methods dominate multimedia forensics (93% of included studies) and mobile and IoT forensics (77%), yet remain the exception in governance (33%), imaging (28%), memory (25%), and data reduction (21%), where classical, deterministic methods still prevail. Second, and more fundamentally, the domains differ not only in the degree of adoption but in its kind: resolved onto a five-level maturity ladder, imaging never advances beyond classical machine learning, multimedia leaps directly to deep learning and explainable methods, mobile and IoT settle at classical machine learning, and governance exhibits a barbell of automation and reasoning with an empty machine-learning middle. Several domains have skipped rungs rather than climbing them. Third, where intelligence enters the classical-heavy domains, it enters as automation before learning, and governance research regulates intelligence, through responsible-AI and privacy frameworks, more than it operationally employs it. Fourth, intelligent contributions, though sporadic before 2019, now dominate the most recent literature, but that momentum is concentrated in two domains rather than distributed across the field.
A second observation concerns integration. Several stages of the investigative process still lack meaningful intelligence support: even where selective imaging and data reduction are well established, little of that work is intelligence-driven, and less still addresses intelligent chain-of-custody management, automated tool validation, cross-domain evidence correlation, or transparent reasoning across tools. Memory forensics and IoT investigations face volatility, heterogeneity, and scale, yet few studies explore how intelligence could support sustained monitoring, anomaly explanation, or continuous verification.
Future work should therefore focus where intelligent support remains limited: (1) intelligent orchestration that coordinates evidence collection, triage, analysis, and governance; (2) explainability that operates across tools rather than within isolated systems; (3) privacy-aware intelligence that adapts to jurisdictional and contextual constraints; (4) data-quality and validation frameworks grounded in established multidimensional principles such as accuracy, completeness, consistency, and interpretability; and (5) integrated reasoning engines capable of aggregating insights across imaging, memory, network, mobile, and multimedia analysis. These directions do not prescribe a single architecture, but they indicate where intelligence can bring coherence to otherwise fragmented processes, and where the least-matured domains identified here have the furthest to advance.
Overall, intelligence in digital forensics should be understood as an evolving set of capabilities with uneven levels of maturity across the investigative landscape. By identifying which domains are well supported by intelligent methods, which remain predominantly classical, and which have advanced in kind rather than merely in degree, this study provides a foundation for building more connected, transparent, and ethically grounded digital forensics systems, advancing intelligence not only as a technical enhancement but as a unifying perspective that strengthens investigative reliability, supports legal defensibility, and enables responsible innovation.