Multimodal Understanding: From Document Image Analysis to Large Language Models
This special issue belongs to the section "Image and Video Processing".
Special Issue Information
Dear Colleagues,
With the rapid digitalization of government, education, healthcare, finance, and scientific research, massive volumes of data are being generated in diverse formats, layouts, and languages. These documents and images contain critical structured and unstructured information. Historically, document image analysis (DIA) served as the cornerstone for extracting this information. Today, however, this field has evolved into a much broader frontier of multimodal understanding, where complex visual perception and advanced textual reasoning closely converge.
In recent years, the research paradigm has shifted significantly with the advent of deep learning, multimodal representation learning, large language models (LLMs), and vision–language foundation models. These advances have not only revolutionized traditional tasks—such as robust scene text recognition, table structure understanding, and key information extraction—but have also unlocked unprecedented capabilities in complex reasoning, generation, and visual question answering.
Despite these breakthroughs, transitioning from foundational image analysis to deploying large-scale models presents new, formidable challenges. Open issues include the parameter-efficient fine-tuning (PEFT) of LLMs for domain-specific tasks, ensuring the robustness and trustworthiness of generative models against vulnerabilities (e.g., jailbreaking), designing lightweight and efficient architectures for high-stakes deployment, and achieving reliable cross-modal and cross-domain generalization.
This Special Issue focuses on recent methodological advances, system-level innovations, and real-world applications at the intersection of visual analysis and large language models. It aims to provide a high-level academic forum for researchers and practitioners to share cutting-edge techniques ranging from advanced document and image representation to the optimization, reasoning, and alignment of LLMs. We seek contributions that address the open challenges of multimodal intelligence, particularly under realistic, large-scale, and resource-constrained deployment conditions.
We invite original research articles and comprehensive reviews addressing both theoretical and practical aspects of document image analysis.
Topics of interest include (but are not limited to) the following:
- Document layout analysis, structural parsing, and complex table structure recognition;
- large language models (LLMs) for document intelligence;
- Handwriting analysis, trajectory recovery, writer identification, and offline signature verification;
- Robust OCR for printed, handwritten, historical, and scene text in the wild;
- Multilingual and low-resource document recognition;
- Form understanding, table recognition, and key information extraction;
- Vision–language models and large language models for document intelligence;
- Document visual question answering (DocVQA);
- Retrieval-augmented generation for document analysis;
- Retrieval-augmented generation and Chain-of-Thought reasoning for complex document analysis;
- Efficient and lightweight document understanding systems;
- Edge and real-time deployment of document AI;
- Benchmarks, datasets, and evaluation protocols;
- Applications in digital archives, finance, healthcare, education, and e-government.
Dr. Hongjian Zhan
Dr. Yu-Jie Xiong
Guest Editors
Manuscript Submission Information
Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.
Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Journal of Imaging is an international peer-reviewed open access monthly journal published by MDPI.
Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 1800 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.
Keywords
- multimodal document intelligence
- document visual question answering
- large language models (LLMs)
- vision–language foundation models
- handwriting and signature verification
- parameter-efficient fine-tuning
- robust representation learning
- table and scene text recognition
- retrieval-augmented generation
- trustworthy AI and security
Benefits of Publishing in a Special Issue
- Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
- Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
- Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
- External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
- Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

