Topic Editors
Artificial Intelligence for Multimedia Analysis, Generation, and Intelligent Services
Topic Information
Dear Colleagues,
We are pleased to invite you to submit your work to this Topic titled Artificial Intelligence for Multimedia Analysis, Generation, and Intelligent Services. Our discussion is centered around a key research question: how can modern artificial intelligence and multimodal foundation models support reliable, efficient, and smart analysis as well as delivery of multimedia information and services?
Nowadays, images, videos, audio, text and their multimodal combinations have become the main carrier for information communication. Additionally, large volumes of multimedia data keep being produced by social media platforms, streaming services, autonomous equipment, medical imaging systems and various IoT devices. This Topic spans the complete technical workflow: starting from multimedia representation and comprehension; moving on to content generation, retrieval, and recommendation; and finally reaching end‑user‑facing intelligent services. Recent progress in multimodal foundation models and generative AI has greatly pushed forward cross‑modal reasoning, content creation and large‑scale vision–language analysis. Even so, real‑world practical deployment still brings many unresolved hurdles, such as model efficiency, trustworthiness, privacy risks and generalization performance.
We welcome original research articles, review papers and perspective manuscripts covering relevant theories, technical approaches and real‑world applications. The scope is grouped into five thematic clusters as follows:
- Multimedia representation and understanding: image, video, audio, and text analysis; self‑supervised learning; multimodal representation learning; and temporal and spatial modeling.
- Multimodal and foundation models: vision–language/audio–language/video–language models; multimodal large language models; multimodal reasoning; model adaptation and efficient fine‑tuning; and multimodal agents.
- Generation, retrieval, and recommendation: generative multimedia content synthesis and editing, cross‑modal retrieval, semantic search, and recommendation and personalization.
- Applications and intelligent services: social‑media‑oriented services, AI‑empowered education, intelligent healthcare services, entertainment scenarios, smart‑city systems, and autonomous and IoT‑driven applications.
- Trustworthy and deployable multimedia AI: model explainability; adversarial robustness; bias and fairness issues; privacy‑preserving multimedia learning (including federated and distributed learning); deepfake detection and provenance tracking for synthetic media; media watermarking and content authenticity verification; lightweight and efficient multimodal models; edge‑AI and on‑device multimedia intelligence; real‑time multimodal data processing; energy‑friendly AI deployment, multimodal datasets, benchmarks, and model evaluation protocols; and weakly‑supervised and self‑supervised learning, domain adaptation, and model generalization.
This Topic hopes to encourage cross‑disciplinary communication, promote practical methodological innovation, and address open challenges along the full multimedia intelligence pipeline. We invite contributions that report new algorithms and hands‑on deployment experience, as well as forward‑looking viewpoint studies.
Prof. Dr. Teng Li
Prof. Dr. Fei Hao
Topic Editors
Keywords
- artificial intelligence
- multimedia analysis and generation
- multimodal foundation models
- deep learning
- generative AI
- vision–language models
- intelligent multimedia services
- cross-modal understanding
- trustworthy AI
- edge multimedia intelligence
Participating Journals
| Journal Name | Impact Factor | CiteScore | Launched Year | First Decision (median) | APC | |
|---|---|---|---|---|---|---|
AI
|
6.5 | 7.3 | 2020 | 20.4 Days | CHF 1800 | Submit |
Analytics
|
- | 4.3 | 2022 | 24.2 Days | CHF 1200 | Submit |
Big Data and Cognitive Computing
|
5.3 | 11.4 | 2017 | 23.3 Days | CHF 1800 | Submit |
Electronics
|
2.9 | 7.0 | 2012 | 14.8 Days | CHF 2400 | Submit |
Information
|
4.3 | 8.2 | 2010 | 18.7 Days | CHF 1800 | Submit |
Multimedia
|
- | - | 2025 | 15.0 days * | CHF 1000 | Submit |
Multimodal Technologies and Interaction
|
3.3 | 6.6 | 2017 | 22.8 Days | CHF 1800 | Submit |
* Median value for all MDPI journals in the first half of 2026.
Preprints.org is a multidisciplinary platform offering a preprint service designed to facilitate the early sharing of your research. It supports and empowers your research journey from the very beginning.
MDPI Topics is collaborating with Preprints.org and has established a direct connection between MDPI journals and the platform. Authors are encouraged to take advantage of this opportunity by posting their preprints at Preprints.org prior to publication:
- Share your research immediately: disseminate your ideas prior to publication and establish priority for your work.
- Safeguard your intellectual contribution: Protect your ideas with a time-stamped preprint that serves as proof of your research timeline.
- Boost visibility and impact: Increase the reach and influence of your research by making it accessible to a global audience.
- Gain early feedback: Receive valuable input and insights from peers before submitting to a journal.
- Ensure broad indexing: Web of Science (Preprint Citation Index), Google Scholar, Crossref, SHARE, PrePubMed, Scilit and Europe PMC.