Skip to Content
PublicationsPublications
  • Opinion
  • Open Access

1 June 2026

Literature Search Query in Academic Databases: Artificial Intelligence Think Tank Guideline for Literature Reviews

1
Department of Business Administration, University of Gothenburg, SE-405 30 Gothenburg, Sweden
2
Faculty of Engineering and Sustainable Development, University of Gävle, SE-801 76 Gävle, Sweden

Abstract

Literature reviews are essential for synthesizing existing knowledge, mapping research domains, identifying intellectual structures, and highlighting research gaps within a field. However, many literature reviews are incomplete because database search strategies are not adequately specified or validated. Search strategies are frequently underreported and undermotivated across the systematic review literature and bibliometrics, while query formulation remains time-consuming, error-prone, and particularly difficult in interdisciplinary or rapidly evolving topics. This article fills that void by developing a guideline for designing a professional topic query in existing academic databases and emphasizing search design as the front-end validity problem in bibliometric research. The article uses the Artificial Intelligence Think Tank framework as a methodological engine and applies it to bibliometric retrieval engineering via structured interaction with generative AI systems and human experts. The paper assists scholars performing bibliometric studies, scientometric analyses, systematic literature reviews, scoping reviews, and hybrid evidence-synthesis projects.

1. Introduction

Literature analyses, bibliometric reviews, systematic reviews, scoping reviews, and related evidence-synthesis studies depend on the search process used to assemble the study corpus. Search design is therefore not only a technical step in data collection; it helps determine which records enter the evidence base and, by extension, what can later be mapped, summarized, or interpreted. PRISMA-S was developed because literature searches are often not reported in sufficient detail to be clearly understood and reproduced, while comparative work on academic search systems shows that databases differ in retrieval qualities such as precision, recall, reproducibility, and user effort (Rethlefsen et al., 2021; Gusenbauer & Haddaway, 2020; Mongeon & Paul-Hus, 2016). The need for a structured protocol is reinforced by persistent reporting weaknesses. PRISMA-S provides a 16-item checklist for reporting systematic literature reviews (Rethlefsen et al., 2021). A recent analysis warns that search strategies are very poorly reported in literature reviews (Rethlefsen et al., 2024). A similar pattern is evident in not only literature reviews but also bibliometric research; Farooq et al. (2023) found that keyword selection, complete search queries, and search space were often inadequately reported, and that search-query quality was frequently poor. Studies (Farooq et al., 2023; MacFarlane et al., 2022; Rethlefsen et al., 2021; Bramer et al., 2018) indicate that search formulation and search reporting remain vulnerable points in review-based research. Search construction is also methodologically demanding. Bramer et al. (2018) describe literature-search development and translation across databases as challenging and propose a systematic approach for building efficient and complete searches. MacFarlane et al. (2022) characterize search-strategy formulation as complex, time-consuming, resource-intensive, and error-prone. These works support the need for a protocol that structures query development rather than entrusting it to undocumented trial and error.
Recent generative artificial intelligence (AI) tools add opportunity and risk to this setting. Gitman et al. (2025) found that AI could identify introduced errors and recreate omitted keywords in drug-harms search strategies, suggesting potential value for search support. Lieberum et al. (2025) conclude that AI applications in systematic reviews are increasing but fully established uses remain rare. Hence, this protocol was developed in that context: to provide a structured workflow for using AI in term generation and challenge rounds while retaining final judgment, validation, and reporting for the researcher. The protocol draws on the Artificial Intelligence Think Tank (AITT) framework, which describes a systematic AI–human problem-solving approach (2025).

2. Proposed Approach

This protocol adapts the AITT framework to the design of publication-grade search queries for academic databases. In this paper, AITT is used as a structured human-in-the-loop workflow for search development rather than as proof that a single tool or model is universally superior. The procedure combines problem definition, multi-AI elicitation, challenge rounds, and expert adjudication. This use is consistent with the original AITT framework, which presents AITT as a systematic AI-assisted problem-solving approach while explicitly retaining human oversight and demanding further real-world verification (Sorooshian, 2025). The search method literature likewise supports the use of structured rather than ad hoc query development because search formulation is complex, iterative, and prone to error if left undocumented (MacFarlane et al., 2022; McGowan et al., 2016). The minimum configuration for the protocol is one expert researcher, one generative AI system, access to the target database, and a search log. A stronger configuration uses two or more AI systems and, when feasible, an additional human validator with experience in database searching or evidence-synthesis methods. This scaling rule is justified because recent work shows that while AI systems can support keyword generation and error detection (Gitman et al., 2025; Lieberum et al., 2025; McGowan et al., 2016), their outputs remain insufficiently reliable for uncritical stand-alone use in review searching (Guimarães et al., 2024). This article offers a four-stage AITT for literature reviews, as Table 1 summarizes. For replication, the protocol requires retention of the following materials: the original master prompt, all AI prompts and raw outputs, model names and access dates, the synthesized term list, the challenged and rejected terms with reasons, all tested query variants, the final validated query, database and collection names, search and export dates, result counts, and the validation log. This record is consistent with the reproducibility emphasis of PRISMA-S and with the evidence that incompletely reported searches are often not reproducible (Rethlefsen et al., 2024; Rethlefsen et al., 2021).
Table 1. Quick-use sheet.
  • Stage 1. Define the search problem and build the master prompt
The researcher first defines the retrieval objective, target database or databases, exact collection or interface, intended use of the records, and the conceptual boundary of the topic. At this stage, the protocol requires the researcher to state what is inside the topic, what is outside it, and whether any preliminary limits are already justified, such as publication years, document types, language, or source restrictions. This emphasis on explicit scoping and documentation is consistent with stepwise search-development methods (MacFarlane et al., 2022; Rethlefsen et al., 2021; Bramer et al., 2018) and with PRISMA-S reporting expectations (Rethlefsen et al., 2021). The master prompt should tell the AI that the task is retrieval design rather than narrative summarization. It should also specify the target database and require database-compatible syntax only. This is important because existing academic databases (including Scopus and Web of Science) use different field tags and operator rules, and apparently similar searches are not always methodologically equivalent across systems. For example, Web of Science documents specify field tags, such as TS = for Topic and PY = for publication year, and define NEAR/x as its proximity operator; Scopus likewise documents the use of field codes, Boolean grouping, and proximity searching in advanced searches.
For instance, for this stage, a database-neutral master prompt can be written as follows: Act as a methodological search strategist. I am designing a publication-grade search query for [TOPIC] in [DATABASE/collection/interface] for [bibliometric study/systematic review/scoping review/hybrid evidence synthesis]. Treat this as a retrieval-methodology task, not a content-summary task. Define the conceptual boundary of the topic; identify the main concept blocks required for a query in [DATABASE]; list synonyms, near-synonyms, spelling variants, acronyms, and related labels for each block; identify terms that are likely to be ambiguous or generate false positives; and recommend an appropriate field choice and any justified restrictions. Return the answer as a table with the following columns: concept block|candidate term|why it belongs|precision risk|suggested handling in [DATABASE].
  • Stage 2. Generate candidate concept blocks and database-specific query strings
The master prompt is then run separately on one or more AI systems. Each output is saved with the model name, access date, and the exact prompt used. The purpose of this stage is term harvesting and structural drafting: the AI is asked to generate candidate concept blocks, synonyms, related expressions, spelling variants, acronyms, potential exclusions, and suggested Boolean and proximity structures. The search method literature supports this type of systematic decomposition into concepts rather than direct construction of one long undifferentiated string (MacFarlane et al., 2022; Bramer et al., 2018). At this stage, the protocol treats the query as a set of conceptual blocks. Terms within the same block are combined with OR, while distinct blocks are combined with AND. Phrase markers, truncation, proximity operators, and filters are then applied only where they are justified by the target database and topic definition. This block-based structure mirrors established systematic-search practice and facilitates later validation because each conceptual decision remains visible (MacFarlane et al., 2022; Bramer et al., 2018; McGowan et al., 2016).
A Stage 2 prompt could be as follows: Using the conceptual blocks for [TOPIC] in [DATABASE], generate one main search query and, if useful, one broader and one narrower sensitivity variant. Use only [DATABASE]-compatible syntax. Show where Boolean operators, proximity operators, exact phrases, truncation, and wildcards should be used. Explain which restrictions should be embedded in the query and which may be left as interface filters. Also indicate likely sources of false positives and false negatives. Return the answer in two parts: exact query string(s), followed by a concise methodological rationale.
  • Stage 3. Challenge, compare, and adjudicate the candidate search
The researcher then synthesizes the initial AI outputs and feeds that synthesis back into the same or other AI systems using challenge prompts. The core challenge questions are as follows: What is missing? What is irrelevant? Which terms are too ambiguous? Which proximity choices are weak? Which database-specific renderings do not preserve conceptual intent? This step is important because omission threatens recall, while indiscriminate synonym expansion threatens precision and topical coherence. Peer Review of Electronic Search Strategies (PRESS) is directly relevant here because its retained checklist elements include translation of the research question, Boolean and proximity operators, subject headings, text words, syntax, and limits or filters (MacFarlane et al., 2022; McGowan et al., 2016). The protocol then requires explicit human adjudication. Each candidate term is retained, modified, moved to a sensitivity analysis, or removed using stated criteria: conceptual relevance, expected contribution to recall, risk to precision, ambiguity, disciplinary appropriateness, and database compatibility. This is also the stage where cross-database translation is tested critically rather than accepted as a simple syntax conversion. Official Web of Science documentation shows that topic, publication-year, and proximity behavior are defined by specific field and operator rules, while Scopus guidance likewise emphasizes structured grouping and proximity handling. Accordingly, the protocol treats database translation as a methodological decision rather than a literal conversion exercise.
Two challenge prompts may be used:
  • For the synthesized search package below, identify omitted synonyms, overlooked disciplinary variants, missing acronyms, adjacent terms worth testing, and any missing validation checks. Return the answer as missing item|why it matters|expected effect on recall|risk to precision|recommended action.
  • For the same candidate search package, identify items that are overly broad, irrelevant, syntactically weak, or likely to damage reproducibility. Return the answer as problematic item|problem type|explanation|recommended revision.
  • Stage 4. Validate and package the final search
After adjudication, the final search is validated and packaged. Validation transcends syntax checking. First, the search is tested against a small set of clearly relevant or seminal papers to verify that known items are retrievable. Second, a sample of retrieved records is inspected to assess visible relevance, borderline cases, and obvious noise. Third, at least one sensitivity analysis is run, such as broader versus narrower proximity windows, broader versus narrower field scopes, or inclusion and exclusion of disputed synonyms. Fourth, all tested strings, dates, result counts, and reasons for revision are preserved in a search log. These requirements align with PRESS, PRISMA-S, and recent reproducibility evidence showing that poor reporting of search methods is a major reason why published searches cannot be reliably rerun (Rethlefsen et al., 2021, 2024; McGowan et al., 2016). The final output of the protocol is a retrieval package rather than merely a final string. The package should contain at least the exact final query, target database and collection, search date, limits or filters used, rationale for field choice, validation decisions, result count before cleaning, model name(s) and access dates for AI-assisted stages, and any notes required for downstream processing. PRISMA-S specifically requires clear reporting of information sources, dates of search, full electronic search strategies, and limits used so that the search can be repeated (Rethlefsen et al., 2024; McGowan et al., 2016).
A Stage 4 audit prompt could be as follows: Act as an external methodological reviewer of the following search package for [TOPIC] in [DATABASE]: [paste final query, limits, validation notes]. Evaluate conceptual adequacy, semantic coverage, precision risk, syntax validity, appropriateness of filters, and reproducibility. For each criterion, recommend, retain, modify, test separately, or remove, and explain why.
The validation of this method is procedural and demonstration-based rather than a full benchmarking exercise against all alternative search-design approaches. This is appropriate because the article aims to present a reproducible protocol for academic database query development rather than to claim universal superiority over other methods. Thus, method validation in this study focused on feasibility and procedural rigor. Two experts with experience in academic database searching and evidence-synthesis methods reviewed the protocol and confirmed that its workflow was feasible for practical use. Their review considered the clarity of the staged procedure, suitability of the prompt structure, database compatibility, and adequacy of the embedded validation and documentation steps.

3. Conclusions

This paper addresses a critical issue in bibliometric and scientometric research: developing reproducible and transparent search queries for academic databases. It emphasizes query formulation as a validity tool, which is critical for the accuracy of analyses derived from bibliometric mapping and literature reviews. While the paper provides operational templates to help researchers and emphasizes the importance of custom query strategies for various databases, it has been criticized for lacking empirical validation against other methods. It suggests potential improvements for future research, such as comparing the AITT with alternative workflows and implementing it systematically across multiple databases. The authors acknowledge that even the best queries can produce poor results if the subsequent data handling is inadequate. Overall, the project aims to elevate search formulation from an informal task to a structured, systematic endeavor, encouraging more diligent and collaborative practices in literature retrieval.

Funding

This research received no external funding.

Data Availability Statement

No new data were created or analyzed in this study.

Acknowledgments

The author used ChatGPT/OpenAI (version 5.5 Thinking, 2026) and AI-assisted tools to refine author-created material and to support manuscript development, including shortening the initial drafted manuscript, improving structure and language, identifying areas for clarification and brainstorming for improvement, and supporting reference checking. The author independently reviewed and verified all AI-assisted outputs, consulted the original sources, and takes full responsibility for the final manuscript.

Conflicts of Interest

The authors declares no conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
AITTArtificial Intelligence Think Tank

References

  1. Bramer, W. M., de Jonge, G. B., Rethlefsen, M. L., Mast, F., & Kleijnen, J. (2018). A systematic approach to searching: An efficient and complete method to develop literature searches. Journal of the Medical Library Association, 106(4), 531–541. [Google Scholar] [CrossRef] [Scilit]
  2. Farooq, U., Nasir, A., & Khan, K. I. (2023). An assessment of the quality of the search strategy: A case of bibliometric studies published in business and economics. Scientometrics, 128, 4855–4874. [Google Scholar] [CrossRef] [Scilit]
  3. Gitman, V., Maxwell, C., & Gamble, J.-M. (2025). Enhancing search strategies for systematic reviews on drug harms: An evaluation of the utility of ChatGPT in error detection and keyword generation. Computers in Biology and Medicine, 193, 110464. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Guimarães, N. S., Joviano-Santos, J. V., Reis, M. G., Chaves, R. R. M., & Observatory of Epidemiology, Nutrition, Health Research (OPENS). (2024). Development of search strategies for systematic reviews in health using ChatGPT: A critical analysis. Journal of Translational Medicine, 22, 1. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Gusenbauer, M., & Haddaway, N. R. (2020). Which academic search systems are suitable for systematic reviews or meta-analyses? Research Synthesis Methods, 11(2), 181–217. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Lieberum, J.-L., Toews, M., Metzendorf, M., Heilmeyer, F., Siemens, W., Haverkamp, C., Böhringer, D., Meerpohl, J. J., & Eisele-Metzger, A. (2025). Large language models for conducting systematic reviews: On the rise, but not yet ready for use—A scoping review. Journal of Clinical Epidemiology, 181, 111746. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. MacFarlane, A., Russell-Rose, T., & Shokraneh, F. (2022). Search strategy formulation for systematic reviews: Issues, challenges and opportunities. Intelligent Systems with Applications, 15, 200091. [Google Scholar] [CrossRef] [Scilit]
  8. McGowan, J., Sampson, M., Salzwedel, D. M., Cogo, E., Foerster, V., & Lefebvre, C. (2016). PRESS peer review of electronic search strategies: 2015 guideline statement. Journal of Clinical Epidemiology, 75, 40–46. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Mongeon, P., & Paul-Hus, A. (2016). The journal coverage of Web of Science and Scopus: A comparative analysis. Scientometrics, 106, 213–228. [Google Scholar] [CrossRef] [Scilit]
  10. Rethlefsen, M. L., Brigham, T. J., Price, C., Moher, D., Bouter, L. M., Kirkham, J. J., Schroter, S., & Zeegers, M. P. (2024). Systematic review search strategies are poorly reported and not reproducible: A cross-sectional meta-research study. Journal of Clinical Epidemiology, 166, 111229. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Rethlefsen, M. L., Kirtley, S., Waffenschmidt, S., Ayala, A. P., Moher, D., Page, M. J., & Koffel, J. B. (2021). PRISMA-S: An extension to the PRISMA statement for reporting literature searches in systematic reviews. Systematic Reviews, 10, 39. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Sorooshian, S. (2025). Artificial intelligence think tank: A modern problem-solving framework. Frontiers in Artificial Intelligence, 8, 1603562. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.