An Open-Source Reproducible Preprocessing Pipeline for Merging Bibliometric Data from Multiple Databases
Abstract
1. Introduction
- Merging of databases provides comprehensive academic literature coverage without missing the most relevant and important articles in the concerned research field.
- The search results will be improved in the merged dataset by identifying the duplicates and multiple DOIs.
- The merged dataset enables a comprehensive analysis that provides valuable insights to the researchers to generate innovative and quality research.
- Accessing the comprehensive and accurate dataset enables data-driven decision-making in the research evaluation.
- Automatic merging of data from multiple databases will save the time and effort of the researchers, enabling them to spend more time on providing quality research work.
2. Literature Review, Problem Statement and Proposed Solution
2.1. Literature Review
2.2. Problem Statement
- A bibliometrix package “mergeDbSources” was already available for merging the databases (Bibliometrix, 2020). It was identified that the articles under ISI were only recorded in the merged dataset by using this package, which may cause the omission of important articles from other databases.
- The code available in Kasaraneni and Rosaline (2024) was executed to merge the articles from the Scopus and WoS databases. The duplicate records were identified based on the DOI. An issue was identified during the removal of duplicate records. It was observed that different databases use either uppercase or lowercase letters to represent data in the DOI column. This issue may impact the duplicate records identification.
- Another issue was identified in the datasets resulting from the keyword search. Multiple DOIs and Titles were generated in the same record, which creates inconsistency in the bibliometric data.
2.3. Proposed Solution
3. Data Collection
4. Methodology
5. Rationale of the Proposed Research Work
- Issue 1: By using the mergeDbSources function, only Institute for Scientific Information (ISI) records (fetched from WoS database) are prioritized. To avoid this consequence, the authors have developed an open-source reproducible preprocessing pipeline to get records from all databases with equal priority.
- Issue 2: DOIs are recorded in small and capital cases in multiple databases. All DOIs are converted to records in small cases for better identification of duplicate records.
- Issue 3: The other issue recognized is that multiple DOIs are identified in one record, which may affect the identification of duplicates. Hence, the records with multiple DOIs are identified and eliminated from the merged dataset.
6. Results and Discussion
- Some records contain both multiple DOIs and multiple titles, which is very important to focus on.
- Some records contain multiple DOIs but no multiple titles.
7. Conclusions
- While performing the merging process, the proposed preprocessing pipeline removes duplicates from all the databases with equal priority by solving the issue of getting only ISI-indexed records.
- DOIs are presented in different cases of alphabets in different databases. This case difference will affect the searching of duplicate records. Hence, to solve this issue, a function is implemented to convert all DOIs to lowercase.
- While merging the records from multiple databases, multiple DOIs and titles were identified and removed from the final merged dataset.Finally, a merged dataset without the noisy data is available for conducting the required analyses.
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| DOI | Digital Object Identifier |
| ISI | Institute for Scientific Information |
| PRISMA | Preferred Reporting Items for Systematic Reviews and Meta-Analyses |
| SJR | SCImago Journal Rank |
| SNIP | Source Normalized Impact per Paper |
| WoS | Web of Science |
References
- Aria, M., & Cuccurullo, C. (2017). bibliometrix: An R-tool for comprehensive science mapping analysis. Journal of Informetrics, 11(4), 959–975. [Google Scholar] [CrossRef] [Scilit]
- Berga, M., Coelho, P., & Ochman, A. (2025). Top 21 data mining tools. Available online: https://www.imaginarycloud.com/blog/data-mining-tools (accessed on 21 June 2025).
- Bibliometrix. (2020). Comprehensive science mapping analysis. Available online: https://www.bibliometrix.org/ (accessed on 17 March 2026).
- Caputo, A., & Kargina, M. (2022). A user-friendly method to merge Scopus and Web of Science data during bibliometric analysis. Journal of Marketing Analytics, 10(1), 82–88. [Google Scholar] [CrossRef] [Scilit]
- Chansanam, W., & Li, C. (2025). KKU-BiblioMerge: A novel tool for multi-database integration in bibliometric analysis. Iberoamerican Journal of Science Measurement and Communication, 5(1), 1–16. [Google Scholar] [CrossRef] [Scilit]
- Clarivate. (2026). Web of science database. Available online: https://access.clarivate.com/login?app=wos (accessed on 17 May 2026).
- Culbert, J. H., Hobert, A., Jahn, N., Haupka, N., Schmidt, M., Donner, P., & Mayr, P. (2025). Reference coverage analysis of OpenAlex compared to Web of Science and Scopus. Scientometrics, 130(4), 2475–2492. [Google Scholar] [CrossRef] [Scilit]
- Donthu, N., Kumar, S., Mukherjee, D., Pandey, N., & Lim, W. M. (2021). How to conduct a bibliometric analysis: An overview and guidelines. Journal of Business Research, 133, 285–296. [Google Scholar] [CrossRef] [Scilit]
- Echchakoui, S. (2020). Why and how to merge Scopus and Web of Science during bibliometric analysis: The case of sales force literature from 1912 to 2019. Journal of Marketing Analytics, 8(3), 165–184. [Google Scholar] [CrossRef] [Scilit]
- Elsevier. (2026). Scopus database. Available online: https://www.scopus.com/pages/home (accessed on 17 May 2026).
- Fang, Y., Li, X., Mohtar, T. M., & Chekima, B. (2025). A bibliometric analysis of work–family balance: Trends, themes, and future directions (2000–2024). Cogent Business & Management, 12(1), 2541041. [Google Scholar] [CrossRef] [Scilit]
- Giorgi, F. M., Ceraolo, C., & Mercatelli, D. (2022). The R language: An engine for bioinformatics and data science. Life, 12(5), 648. [Google Scholar] [CrossRef] [Scilit]
- Gong, Y., Liu, G., Xue, Y., Li, R., & Meng, L. (2023). A survey on dataset quality in machine learning. Information and Software Technology, 162, 107268. [Google Scholar] [CrossRef] [Scilit]
- Groos, O. V., & Pritchard, A. (1969). Documentation notes. Journal of Documentation, 25(4), 344–349. [Google Scholar] [CrossRef] [Scilit]
- Kara, B. C., Şahin, A., & Dirsehan, T. (2025). BibexPy: Harmonizing the bibliometric symphony of Scopus and Web of Science. SoftwareX, 30, 102098. [Google Scholar] [CrossRef] [Scilit]
- Kasaraneni, H., & Rosaline, S. (2024). Automatic merging of Scopus and Web of Science data for simplified and effective bibliometric analysis. Annals of Data Science, 11(3), 785–802. [Google Scholar] [CrossRef] [Scilit]
- Lens. (2026). Lens database. Available online: https://www.lens.org/ (accessed on 17 May 2026).
- Moher, D., Liberati, A., Tetzlaff, J., Altman, D. G., & The PRISMA Group. (2009). Preferred reporting items for systematic reviews and meta-analyses: The PRISMA statement. PLoS Medicine, 6(7), e1000097. [Google Scholar] [CrossRef] [Scilit]
- Nosek, B. A., Alter, G., Banks, G. C., Borsboom, D., Bowman, S. D., Breckler, S. J., Buck, S., Chambers, C. D., Chin, G., Christensen, G., Contestabile, M., Dafoe, A., Eich, E., Freese, J., Glennerster, R., Goroff, D., Green, D. P., Hesse, B., Humphreys, M., … Yarkoni, T. (2015). Promoting an open research culture. Science, 348(6242), 1422–1425. [Google Scholar] [CrossRef] [Scilit]
- Shi, Y., Tian, Y., Kou, G., Peng, Y., & Li, J. (2011). Optimization based data mining: Theory and applications. Springer. [Google Scholar] [CrossRef] [Scilit]
- Singh, V. K., Singh, P., Karmakar, M., Leta, J., & Mayr, P. (2021). The journal coverage of Web of Science, Scopus and Dimensions: A comparative analysis. Scientometrics, 126(6), 5113–5142. [Google Scholar] [CrossRef] [Scilit]
- Tranfield, D., Denyer, D., & Smart, P. (2003). Towards a methodology for developing evidence-informed management knowledge by means of systematic review. British Journal of Management, 14(3), 207–222. [Google Scholar] [CrossRef] [Scilit]
- Turgel, I. D., & Chernova, O. A. (2024). Open science alternatives to Scopus and the Web of Science: A case study in regional resilience. Publications, 12(4), 43. [Google Scholar] [CrossRef] [Scilit]
- Vera-Baceta, M.-A., Thelwall, M., & Kousha, K. (2019). Web of Science and Scopus language coverage. Scientometrics, 121(3), 1803–1813. [Google Scholar] [CrossRef] [Scilit]
- Vieira, G. A., & Leta, J. (2024). biblioverlap: An R package for document matching across bibliographic datasets. Scientometrics, 129(7), 4513–4527. [Google Scholar] [CrossRef] [Scilit]
- Visser, M., Van Eck, N. J., & Waltman, L. (2021). Large-scale comparison of bibliographic data sources: Scopus, Web of Science, Dimensions, Crossref, and Microsoft Academic. Quantitative Science Studies, 2(1), 20–41. [Google Scholar] [CrossRef] [Scilit]
- Zhu, J., & Liu, W. (2020). A tale of two databases: The use of Web of Science and Scopus in academic papers. arXiv, arXiv:2002.02608. [Google Scholar] [CrossRef] [Scilit]









| Reference | Merging Process | Years Considered | Databases Considered |
|---|---|---|---|
| Visser et al. (2021) | Manual | 2008–2017 | WoS, Scopus, Crossref, Dimensions, and Microsoft Academic |
| Singh et al. (2021) | Manual | 2010–2018 | Scopus, WoS, and Dimensions |
| Zhu and Liu (2020) | Manual | 2004–2018 | Scopus, WoS |
| Caputo and Kargina (2022) | Bibliometrix and Excel (Automatic up to some extent) | Not available | Scopus, WoS |
| Vera-Baceta et al. (2019) | Advanced query in the database | 2018 | Scopus, WoS |
| Kasaraneni and Rosaline (2024) | Automatic | 2022 | Scopus, WoS |
| Echchakoui (2020) | Manual process, Bibliometrix and Excel | 1912–2019 | Scopus, WoS |
| Source | DOI | Document Title | ISSN | Source | Document Type | Authors’ Keywords | Publisher | Authors |
|---|---|---|---|---|---|---|---|---|
| Scopus | DI | TI | SN | SO | DT | DE | PU | AU |
| WoS | DI | TI | SN | SO | DT | DE | PU | AU |
| Lens | DI | TI | ISSN | SO | DT | DE | PU | AU |
| Database | Search Strategy |
|---|---|
| WoS | “brain tumor detection” OR “brain tumor classification” (All Fields) and 2013 or 2014 or 2015 or 2016 or 2017 or 2018 or 2019 or 2020 or 2021 or 2022 or 2023 (Publication Years) and Article or Review Article (Document Types) and English (Languages) |
| Scopus | TITLE-ABS-KEY (“brain tumor detection” OR “brain tumor classification”) AND PUBYEAR > 2012 AND PUBYEAR < 2024 AND (LIMIT-TO (DOCTYPE, “ar”) OR LIMIT-TO (DOCTYPE, “re”) ) AND (LIMIT-TO (PUBSTAGE, “final”) ) AND (LIMIT-TO (SRCTYPE, “j”) ) AND (LIMIT-TO (LANGUAGE, “English”) ) |
| Lens | Scholarly Works (868) = (“brain tumor detection” OR “brain tumor classification”) Filters: Year Published = (2013–2023) Publication Type = (journal article, conference proceedings article) External ID Type = (DOI) |
| Steps |
|---|
| 1. Import required R libraries to achieve the desired functionality. list.of.packages ← c (“bibliometrix”, “tidyverse”, “readxl”, “writexl”, “openxlsx”) |
| 2. Choose Scopus .bib file. read, convert, and save it as an Excel file scopusbibFile ← file.choose () scopusbibFile ← convert2df (file = scopusbibFile, dbsource = ‘scopus’, format = “bibtex”) scopus_ExcelFile ← write.xlsx (scopusbibFile, “scopus.xlsx”) |
| 3. Choose WOS .bib file. read, convert, and save it as an Excel file wosbibFile ← file.choose () wosbibFile ← convert2df (file = wosbibFile, dbsource = ‘isi’, format = “bibtex”) wos_ExcelFile ← write.xlsx (wosbibFile, “wos.xlsx”) |
| 4. Choose Lens .csv file. read, convert, and save it as an Excel file lenscsvFile ← file.choose () lenscsvFile ← convert2df (file = lenscsvFile, dbsource = ‘lens’, format = “csv”) lens_ExcelFile ← write.xlsx (lenscsvFile, “lens.xlsx”) |
| 5. Merge the obtained three Excel files as a singe file (without duplicates) using mergeDBSources function. file_merged ← mergeDbSources (scopusbibFile, wosbibFile, lenscsvFile, remove.duplicated = T) write.xlsx (file_merged, “file_merged.xlsx”) |
| Steps |
|---|
| 1. Import required R libraries to achieve the desired functionality. list.of.packages ← c (“bibliometrix”, “tidyverse”, “readxl”, “writexl”, “openxlsx”, “stringr”, “janitor”) |
| 2. Read Scopus, WoS, and Lens data Excel files and save them into the respective objects. scopus_file ← file.choose () wos_file ← file.choose () lens_file ← file.choose () |
| 3. Retrieve column names from each file and save them into the respective objects. cols_scopus ← colnames (scopus_file) cols_wos ← colnames (wos_file) cols_lens ← colnames (lens_file) |
| 4. Retrieve common column names from these columns and save them in an object. common_cols ← intersect (cols_scopus, cols_wos, cols_lens) |
| 5. Select data from each file based on the common columns identified and save as a sub-dataset. subdata_scopus ← select (scopus_file, common_cols) subdata_wos ← select (wos_file, common_cols) subdata_lens ← select (lens_file, common_cols) |
| 6. Merge all sub-datasets. merged_data ← rbind (rbind (subdata_scopus, subdata_wos), subdata_lens) |
| 7. Convert DOIs to lowercase. The common column for DOI is DI. This column name is used for case conversion. merged_data$DI ← as.character (merged_data$DI) merged_data$DI ← tolower (merged_data$DI) |
| 8. Find the duplicate records based on DOI and remove them. Keep the distinct records. # Duplicate records dup_merged_data ← merged_data[duplicated (merged_data$DI), ] cat (“No. of Duplicate Records =”, print (nrow (dup_merged_data))) print (dup_merged_data) write_xlsx (dup_merged_data, “Duplicates.xlsx”) # Unique records unique_data ← distinct (merged_data, DI, .keep_all = T) write_xlsx (unique_data, “merged_data_after_deduplication.xlsx”) |
| 9. Search for multiple DOIs and blanks in DI column using a pattern and remove the records with multiple DOIs. pattern_doi_alphabet ← c (“10.\\d4,9/[-._;()/:a-z0-9A-Z]+10.\\d4,9/[-._;()/:a-z0-9A-Z]+”) # Searching for multiple DOIs multiple_doi_alphabet ← multiple_doi_alphabet % > % remove_empty(“rows”) write_xlsx (multiple_doi_alphabet, “multiple_doi_alphabet.xlsx”) print (nrow (multiple_doi_alphabet)) pattern_doi_whitespace ← 10\\.[^\\ s]+\\ s+[^\\ s]+ multiple_doi_whitespace ← unique_data[str_detect (unique_data$DI, pattern_doi_whitespace), ] multiple_doi_whitespace ← multiple_doi_whitespace % > % remove_empty(“rows”) write_xlsx (multiple_doi_whitespace, “multiple_doi_whitespace.xlsx”) print (nrow (multiple_doi_whitespace)) all_patterns ← paste (pattern_doi_alphabet, pattern_doi_whitespace, sep = “|”) |
| 10. Save unique data after removing the records with multiple DOIs. unique_data ← unique_data[!str_detect (unique_data$DI, all_patterns), ] write_xlsx (unique_data, “merged_data_after_removing_multiple_DOIs.xlsx”) print (nrow (unique_data)) |
| 11. Tracking processing time. end.time ← Sys.time () time.taken ← end.time − start.time print (paste0 (“Start Time”, start.time)) print (paste0 (“End Time”, end.time)) print (time.taken) |
| Description | Quantitative Details |
|---|---|
| Number of records in Scopus file | 1138 |
| Number of records in WOS file | 839 |
| Number of records in Lens file | 868 |
| Number of records in the merged file before deduplication | 2845 |
| Number of duplicate records in the merged file | 910 |
| Number of records in the merged file after deduplication | 1935 |
| Number of records with multiple DOIs (considering alphabets in DOI) | 8 |
| Number of records with multiple DOIs (considering whitespace in DOI) | 8 |
| Number of records in the merged file after removing records with multiple DOIs | 1927 |
| Total Processing time | “Start Time 2026-05-17 06:08:07.319666” “End Time 2026-05-17 06:08:09.519493” Time difference of 2.199827 s |
| Description | Quantitative Details |
|---|---|
| Number of records in Scopus file | 19,999 |
| Number of records in WOS file | 20,000 |
| Number of records in Lens file | 10,000 |
| Number of records in the merged file before deduplication | 49,999 |
| Number of duplicate records in the merged file | 12,734 |
| Number of records in the merged file after deduplication | 37,265 |
| Number of records with multiple DOIs (considering alphabets in DOI) | 7 |
| Number of records with multiple DOIs (considering whitespace in DOI) | 7 |
| Number of records in the merged file after removing records with multiple DOIs | 37,258 |
| Total processing time | “Start Time 2026-05-17 06:22:16.413374” “End Time 2026-05-17 06:22:37.726767” Time difference of 21.31339 s |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Purna Prakash, K.; HimaJyothi, K.; Rosaline, S.; Venkata Pavan Kumar, Y.; Reddy, G.P.; Mukkapati, N. An Open-Source Reproducible Preprocessing Pipeline for Merging Bibliometric Data from Multiple Databases. Publications 2026, 14, 34. https://doi.org/10.3390/publications14020034
Purna Prakash K, HimaJyothi K, Rosaline S, Venkata Pavan Kumar Y, Reddy GP, Mukkapati N. An Open-Source Reproducible Preprocessing Pipeline for Merging Bibliometric Data from Multiple Databases. Publications. 2026; 14(2):34. https://doi.org/10.3390/publications14020034
Chicago/Turabian StylePurna Prakash, Kasaraneni, Kasaraneni HimaJyothi, Salini Rosaline, Yellapragada Venkata Pavan Kumar, Gogulamudi Pradeep Reddy, and Naveen Mukkapati. 2026. "An Open-Source Reproducible Preprocessing Pipeline for Merging Bibliometric Data from Multiple Databases" Publications 14, no. 2: 34. https://doi.org/10.3390/publications14020034
APA StylePurna Prakash, K., HimaJyothi, K., Rosaline, S., Venkata Pavan Kumar, Y., Reddy, G. P., & Mukkapati, N. (2026). An Open-Source Reproducible Preprocessing Pipeline for Merging Bibliometric Data from Multiple Databases. Publications, 14(2), 34. https://doi.org/10.3390/publications14020034

