Opportunities and Challenges in Big Data Analytics for Decision Making: An Integrated Framework
Abstract
1. Introduction
2. Conceptual Background and Framework Development
2.1. BDA Capability and Decision Making
2.2. Conceptual Synthesis and Framework Development
2.3. Socio-Technical Mechanisms and Contextual Conditions
3. Design Dimensions of Big Data Analytics
3.1. Processing Dimension
3.2. Platform Dimension
3.3. Analytics Dimension
3.4. Governance Dimension
4. Implementation Requirements and Constraints
4.1. Data Quality
4.2. Scalability and Cost
4.3. Security and Privacy as Distinct Risk Dimensions
4.4. Interoperability
4.5. Explainability and Bias
4.6. Skills Gap
4.7. Responsible AI and Sustainability
4.8. Organizational, Task, and Institutional Conditions
5. Architectural Rationale: From Traditional to Modern BDA
5.1. Traditional Architectural Logic
5.1.1. On-Premises Clusters
5.1.2. Batch Processing
5.1.3. Hadoop/MapReduce
5.1.4. Limited Governance
5.2. Modern Architectural Logic
5.2.1. Cloud-Native Analytics
5.2.2. Real-Time Processing
5.2.3. Lakehouse Architecture
5.2.4. MLOps and DataOps
5.2.5. Integrated Governance
5.3. Design Propositions
6. Proposed Integrated Big Data Analytics Decision-Making Framework
6.1. Conceptual Foundation
6.2. Architectural Structure of the Framework
6.3. AI Enablement Across the Framework
6.4. Cross-Cutting Governance, Human Oversight, and Organizational Readiness
6.5. Feedback, Monitoring, and Continuous Adaptation
6.6. Framework Evaluation and Validation Strategy
6.7. Comparison with Existing Big Data Analytics Frameworks
6.8. Boundary Conditions, Scalability, and Configurability
6.9. Practical Feasibility, Implementation Challenges, and Open-Source Pathways
7. Conclusions and Future Directions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Chen, H.; Chiang, R.H.L.; Storey, V.C. Business intelligence and analytics: From big data to big impact. MIS Q. 2012, 36, 1165–1188. [Google Scholar] [CrossRef] [Scilit]
- Labrinidis, A.; Jagadish, H.V. Challenges and opportunities with big data. Proc. VLDB Endow. 2012, 5, 2032–2033. [Google Scholar] [CrossRef] [Scilit]
- Agarwal, R.; Dhar, V. Editorial-Big data, data science, and analytics: The opportunity and challenge for IS research. Inf. Syst. Res. 2014, 25, 443–448. [Google Scholar] [CrossRef] [Scilit]
- Badshah, A.; Daud, A.; Alharbey, R.; Banjar, A.; Bukhari, A.; Alshemaimri, B. Big data applications: Overview, challenges and future. Artif. Intell. Rev. 2024, 57, 290. [Google Scholar] [CrossRef] [Scilit]
- Jamarani, A.; Haddadi, S.; Sarvizadeh, R.; Haghi Kashani, M.; Akbari, M.; Moradi, S. Big data and predictive analytics: A systematic review of applications. Artif. Intell. Rev. 2024, 57, 176. [Google Scholar] [CrossRef] [Scilit]
- Tosi, D.; Kokaj, R.; Roccetti, M. 15 years of Big Data: A systematic literature review. J. Big Data 2024, 11, 73. [Google Scholar] [CrossRef] [Scilit]
- Huynh, M.T.; Nippa, M.; Aichner, T. Big data analytics capabilities: Patchwork or progress. A systematic review of the status quo and implications for future research. Technol. Forecast. Soc. Change 2023, 197, 122884. [Google Scholar] [CrossRef] [Scilit]
- Rodepeter, E.; Gschnaidtner, C.; Hottenrott, H. Big data-based management decisions and start-up performance. Small Bus. Econ. 2026, 67, 361–399. [Google Scholar] [CrossRef] [Scilit]
- Wixom, B.; Yen, B.; Relich, M. Maximizing value from business analytics. MIS Q. Exec. 2013, 12, 111–123. [Google Scholar]
- Elgendy, N.; Elragal, A. Big data analytics in support of the decision-making process. Procedia Comput. Sci. 2016, 100, 1071–1084. [Google Scholar] [CrossRef] [Scilit]
- Chatterjee, S.; Chaudhuri, R.; Gupta, S.; Sivarajah, U.; Bag, S. Assessing the impact of big data analytics on decision-making processes, forecasting, and performance of a firm. Technol. Forecast. Soc. Change 2023, 196, 122824. [Google Scholar] [CrossRef] [Scilit]
- Kreuzberger, D.; Kuhl, N.; Hirschl, S. Machine Learning Operations (MLOps): Overview, definition, and architecture. IEEE Access 2023, 11, 31866–31879. [Google Scholar] [CrossRef] [Scilit]
- Harby, A.A.; Zulkernine, F. Data lakehouse: A survey and experimental study. Inf. Syst. 2025, 127, 102460. [Google Scholar] [CrossRef] [Scilit]
- Rajan, A.A.; Vetriselvi, V. Systematic survey: Secure and privacy-preserving big data analytics in cloud. J. Comput. Inf. Syst. 2024, 64, 136–156. [Google Scholar] [CrossRef] [Scilit]
- Papagiannidis, E.; Mikalef, P.; Conboy, K. Responsible artificial intelligence governance: A review and research framework. J. Strateg. Inf. Syst. 2025, 34, 101885. [Google Scholar] [CrossRef] [Scilit]
- Shmueli, G.; Koppius, O.R. Predictive analytics in information systems research. MIS Q. 2011, 35, 553–572. [Google Scholar] [CrossRef] [Scilit]
- Chen, C.L.P.; Zhang, C.Y. Data-intensive applications, challenges, techniques and technologies: A survey on Big Data. Inf. Sci. 2014, 275, 314–347. [Google Scholar] [CrossRef] [Scilit]
- Gupta, M.; George, J.F. Toward the development of a big data analytics capability. Inf. Manag. 2016, 53, 1049–1064. [Google Scholar] [CrossRef] [Scilit]
- Ghasemaghaei, M.; Hassanein, K.; Turel, O. Increasing firm agility through the use of data analytics: The role of fit. Decis. Support Syst. 2018, 101, 95–105. [Google Scholar] [CrossRef] [Scilit]
- Phan, T.; Baird, K. The use of big data analytics in performance management: The antecedents and role in enhancing performance measurement system effectiveness. J. Manag. Control 2026, 37, 111–139. [Google Scholar] [CrossRef] [Scilit]
- Janssen, M.; van der Voort, H.; Wahyudi, A. Factors influencing big data decision-making quality. J. Bus. Res. 2017, 70, 338–345. [Google Scholar] [CrossRef] [Scilit]
- Muller, O.; Fay, M.; vom Brocke, J. The effect of big data and analytics on firm performance: An econometric analysis considering industry characteristics. J. Manag. Inf. Syst. 2018, 35, 488–509. [Google Scholar] [CrossRef] [Scilit]
- Ghasemaghaei, M. Does data analytics use improve firm decision making quality. The role of knowledge sharing and data analytics competency. Decis. Support Syst. 2019, 120, 14–24. [Google Scholar] [CrossRef] [Scilit]
- Chen, M.; Mao, S.; Liu, Y. Big data: A survey. Mob. Netw. Appl. 2014, 19, 171–209. [Google Scholar] [CrossRef] [Scilit]
- Gandomi, A.; Haider, M. Beyond the hype: Big data concepts, methods, and analytics. Int. J. Inf. Manag. 2015, 35, 137–144. [Google Scholar] [CrossRef] [Scilit]
- Oussous, A.; Benjelloun, F.Z.; Lahcen, A.A.; Belfkih, S. Big data technologies: A survey. J. King Saud Univ. Comput. Inf. Sci. 2018, 30, 431–448. [Google Scholar] [CrossRef] [Scilit]
- Elgendy, N.; Elragal, A. Big data analytics: A literature review paper. In Advances in Data Mining. Applications and Theoretical Aspects; Springer: Cham, Switzerland, 2014; pp. 214–227. [Google Scholar]
- Provost, F.; Fawcett, T. Data science and its relationship to big data and data-driven decision making. Big Data 2013, 1, 51–59. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, Y.; Kung, L.; Byrd, T.A. Big data analytics: Understanding its capabilities and potential benefits for healthcare organizations. Technol. Forecast. Soc. Change 2018, 126, 3–13. [Google Scholar] [CrossRef] [Scilit]
- Akter, S.; Wamba, S.F.; Gunasekaran, A.; Dubey, R.; Childe, S.J. How to improve firm performance using big data analytics capability and business strategy alignment. Int. J. Prod. Econ. 2016, 182, 113–131. [Google Scholar] [CrossRef] [Scilit]
- Cao, J. Intelligent decision-making in business management: Integrating artificial intelligence and big data analytics for strategic optimization in enterprise operations. Sustain. Comput. Inform. Syst. 2026, 51, 101382. [Google Scholar] [CrossRef] [Scilit]
- Amou Najafabadi, F.A.; Bogner, J.; Gerostathopoulos, I.; Lago, P. An architectural perspective on MLOps: Structures, processes, tools, and stakeholders. Inf. Softw. Technol. 2026, 193, 108029. [Google Scholar] [CrossRef] [Scilit]
- Abraham, R.; Schneider, J.; vom Brocke, J. Data governance: A conceptual framework, structured review, and research agenda. Int. J. Inf. Manag. 2019, 49, 424–438. [Google Scholar] [CrossRef] [Scilit]
- Xia, H.; Chen, H.; Zhang, J.Z.; Kamal, M.M. Exploring the impact of responsible AI governance on corporate performance: A quasi-natural experiment. Technol. Forecast. Soc. Change 2026, 223, 124425. [Google Scholar] [CrossRef] [Scilit]
- Mikalef, P.; Pappas, I.O.; Krogstie, J.; Giannakos, M. Big data analytics capabilities: A systematic literature review and research agenda. Inf. Syst. e-Bus. Manag. 2018, 16, 547–578. [Google Scholar] [CrossRef] [Scilit]
- Mikalef, P.; Krogstie, J. Examining the interplay between big data analytics and contextual factors in driving process innovation capabilities. Eur. J. Inf. Syst. 2020, 29, 260–287. [Google Scholar] [CrossRef] [Scilit]
- Hashem, I.A.T.; Yaqoob, I.; Anuar, N.B.; Mokhtar, S.; Gani, A.; Khan, S.U. The rise of big data on cloud computing: Review and open research issues. Inf. Syst. 2015, 47, 98–115. [Google Scholar] [CrossRef] [Scilit]
- Assuncao, M.D.; Calheiros, R.N.; Bianchi, S.; Netto, M.A.S.; Buyya, R. Big Data computing and clouds: Trends and future directions. J. Parallel Distrib. Comput. 2015, 79–80, 3–15. [Google Scholar] [CrossRef] [Scilit]
- Kune, R.; Konugurthi, P.K.; Agarwal, A.; Chillarige, R.R.; Buyya, R. The anatomy of big data computing. Softw. Pract. Exp. 2016, 46, 79–105. [Google Scholar] [CrossRef] [Scilit]
- Yaqoob, I.; Hashem, I.A.T.; Gani, A.; Mokhtar, S.; Ahmed, E.; Anuar, N.B.; Vasilakos, A.V. Big data: From beginning to future. Int. J. Inf. Manag. 2016, 36, 1231–1247. [Google Scholar] [CrossRef] [Scilit]
- Marjani, M.; Nasaruddin, F.; Gani, A.; Karim, A.; Hashem, I.A.T.; Siddiqa, A.; Yaqoob, I. Big IoT data analytics: Architecture, opportunities, and open research challenges. IEEE Access 2017, 5, 5247–5261. [Google Scholar] [CrossRef] [Scilit]
- Marcu, O.-C.; Bouvry, P. Big Data Stream Processing; Technical Report; University of Luxembourg: Luxembourg, 2024. [Google Scholar]
- Dean, J.; Ghemawat, S. MapReduce: Simplified data processing on large clusters. Commun. ACM 2008, 51, 107–113. [Google Scholar]
- Sakr, S.; Liu, A.; Batista, D.M.; Alomari, M. A survey of large scale data management approaches in cloud environments. IEEE Commun. Surv. Tutor. 2011, 13, 311–336. [Google Scholar] [CrossRef] [Scilit]
- Henning, S.; Hasselbring, W. Benchmarking scalability of stream processing frameworks deployed as microservices in the cloud. J. Syst. Softw. 2024, 208, 111879. [Google Scholar] [CrossRef] [Scilit]
- Almeida, A.; Brás, S.; Sargento, S.; Pinto, F.C. Time series big data: A survey on data stream frameworks, analysis and algorithms. J. Big Data 2023, 10, 83. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zaharia, M.; Xin, R.S.; Wendell, P.; Das, T.; Armbrust, M.; Dave, A.; Meng, X.; Rosen, J.; Venkataraman, S.; Franklin, M.J.; et al. Apache Spark: A unified engine for big data processing. Commun. ACM 2016, 59, 56–65. [Google Scholar]
- Carbone, P.; Katsifodimos, A.; Ewen, S.; Markl, V.; Haridi, S.; Tzoumas, K. Apache Flink: Stream and batch processing in a single engine. IEEE Data Eng. Bull. 2015, 38, 28–38. [Google Scholar]
- Ogrizović, M.; Drašković, D.; Bojić, D. Quality assurance strategies for machine learning applications in big data analytics: An overview. J. Big Data 2024, 11, 156. [Google Scholar] [CrossRef] [Scilit]
- Grolinger, K.; Higashino, W.A.; Tiwari, A.; Capretz, M.A.M. Data management in cloud environments: NoSQL and NewSQL data stores. J. Cloud Comput. 2013, 2, 22. [Google Scholar] [CrossRef] [Scilit]
- Nargesian, F.; Zhu, E.; Miller, R.J.; Pu, K.Q.; Arocena, P.C. Data Lake management: Challenges and opportunities. Proc. VLDB Endow. 2019, 12, 1986–1989. [Google Scholar]
- Armbrust, M.; Ghodsi, A.; Xin, R.S.; Zaharia, M. Lakehouse: A new generation of open platforms that unify data warehousing and advanced analytics. Proc. CIDR 2021, 8, 28. [Google Scholar]
- Pohl, M.; Wijemanne, N.D.; Staegemann, D.; Haertel, C.; Daase, C.; Dreschel, D.; Walia, D.S.; Osterthun, A.; Reibert, J.; Turowski, K. Data Lakehouse for Time Series Data: A Systematic Literature Review. In Proceedings of the 2024 IEEE International Conference on Big Data (BigData), Washington, DC, USA, 15–18 December 2024; pp. 5833–5842. [Google Scholar] [CrossRef] [Scilit]
- Ji, S.; Li, Q.; Cao, W.; Zhang, P.; Muccini, H. Quality assurance technologies of big data applications: A systematic literature review. Appl. Sci. 2020, 10, 8052. [Google Scholar] [CrossRef] [Scilit]
- Buyya, R.; Ilager, S.; Arroba, P. Energy-efficiency and sustainability in new generation cloud computing: A vision and directions for integrated management of data centre resources and workloads. Softw. Pract. Exp. 2024, 54, 24–38. [Google Scholar] [CrossRef] [Scilit]
- Wu, X.; Zhu, X.; Wu, G.Q.; Ding, W. Data mining with big data. IEEE Trans. Knowl. Data Eng. 2014, 26, 97–107. [Google Scholar] [CrossRef] [Scilit]
- Qiu, J.; Wu, Q.; Ding, G.; Xu, Y.; Feng, S. A survey of machine learning for big data processing. EURASIP J. Adv. Signal Process. 2016, 2016, 67. [Google Scholar] [CrossRef] [Scilit]
- Najafabadi, M.M.; Villanustre, F.; Khoshgoftaar, T.M.; Seliya, N.; Wald, R.; Muharemagic, E. Deep learning applications and challenges in big data analytics. J. Big Data 2015, 2, 1. [Google Scholar] [CrossRef] [Scilit]
- Jordan, M.I.; Mitchell, T.M. Machine learning: Trends, perspectives, and prospects. Science 2015, 349, 255–260. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Witten, I.H.; Frank, E.; Hall, M.A.; Pal, C.J. Data Mining: Practical Machine Learning Tools and Techniques, 4th ed.; Morgan Kaufmann: Burlington, MA, USA, 2016. [Google Scholar]
- Castro, A.; Villagrá, V.A.; García, P.; Rivera, D.; Toledo, D. An ontological-based model to data governance for big data. IEEE Access 2021, 9, 109943–109959. [Google Scholar] [CrossRef] [Scilit]
- Khatri, V.; Brown, C.V. Designing data governance. Commun. ACM 2010, 53, 148–152. [Google Scholar] [CrossRef] [Scilit]
- Otto, B. A morphology of the organisation of data governance. In Proceedings of the 19th European Conference on Information Systems (ECIS), Helsinki, Finland, 9–11 June 2011; p. 272. [Google Scholar]
- Tallon, P.P.; Ramirez, R.V.; Short, J.E. The information artifact in IT governance: Toward a theory of information governance. J. Manag. Inf. Syst. 2013, 30, 141–178. [Google Scholar] [CrossRef] [Scilit]
- Bližnák, K.; Munk, M.; Pilková, A. A systematic review of recent literature on data governance (2017–2023). IEEE Access 2024, 12, 149875–149888. [Google Scholar] [CrossRef] [Scilit]
- Gieß, A.; Hutterer, A. The future of data management: A delimitation of data platforms, data spaces, data meshes, and data fabrics. Inf. Syst. e-Bus. Manag. 2025, 23, 971–997. [Google Scholar] [CrossRef] [Scilit]
- Armbrust, M.; Fox, A.; Griffith, R.; Joseph, A.D.; Katz, R.; Konwinski, A.; Lee, G.; Patterson, D.; Rabkin, A.; Stoica, I.; et al. A view of cloud computing. Commun. ACM 2010, 53, 50–58. [Google Scholar] [CrossRef] [Scilit]
- AbouZaid, A.; Barclay, P.J.; Chrysoulas, C.; Pitropakis, N. Building a modern data platform based on the data lakehouse architecture and cloud-native ecosystem. Discov. Appl. Sci. 2025, 7, 166. [Google Scholar] [CrossRef] [Scilit]
- Tonnarelli, M.; Kumara, I.; Driessen, S.; Tamburri, D.A.; van den Heuvel, W.J.; Oor, P. Data catalog tools: A systematic multivocal literature review. J. Syst. Softw. 2025, 230, 112584. [Google Scholar] [CrossRef] [Scilit]
- Kreps, J.; Narkhede, N.; Rao, J. Kafka: A distributed messaging system for log processing. Proc. NetDB 2011, 11, 1–7. [Google Scholar]
- Shvachko, K.; Kuang, H.; Radia, S.; Chansler, R. The Hadoop Distributed File System. In Proceedings of the 2010 IEEE 26th Symposium on Mass Storage Systems and Technologies (MSST), Incline Village, NV, USA, 3–7 May 2010; pp. 1–10. [Google Scholar]
- Wixom, B.H.; Watson, H.J. The BI-based organization. Int. J. Bus. Intell. Res. 2010, 1, 13–28. [Google Scholar] [CrossRef] [Scilit]
- Sivarajah, U.; Kamal, M.M.; Irani, Z.; Weerakkody, V. Critical analysis of big data challenges and analytical methods. J. Bus. Res. 2017, 70, 263–286. [Google Scholar] [CrossRef] [Scilit]
- Demchenko, Y.; Ngo, C.; de Laat, C.; Membrey, P.; Gordijenko, D. Big security for big data: Addressing security challenges for the big data infrastructure. In Proceedings of the Workshop on Secure Data Management; Springer: Cham, Switzerland, 2014; pp. 76–85. [Google Scholar]
- Katal, A.; Wazid, M.; Goudar, R.H. Big data: Issues, challenges, tools and good practices. In Proceedings of the 2013 Sixth International Conference on Contemporary Computing (IC3), Noida, India, 8–10 August 2013; pp. 404–409. [Google Scholar]
- Khan, N.; Yaqoob, I.; Hashem, I.A.T.; Inayat, Z.; Ali, W.; Alam, M.; Shiraz, M.; Gani, A. Big data: Survey, technologies, opportunities, and challenges. Sci. World J. 2014, 2014, 712826. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Batini, C.; Cappiello, C.; Francalanci, C.; Maurino, A. Methodologies for data quality assessment and improvement. ACM Comput. Surv. 2009, 41, 1–52. [Google Scholar] [CrossRef] [Scilit]
- Cai, L.; Zhu, Y. The challenges of data quality and data quality assessment in the big data era. Data Sci. J. 2015, 14, 2. [Google Scholar] [CrossRef] [Scilit]
- Merino, J.; Caballero, I.; Rivas, B.; Serrano, M.; Piattini, M. A data quality in use model for big data. Future Gener. Comput. Syst. 2016, 63, 123–130. [Google Scholar] [CrossRef] [Scilit]
- Taleb, I.; Serhani, M.A.; Dssouli, R. Big data quality: A survey. In Proceedings of the IEEE International Congress on Big Data, San Francisco, CA, USA, 2–7 July 2018; pp. 166–173. [Google Scholar]
- Begoli, E.; Goethert, I.; Knight, K. A lakehouse architecture for the management and analysis of heterogeneous data for analytics and AI. In Proceedings of the 2021 IEEE International Conference on Big Data (Big Data), Orlando, FL, USA, 15–18 December 2021; pp. 5317–5326. [Google Scholar]
- Dinh, H.T.; Lee, C.; Niyato, D.; Wang, P. A survey of mobile cloud computing: Architecture, applications, and approaches. Wirel. Commun. Mob. Comput. 2013, 13, 1587–1611. [Google Scholar] [CrossRef] [Scilit]
- Ranjan, R. Streaming big data processing in datacenter clouds. IEEE Cloud Comput. 2014, 1, 78–83. [Google Scholar] [CrossRef] [Scilit]
- Buyya, R.; Srirama, S.N.; Casale, G.; Calheiros, R.; Simmhan, Y.; Varghese, B.; Gelenbe, E.; Javadi, B.; Vaquero, L.M.; Netto, M.A.S.; et al. A manifesto for future generation cloud computing: Research directions for the next decade. ACM Comput. Surv. 2019, 51, 1–38. [Google Scholar]
- Toosi, A.N.; Calheiros, R.N.; Buyya, R. Interconnected cloud computing environments: Challenges, taxonomy, and survey. ACM Comput. Surv. 2014, 47, 7. [Google Scholar] [CrossRef] [Scilit]
- Hai, R.; Geisler, S.; Quix, C. Constance: An intelligent data lake system. In Proceedings of the 2016 International Conference on Management of Data (SIGMOD ‘16), San Francisco, CA, USA, 26 June–1 July 2016; pp. 2097–2100. [Google Scholar] [CrossRef] [Scilit]
- Zuech, R.; Khoshgoftaar, T.M.; Wald, R. Intrusion detection and big heterogeneous data: A survey. J. Big Data 2015, 2, 3. [Google Scholar] [CrossRef] [Scilit]
- Kshetri, N. Big data’s impact on privacy, security and consumer welfare. Telecommun. Policy 2014, 38, 1134–1145. [Google Scholar] [CrossRef] [Scilit]
- Suthaharan, S. Big data classification: Problems and challenges in network intrusion prediction with machine learning. ACM SIGMETRICS Perform. Eval. Rev. 2014, 41, 70–73. [Google Scholar]
- Dwork, C. Differential privacy. In Automata, Languages and Programming; Springer: Berlin, Germany, 2006; pp. 1–12. [Google Scholar]
- Kairouz, P.; McMahan, H.B.; Avent, B.; Bellet, A.; Bennis, M.; Bhagoji, A.N.; Bonawitz, K.; Charles, Z.; Cormode, G.; Cummings, R.; et al. Advances and open problems in federated learning. Found. Trends Mach. Learn. 2021, 14, 1–210. [Google Scholar] [CrossRef] [Scilit]
- Li, T.; Sahu, A.K.; Talwalkar, A.; Smith, V. Federated learning: Challenges, methods, and future directions. IEEE Signal Process Mag. 2020, 37, 50–60. [Google Scholar] [CrossRef] [Scilit]
- Halevy, A.; Korn, F.; Noy, N.F.; Olston, C.; Polyzotis, N.; Roy, S.; Whang, S.E. Goods: Organizing Google’s datasets. In Proceedings of the 2016 International Conference on Management of Data; Association for Computing Machinery: New York, NY, USA, 2016; pp. 795–806. [Google Scholar]
- Dong, X.L.; Srivastava, D. Big data integration. Synth. Lect. Data Manag. 2015, 7, 1–198. [Google Scholar] [CrossRef] [Scilit]
- Lenzerini, M. Data integration: A theoretical perspective. In Proceedings of the Twenty-First ACM SIGMOD-SIGACT-SIGART Symposium on Principles of Database Systems; Association for Computing Machinery: New York, NY, USA, 2002; pp. 233–246. [Google Scholar]
- Stonebraker, M.; Ilyas, I.F. Data integration: The current status and the way forward. IEEE Data Eng. Bull. 2018, 41, 3–9. [Google Scholar]
- Doan, A.; Halevy, A.; Ives, Z. Principles of Data Integration; Morgan Kaufmann: Waltham, MA, USA, 2012. [Google Scholar]
- Guidotti, R.; Monreale, A.; Ruggieri, S.; Turini, F.; Giannotti, F.; Pedreschi, D. A survey of methods for explaining black box models. ACM Comput. Surv. 2018, 51, 1–42. [Google Scholar] [CrossRef] [Scilit]
- Barredo Arrieta, A.; Diaz-Rodriguez, N.; Del Ser, J.; Bennetot, A.; Tabik, S.; Barbado, A.; Garcia, S.; Gil-Lopez, S.; Molina, D.; Benjamins, R.; et al. Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Inf. Fusion 2020, 58, 82–115. [Google Scholar] [CrossRef] [Scilit]
- Rudin, C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat. Mach. Intell. 2019, 1, 206–215. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mehrabi, N.; Morstatter, F.; Saxena, N.; Lerman, K.; Galstyan, A. A survey on bias and fairness in machine learning. ACM Comput. Surv. 2021, 54, 1–35. [Google Scholar] [CrossRef] [Scilit]
- Holstein, K.; Vaughan, J.W.; Daume, H.; Dudik, M.; Wallach, H. Improving fairness in machine learning systems: What do industry practitioners need. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, 2019; pp. 1–16. [Google Scholar]
- Mittelstadt, B.; Russell, C.; Wachter, S. Explaining explanations in AI. In Proceedings of the Conference on Fairness, Accountability, and Transparency; Association for Computing Machinery: New York, NY, USA, 2019; pp. 279–288. [Google Scholar]
- Raji, I.D.; Smart, A.; White, R.N.; Mitchell, M.; Gebru, T.; Hutchinson, B.; Smith-Loud, J.; Theron, D.; Barnes, P. Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency; Association for Computing Machinery: New York, NY, USA, 2020; pp. 33–44. [Google Scholar]
- Davenport, T.H.; Patil, D.J. Data scientist: The sexiest job of the 21st century. Harv. Bus. Rev. 2012, 90, 70–76. [Google Scholar] [PubMed]
- Chiang, R.H.L.; Goes, P.; Stohr, E.A. Business intelligence and analytics education, and program development: A unique opportunity for the information systems discipline. ACM Trans. Manag. Inf. Syst. 2012, 3, 1–13. [Google Scholar]
- Debortoli, S.; Muller, O.; vom Brocke, J. Comparing business intelligence and big data skills. Bus. Inf. Syst. Eng. 2014, 6, 289–300. [Google Scholar] [CrossRef] [Scilit]
- Vidgen, R.; Shaw, S.; Grant, D.B. Management challenges in creating value from business analytics. Eur. J. Oper. Res. 2017, 261, 626–639. [Google Scholar] [CrossRef] [Scilit]
- Jobin, A.; Ienca, M.; Vayena, E. The global landscape of AI ethics guidelines. Nat. Mach. Intell. 2019, 1, 389–399. [Google Scholar] [CrossRef] [Scilit]
- Floridi, L.; Cowls, J. A unified framework of five principles for AI in society. Harv. Data Sci. Rev. 2019, 1, 1–15. [Google Scholar]
- Mittelstadt, B. Principles alone cannot guarantee ethical AI. Nat. Mach. Intell. 2019, 1, 501–507. [Google Scholar] [CrossRef] [Scilit]
- Schwartz, R.; Dodge, J.; Smith, N.A.; Etzioni, O. Green AI. Commun. ACM 2020, 63, 54–63. [Google Scholar] [CrossRef] [Scilit]
- Hu, H.; Wen, Y.; Chua, T.S.; Li, X. Toward scalable systems for big data analytics: A technology tutorial. IEEE Access 2014, 2, 652–687. [Google Scholar] [CrossRef] [Scilit]
- Inmon, W.H. Building the Data Warehouse, 4th ed.; Wiley: Indianapolis, Indiana, 2005. [Google Scholar]
- Kimball, R.; Ross, M. The Data Warehouse Toolkit, 3rd ed.; Wiley: Indianapolis, Indiana, 2013. [Google Scholar]
- Chaudhuri, S.; Dayal, U. An overview of data warehousing and OLAP technology. ACM SIGMOD Rec. 1997, 26, 65–74. [Google Scholar] [CrossRef] [Scilit]
- Stonebraker, M.; Abadi, D.J.; DeWitt, D.J.; Madden, S.; Paulson, E.; Pavlo, A.; Rasin, A. MapReduce and parallel DBMSs: Friends or foes. Commun. ACM 2010, 53, 64–71. [Google Scholar]
- Pavlo, A.; Paulson, E.; Rasin, A.; Abadi, D.J.; DeWitt, D.J.; Madden, S.; Stonebraker, M. A comparison of approaches to large-scale data analysis. In Proceedings of the 2009 ACM SIGMOD International Conference on Management of Data; Association for Computing Machinery: New York, NY, USA, 2009; pp. 165–178. [Google Scholar]
- Zhang, Q.; Cheng, L.; Boutaba, R. Cloud computing: State-of-the-art and research challenges. J. Internet Serv. Appl. 2010, 1, 7–18. [Google Scholar] [CrossRef] [Scilit]
- Vellido, A. The importance of interpretability and visualization in machine learning for applications in medicine and health care. Neural Comput. Appl. 2020, 32, 18069–18083. [Google Scholar] [CrossRef] [Scilit]
- Zaharia, M.; Chowdhury, M.; Franklin, M.J.; Shenker, S.; Stoica, I. Spark: Cluster computing with working sets. In Proceedings of the 2nd USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 10); USENIX Association: Berkeley, CA, USA, 2010; pp. 1–7. [Google Scholar]
- Bu, Y.; Howe, B.; Balazinska, M.; Ernst, M.D. HaLoop: Efficient iterative data processing on large clusters. Proc. VLDB Endow. 2010, 3, 285–296. [Google Scholar]
- Weber, K.; Otto, B.; Osterle, H. One size does not fit all: A contingency approach to data governance. J. Data Inf. Qual. 2009, 1, 1–27. [Google Scholar] [CrossRef] [Scilit]
- Kolajo, T.; Daramola, O.; Adebiyi, A. Big data stream analysis: A systematic literature review. J. Big Data 2019, 6, 47. [Google Scholar] [CrossRef] [Scilit]
- Sculley, D.; Holt, G.; Golovin, D.; Davydov, E.; Phillips, T.; Ebner, D.; Chaudhary, V.; Young, M.; Crespo, J.F.; Dennison, D. Hidden technical debt in machine learning systems. Adv. Neural Inf. Process. Syst. 2015, 28, 2503–2511. [Google Scholar]
- Amershi, S.; Begel, A.; Bird, C.; DeLine, R.; Gall, H.; Kamar, E.; Nagappan, N.; Nushi, B.; Zimmermann, T. Software engineering for machine learning: A case study. In Proceedings of the 2019 IEEE/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP); IEEE: Piscataway, NJ, USA, 2019; pp. 291–300. [Google Scholar]
- Juma’h, A.H.; Li, Y. The effects of auditors’ knowledge, professional skepticism, and perceived adequacy of accounting standards on their intention to use blockchain. Int. J. Account. Inf. Syst. 2023, 51, 100650. [Google Scholar] [CrossRef] [Scilit]
- Hossain, M.K.; Srivastava, A.; Oliver, G.C.; Islam, M.E.; Jahan, N.A.; Karim, R.; Kanij, T.; Mahdi, T.H. Adoption of artificial intelligence and big data analytics: An organizational readiness perspective of the textile and garment industry in Bangladesh. Bus. Process Manag. J. 2024, 30, 2665–2683. [Google Scholar] [CrossRef] [Scilit]
- Alnsour, Y.; Juma’h, A.H. The effect of political environment on security and privacy of contact tracing apps evaluation. Int. J. Public Sect. Manag. 2024, 37, 864–879. [Google Scholar] [CrossRef] [Scilit]

| Decision-Making Phase | Main Analytical Need | Suitable Big Data Support | Illustrative Scholarly Support |
|---|---|---|---|
| Intelligence | Recognizing problems, opportunities, anomalies and environmental signals. | Data collection, integration, dashboards, descriptive analytics and early-warning indicators. | Chen et al. [1]; Chatterjee et al. [11] |
| Design | Formulating alternatives, scenarios and possible courses of action. | Predictive analytics, simulation, data mining and exploratory modeling. | Shmueli and Koppius [16]; Jamarani et al. [5] |
| Choice | Selecting the most suitable alternative under uncertainty and constraints. | Prescriptive analytics, optimization, ranking models and evidence-based comparison. | Elgendy and Elragal [10]; Cao [31] |
| Implementation | Deploying decisions, monitoring outcomes and adjusting actions. | Streaming analytics, operational dashboards, MLOps and performance monitoring. | Ghasemaghaei [23]; Amou Najafabadi et al. [32] |
| Learning and control | Evaluating decision outcomes and improving future decision rules. | Governance, audit trails, feedback loops, model monitoring and accountability mechanisms. | Abraham et al. [33]; Xia et al. [34] |
| Category | Representative Tools/Platforms | Main Use | Decision-Making Value |
|---|---|---|---|
| Distributed processing | Apache Spark, Hadoop MapReduce | Large-scale batch and distributed data processing | Supports analysis of large historical datasets |
| Stream analytics | Apache Kafka, Apache Flink, Spark Structured Streaming | Real-time data pipelines and continuous analytics | Supports fast operational and risk-related decisions |
| Cloud analytics | BigQuery, Redshift, Azure Synapse, Microsoft Fabric | Elastic storage, integration, and analytics services | Reduces infrastructure barriers and improves scalability |
| Lakehouse systems | Databricks, Delta Lake, Apache Iceberg, Apache Hudi | Unified data lake and warehouse capabilities | Supports flexible analytics and machine learning |
| Visualization and BI | Power BI, Tableau, Looker, MicroStrategy | Dashboards, reports, and visual exploration | Improves interpretability and communication |
| Machine learning and MLOps | H2O.ai, MLflow, Kubeflow, Azure ML | Model development, deployment, and monitoring | Supports predictive and prescriptive decisions |
| Governance tools | Microsoft Purview, Collibra, Apache Atlas | Data cataloging, lineage, quality, and access management | Improves trust, compliance, and accountability |
| Taxonomy Approach | Main Advantages | Main Disadvantages | Representative Scholarly Support |
|---|---|---|---|
| Processing-oriented approaches | They support large-scale data transformation, parallel execution and analysis of historical data. They are useful when organizations need to process massive datasets that cannot be handled by single-machine systems. | They may require complex cluster configuration, specialized technical skills and high resource consumption. Batch-oriented processing can also delay decisions when real-time response is required. | MapReduce introduced distributed batch processing [43], while recent stream-processing reviews emphasize the move toward continuous analytics [42]. |
| Platform-oriented approaches | They provide scalable infrastructure for storage, computation, analytics services and deployment. Cloud platforms also reduce the need for organizations to own all hardware resources internally. | They can create dependence on providers, unpredictable costs, data sovereignty concerns and migration difficulties. Performance and governance also depend on platform configuration. | Cloud elasticity was framed by Armbrust et al. [67], while recent cloud-native Lakehouse work emphasizes portable, resilient platforms [68]. |
| Analytics-oriented approaches | They convert data into descriptive, predictive and prescriptive insights. They are valuable for classification, forecasting, anomaly detection, recommendation and decision support. | Their effectiveness depends on data quality, model selection, interpretability and domain expertise. Complex models may produce accurate output but weak explanations for decision makers. | Machine learning became central to analytics [59], and recent reviews highlight predictive analytics applications in big data environments [5]. |
| Governance-oriented approaches | They improve trust, accountability, data quality, privacy protection and regulatory readiness. They also clarify ownership and responsibility across the analytics lifecycle. | They can slow implementation when governance procedures are heavy, fragmented or not aligned with organizational culture. Governance requires continuous coordination between technical and managerial teams. | Data governance clarifies decision rights and accountability [62], while recent work extends governance to data catalogs and responsible AI [34,69]. |
| Lifecycle Layer | Typical Techniques or Functions | Representative Tools/Platforms | Decision Output |
|---|---|---|---|
| Data ingestion | Batch ingestion, event capture, sensor data capture and log collection. | Kafka, ETL/ELT pipelines, APIs and data connectors. | Timely access to operational and external signals [42,70]. |
| Storage and organization | Distributed file storage, cloud object storage, data lakes and Lakehouse organization. | HDFS, cloud storage, Delta Lake, Apache Iceberg and Lakehouse platforms. | Reliable historical and Lakehouse-based data foundation for analysis and reporting [13,71]. |
| Processing | Batch processing, iterative processing, in-memory computation and stream processing. | MapReduce, Spark, Flink and cloud processing services. | Faster transformation of raw data into analytical variables [45,47]. |
| Analytics and modeling | Descriptive, diagnostic, predictive, prescriptive, machine learning and deep learning analytics. | R, Python ecosystems, H2O.ai, ML platforms and MLOps pipelines. | Patterns, predictions, classifications and recommended actions [31,58]. |
| Visualization and communication | Dashboards, scorecards, visual exploration and self-service analytics. | Power BI, Tableau, Looker and business intelligence platforms. | Understandable insights for managers and domain experts [20,72]. |
| Governance and control | Metadata management, lineage, access control, data quality monitoring and compliance. | Data catalogs, governance platforms and policy-based access tools. | Trustworthy and accountable decision support [15,62]. |
| Taxonomy Challenge | Nature of the Challenge | Effect on Decision Making | Recommended Response |
|---|---|---|---|
| Data quality | Incomplete, noisy, inconsistent, duplicated, outdated, or biased data from multiple sources | Weakens forecasting, risk detection, reporting accuracy, and confidence in analytical outputs | Apply data profiling, cleaning, validation, lineage, metadata management, and continuous quality monitoring |
| Scalability and cost | Growing storage, compute, network, and real-time processing demand | Delays insights or increases operational cost when analytics workloads are not optimized | Use elastic architecture, workload optimization, lifecycle policies, cost monitoring, and appropriate platform selection |
| Security | Unauthorized access, manipulation, disruption, insecure interfaces, and attacks across distributed data and analytics infrastructure | Threatens confidentiality, integrity, availability, reliability, and willingness to depend on analytical systems | Use identity and access management, encryption, secure pipelines, monitoring, incident response, audit trails, and resilience controls |
| Privacy | Excessive collection, linkage, inference, re-identification, secondary use, or inappropriate sharing of sensitive data | Reduces stakeholder trust, constrains legitimate data use, and creates ethical and regulatory exposure | Apply purpose limitation, data minimization, privacy assessment, anonymization/pseudonymization, differential privacy, retention rules, and governed data sharing |
| Interoperability | Different systems, schemas, formats, tools, and repositories may not exchange data effectively | Produces fragmented evidence and weakens cross-functional decision making | Use APIs, metadata standards, semantic integration, data catalogs, common models, and lakehouse architecture |
| Explainability and bias | Complex models may be difficult to interpret and may reproduce biased patterns from data | Reduces trust, accountability, fairness, and acceptance of analytics-supported decisions | Use explainable AI, bias testing, model documentation, validation, human review, and fairness monitoring |
| Skills gap | Organizations may lack the technical, analytical, domain, and managerial skills required for BDA | Limits the ability to interpret results and transform analytics into action | Develop data literacy, interdisciplinary teams, training programs, and communication between technical and business users |
| Responsible AI and sustainability | AI-driven analytics raises ethical, social, accountability, and environmental concerns | May create harmful, opaque, costly, or unsustainable decision systems | Adopt responsible AI governance, sustainability metrics, workload optimization, human oversight, and transparent accountability mechanisms |
| Challenge Area | Possible Response | Advantages | Limitations |
|---|---|---|---|
| Data quality | Data profiling, cleansing, validation, metadata management and lineage tracking. | Improves the reliability of dashboards, models and forecasts. It also reduces misleading results caused by missing, duplicate or inconsistent data [49,77]. | Quality improvement can be costly and require continuous effort because big data sources change frequently and may be generated outside organizational control. |
| Scalability and cost | Elastic cloud resources, workload monitoring and cost-aware architecture design. | Allows organizations to scale storage and computation according to demand rather than fixed capacity, especially when cloud resources are managed efficiently [55]. | Elasticity may make costs unpredictable when workloads, queries or data movement are not controlled. |
| Security | Encryption, identity and access control, secure pipelines, monitoring, incident response, and resilience controls. | Protects the confidentiality, integrity, availability, and reliability of data and analytical services [14,74,87,88,89]. | Strong controls can add cost and operational complexity and must be calibrated to risk and criticality. |
| Privacy | Data minimization, purpose limitation, anonymization/pseudonymization, differential privacy, federated approaches, and governed sharing. | Supports legitimate use of sensitive data and can strengthen stakeholder trust and regulatory acceptability [14,90,91,92]. | Privacy protection can reduce analytical granularity or data availability and may require context-specific trade-offs. |
| Interoperability | Schema mapping, data integration tools, common data models and semantic alignment. | Supports integration across heterogeneous systems and reduces fragmentation between operational and analytical data [69,94]. | Integration remains difficult when sources have different formats, meanings, update cycles and ownership structures. |
| Explainability and bias | Explainable AI methods, fairness assessment, model documentation and human review. | Improves transparency and allows decision makers to understand and challenge model outputs [34,99]. | Explanations may be incomplete, and bias mitigation requires social, organizational, and technical assessments. |
| Skills gap | Training, interdisciplinary teams and collaboration between domain experts, data engineers and analysts. | Improves the ability to translate analytical results into practical decisions. Big data and BI skills require both technical and business knowledge [20,107]. | Training requires time and resources, and organizations may still face shortages in advanced data engineering and analytics expertise. |
| Responsible AI and sustainability | AI governance, model audit, energy-aware design and responsible data practices. | Supports accountable analytics and reduces environmental and ethical risks. Responsible AI and Green AI emphasize governance, computational cost and sustainability [15,112]. | Responsible AI and sustainability metrics are still developing, and organizations may struggle to balance accuracy, speed, cost and ethical obligations. |
| Challenge Area | Primary Risk to Analytics | Decision-Making Consequence | Priority Response |
|---|---|---|---|
| Data quality | Incomplete, inconsistent, duplicated or outdated data. | Incorrect forecasts, misleading dashboards and weak confidence in analytical results. | Establish quality rules, profiling, validation and continuous monitoring [49,77]. |
| Security | Unauthorized access, manipulation, service disruption, weak authentication, and compromised analytical pipelines. | Unreliable or unavailable evidence, operational interruption, and lower confidence in analytics-dependent decisions. | Prioritize access control, encryption, monitoring, secure development, incident response, and resilience according to system criticality [14,74,87,88,89]. |
| Privacy | Over-collection, re-identification, secondary use, inappropriate inference, or sharing of sensitive data. | Reduced stakeholder trust, restrictions on data use, legal exposure, and lower legitimacy of analytics-supported decisions. | Apply privacy-by-design, minimization, purpose limitation, privacy-preserving analytics, retention controls, and governed data sharing [14,90,91,92]. |
| Interoperability | Fragmented systems, inconsistent schemas and semantic mismatch between data sources. | Partial decision views and difficulty combining operational, customer and external data. | Use common metadata, APIs, semantic integration and catalog-supported governance [69,97]. |
| Explainability and bias | Opaque models, biased training data and difficulty justifying automated recommendations. | Unfair, untrusted or non-actionable decisions in high-impact domains. | Apply model documentation, explainable methods, bias testing and human review [34,98]. |
| Scalability and cost | Processing delays, uncontrolled cloud spending and performance bottlenecks. | Slow decisions, delayed reporting and poor operational responsiveness. | Adopt elastic architecture, workload monitoring and cost-aware platform design [45,113]. |
| Sustainability | High energy consumption caused by large-scale storage, training and processing. | Conflict between analytics value and environmental responsibility. | Apply green AI principles, efficient modeling and workload optimization [55,112]. |
| Taxonomy Item | Traditional Approach | Modern Approach | Decision-Making Implication |
|---|---|---|---|
| Infrastructure | On-premises clusters managed internally | Cloud-native or hybrid infrastructure with elastic resources | Modern systems improve scalability and deployment speed; traditional systems offer direct control. |
| Processing model | Batch processing at scheduled intervals | Batch, stream, and hybrid real-time processing | Modern systems reduce decision latency and support faster operational response. |
| Core technology | Hadoop/MapReduce and distributed file systems | Spark, Kafka, Flink, cloud analytics services, and Lakehouse platforms | Modern tools support interactive, iterative, and continuous analytics more effectively. |
| Data architecture | Separate storage, processing, and reporting layers | Integrated Lakehouse architecture linking storage, analytics, ML, and BI | Integrated architecture improves data reuse and reduces fragmentation. |
| Governance | Limited or fragmented governance across systems | Integrated catalogs, lineage, access control, quality rules, and model governance | Stronger governance increases trust, compliance, and accountability. |
| Analytics capability | Mostly historical reporting and offline analysis | Predictive, prescriptive, AI-assisted, and real-time analytics | Modern analytics supports proactive and adaptive decision making. |
| Operational practices | Manual pipeline and model management | DataOps and MLOps for automation, monitoring, and versioning | Operational discipline improves reliability and reproducibility of analytical outputs. |
| Main limitation | High maintenance effort, slower scaling, and delayed insights | Cloud cost, vendor dependency, governance complexity, and skills requirements | Tool selection should match data sensitivity, budget, speed, and governance maturity. |
| Environment or Approach | Advantages | Disadvantages | Suitable Use |
|---|---|---|---|
| On-premises Hadoop/MapReduce | Provides local control over infrastructure and can process very large datasets using distributed storage and computation [71]. | Requires investment of hardware, cluster administration, tuning and specialized staff. It is less flexible when workloads change quickly. | Large historical datasets, batch jobs and organizations with strict internal infrastructure requirements [6,71]. |
| Parallel DBMS and data warehouse | Offers mature query optimization, structured reporting and strong support for business intelligence. | Less suitable for highly unstructured, semi-structured or rapidly changing data sources. It may also be expensive on a very large scale. | Structured reporting, OLAP workloads and controlled enterprise data environments [13,117]. |
| Cloud analytics | Provides elastic computing, managed services and faster deployment without full ownership of physical infrastructure. | Creates concerns related to provider dependency, data transfer costs, compliance and cloud security. | Organizations requiring scalable analytics services and variable workloads [37,68]. |
| Stream processing | Supports near-real-time monitoring, anomaly detection and rapid response to changing events. | Requires careful management of latency, fault tolerance, ordering, state and continuous data quality. | Fraud detection, IoT monitoring, cybersecurity, transportation and time-sensitive decisions [42,124]. |
| Lakehouse architecture | Combines data lake flexibility with warehouse-style management, supporting BI, data science and machine learning on shared data. | Still requires strong governance, metadata management and architecture maturity to avoid turning into an unmanaged data lake. | Unified analytics platforms, mixed structured/unstructured data and AI-ready data environments [13,52]. |
| MLOps and DataOps | Improves model deployment, monitoring, reproducibility and collaboration between data science and operations teams. | Adds process complexity and requires automation maturity, version control and continuous monitoring. | Production analytics, deployed ML systems and reliable model lifecycle management [12,32]. |
| Integrated governance | Clarifies ownership, standards, access rights, quality controls and accountability across the analytics lifecycle. | May be difficult to implement uniformly across departments, platforms and external data sources. | Regulated sectors, multi-platform environments and high-impact analytics decisions [15,33]. |
| Architecture Generation | Dominant Design Logic | Typical Limitation | Current Relevance |
|---|---|---|---|
| Enterprise data warehouse | Centralized, structured and schema-driven analytical repository. | Less flexible for unstructured, streaming and high-variety data. | Still useful for governed reporting and stable business intelligence [13,114]. |
| MapReduce and Hadoop ecosystem | Distributed storage and batch processing on commodity clusters. | High latency and limited suitability for iterative machine learning and interactive analytics. | Important historical foundation for large-scale processing [6,43]. |
| In-memory and iterative processing | Faster processing through memory-based computation and reusable working sets. | Requires careful resource management and cluster tuning. | Supports machine learning, iterative analytics and large-scale data engineering [45,47]. |
| Cloud-native analytics | Elastic services, managed infrastructure and scalable storage-compute separation. | Vendor dependency, cost governance and data sovereignty concerns. | Dominant model for flexible analytics deployment and rapid scaling [67,68]. |
| Streaming and real-time analytics | Continuous processing of events, logs, transactions and sensor streams. | Requires fault tolerance, low-latency design and continuous monitoring. | Essential for fraud detection, IoT, cybersecurity and operational decision making [42,124]. |
| Lakehouse and AI-enabled platforms | Integration of data lake flexibility with warehouse governance and AI/ML workloads. | Still evolving in relation to standards, governance and performance benchmarking. | Promising direction for unified analytics, machine learning and data governance [13,52]. |
| Evaluation Domain | Illustrative Indicators | Validation Method | Expected Evidence |
|---|---|---|---|
| Technical architecture | Latency, throughput, reliability, interoperability, resource use, cost | Benchmarking; proof-of-concept deployment | Evidence that the selected architecture meets workload and timing requirements |
| Analytics and AI | Accuracy or utility, robustness, drift, interpretability, fairness | Offline validation; stress testing; model monitoring | Evidence that analytical outputs remain useful, stable, and understandable |
| Human and organizational | Usability, trust, data literacy, role clarity, readiness, adoption | Expert review; user studies; surveys | Evidence that users can interpret, challenge, and act on analytical outputs |
| Governance, security and privacy | Auditability, access control, incidents, compliance, privacy risk, accountability | Control assessment; audit; threat/privacy review | Evidence that data and models are used within acceptable risk and accountability boundaries |
| Decision process | Decision speed, quality, error rate, adoption, reversibility | Scenario tests; controlled studies; case comparison | Evidence that architecture supports better decision processes in the target context |
| Organizational outcomes | Cost, productivity, risk reduction, service quality, resilience, realized value | Longitudinal case study; before/after comparison | Evidence of sustained value beyond technical performance |
| Framework Stream | Primary Focus | Typical Limitation | Distinctive Contribution of the Proposed Framework |
|---|---|---|---|
| Technology-oriented BDA frameworks | Data acquisition, distributed storage, processing engines, and analytical tools. | Limited integration with managerial decision stages, embedded governance, and post-deployment learning. | Links the technical analytics lifecycle directly to decision formulation, implementation, monitoring, and feedback [24,26,37]. |
| BDA capability frameworks | Technological, human, organizational, and intangible resources that enable analytics value. | Explain required capabilities more clearly than the operational interactions among pipelines, models, governance, and decisions. | Translates capability requirements into an end-to-end socio-technical reference architecture [7,18,30,35]. |
| Lakehouse and cloud-native architectures | Unified storage, reliable data management, scalable computation, and shared BI/ML workloads. | Primarily platform-centered; decision processes, human interpretation, and organizational learning remain outside the core architecture. | Use Lakehouse or cloud services as one architectural layer within a wider governed decision system [13,52,66,68]. |
| Stream-processing frameworks | Continuous ingestion and low-latency processing for event-driven analytics. | Optimize processing speed but do not fully specify how real-time outputs are interpreted, governed, acted upon, and evaluated. | Combines batch, streaming, and hybrid processing with interpretation, human oversight, and outcome monitoring [42,45,46]. |
| MLOps and model-lifecycle frameworks | Model deployment, versioning, reproducibility, monitoring, maintenance, and technical reliability. | Model-centered scope may underrepresent data governance, managerial choice, accountability, and feedback from implemented decisions. | Connects DataOps/MLOps with explainability, governance, decision accountability, and closed-loop adaptation [12,15,32]. |
| Data and responsible-AI governance frameworks | Ownership, metadata, quality, privacy, security, fairness, accountability, and compliance. | Governance is often treated as a specialized control domain rather than a mechanism embedded across the entire decision architecture. | Positions governance and human oversight as cross-cutting controls operating across all technical and managerial layers [15,34,65,69]. |
| Technology-adoption and organizational-readiness models | Knowledge, perceived benefits and barriers, leadership, human capability, financial resources, engagement, and contextual conditions shaping adoption. | Often explain willingness or readiness to adopt a technology without specifying how adoption conditions interact with end-to-end BDA architecture and post-decision learning. | Uses readiness, task characteristics, and institutional conditions as configurational factors that moderate how the seven-stage architecture is adopted and converted into decision value [127,128]. |
| Framework Level | Illustrative Open-Source Options | Cost-Conscious Implementation Approach | Practical Challenges |
|---|---|---|---|
| Data sources | PostgreSQL, MariaDB, Apache Cassandra, object/file stores | Reuse existing operational sources; expose only required data through controlled connectors | Data ownership, source quality, schema drift, and access permissions |
| Ingestion and integration | Apache Kafka, Apache NiFi, Airbyte | Start with scheduled ingestion; add streaming only where decision latency requires it | Connector maintenance, semantic mismatch, duplicate events, metadata capture |
| Storage and platform | HDFS, MinIO, Apache Iceberg, Delta Lake | Use commodity or cloud object storage; separate hot and archival data | Capacity planning, backup, data lifecycle, metadata and governance overhead |
| Processing | Apache Spark, Apache Flink | Use batch processing for non-urgent workloads; reserve streaming resources for time-sensitive cases | Cluster tuning, state management, fault tolerance, compute cost |
| Analytics and AI | Python/scikit-learn, H2O.ai, MLflow, Kubeflow | Begin with interpretable baseline models; automate deployment only after stable validation | Model drift, reproducibility, explainability, skills, GPU/compute demand for advanced AI |
| Visualization and interpretation | Apache Superset, Grafana | Provide role-specific dashboards and alerts rather than broad unrestricted reporting | Poor visual design, information overload, misinterpretation of uncertainty |
| Governance and control | Apache Atlas, OpenMetadata, Keycloak, OpenBao | Prioritize identity/access control, metadata, lineage, logging, and high-risk data first | Configuration complexity, policy ownership, audit effort, cross-platform consistency |
| Decision, action and learning | Workflow tools plus existing ticketing/ERP/BI processes | Integrate recommendations into existing approval and monitoring processes before adding automation | Change management, accountability, adoption, feedback quality |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Alma’aitah, W.Z.; AL-Aswadi, F.N.; Quraan, A.; Karim, N.A.; Alahmer, H.; Mustafa, M.Y. Opportunities and Challenges in Big Data Analytics for Decision Making: An Integrated Framework. Computers 2026, 15, 584. https://doi.org/10.3390/computers15090584
Alma’aitah WZ, AL-Aswadi FN, Quraan A, Karim NA, Alahmer H, Mustafa MY. Opportunities and Challenges in Big Data Analytics for Decision Making: An Integrated Framework. Computers. 2026; 15(9):584. https://doi.org/10.3390/computers15090584
Chicago/Turabian StyleAlma’aitah, Wafa’ Za’al, Fatima N. AL-Aswadi, Addy Quraan, Nader Abdel Karim, Hussein Alahmer, and Mohamad Y. Mustafa. 2026. "Opportunities and Challenges in Big Data Analytics for Decision Making: An Integrated Framework" Computers 15, no. 9: 584. https://doi.org/10.3390/computers15090584
APA StyleAlma’aitah, W. Z., AL-Aswadi, F. N., Quraan, A., Karim, N. A., Alahmer, H., & Mustafa, M. Y. (2026). Opportunities and Challenges in Big Data Analytics for Decision Making: An Integrated Framework. Computers, 15(9), 584. https://doi.org/10.3390/computers15090584

