Distilling the Complexity of Agent-Based Simulations into Textual Explanations via Large Language Models
Abstract
1. Introduction

- We design, implement, and evaluate a staged design-of-experiments (DoE) methodology to analyze and optimize multi-component LLM-based reporting systems. By combining aggregated ANOVA screening, factorial variance decomposition, and direct level-dominance comparisons, the approach attributes performance variability to individual factors and interaction effects, thereby providing a systematic and generalizable workflow for configuring simulation-to-text methods.
- We analyze how summarization strategies, prompt elements, and the identity of the large language model jointly influence the quality of generated reports, showing in particular that interactions between summarization algorithms and LLMs explain a substantial share of performance variability.
- We perform a comparative empirical evaluation across multiple ABMs using four recent large language models (Google’s Gemini 3.1 Pro Preview, Alibaba’s Qwen 3.5 27B, MoonshotAI’s Kimi K2.5, and Anthropic’s Claude Opus 4.6), providing practical guidance on evidence presentation, summarization choices, and model selection for simulation reporting tasks.
2. Background
2.1. Conveying Agent-Based Models to End-Users
2.2. LLMs and Simulation
2.3. Evaluating LLM-Generated Reports
2.4. Prompt Engineering and Experimental Design
3. Materials and Methods
3.1. Overview
3.2. Data Extraction
3.3. Report Generation by Prompt Engineering
3.3.1. Prompts and Factorial Design


3.3.2. Post-Processing
3.4. Evaluation: Comparing LLM Outputs with Reports and Quantifying Readability
- A summary was jointly written by the authors for each ABM to serve as a ground truth reference. We wrote the summary given the same information as was available to the LLMs, including the time series from each simulation run. These summaries were all written prior to generating text via LLMs, thus ensuring that the ground truth reference provided by our summary was not ‘contaminated’ by sentences that we may read in the LLM output. We refer to these summaries as ‘authors’ in our results and on the files provided in Supplementary Material S2.
- A summary was written for each ABM by independent associate professors with several peer-reviewed publications on agent-based modeling. That is, the summary for each ABM was written by a different modeler. To ensure consistency in the task, we created the same onboarding form for the modelers, varying only the content of the ABM. From hereon, these summaries are labeled as ‘modelers’ in our results and in Supplementary Material S2, which also contains the onboarding forms.
- We generated a full report and a summary as follows. We adapt the prompts shown in prior sections to rely on articles instead of the ABM documentation, and we use a state-of-the-art LLM that we did not use in our experimental setup (to avoid the output of an LLM being compared to itself) to generate the description of each plot. As a means to also demonstrate that reports could be produced without documentation from the ComSES.net repository or the NetLogo code, we used contextual augmentation by providing the peer-reviewed article for each ABM in the prompt. We corrected the descriptions of each plot to ensure factuality. Then, these descriptions were combined into a long report (note that it is the only material for comparison that is not a summary), which we call ‘long LLM’ from hereon. Finally, these reports were summarized into a short form using the same state-of-the-art LLM, resulting in the document named ‘short LLM’. We used OpenAI’s GPT 5.2 Thinking (released in December 2025) for this task. Our prompts and the LLM’s raw responses are shown in Supplementary Material S2, where our manual corrections are highlighted. For the fauna model, we made 6 corrections to the values reported in the curve, and we added two sentences on noticeable trends (out of 14 plots). For the milk model, we expanded two sentences and corrected fifteen values (out of 12 plots). No changes were needed for the grazing model.
3.5. Screening and Optimization Design
4. Results
4.1. Experimental Setup
- The Milk Adoption ABM from Gibson et al. [108], which was designed to evaluate different mechanisms of consumer decision making in milk choice substitution to better understand and replicate historical trends in British milk consumption as a step toward modeling sustainable dietary shifts.
- The extended RAGE model [109], which explores how different types of learning (learning-by-doing and social learning) and varying attitudes toward risk influence smallholder farmers’ decisions about livestock management and the resulting outcomes for their livelihoods and the environment.
4.2. Screening Stage: Moonshot AI’s Kimi K2.5 and Alibaba’s Qwen 3.5 27B
4.2.1. Impact of Experimental Design Factors
4.2.2. Optimal Performances
- We aggregated ANOVA results across evaluation metrics and reference formats to determine how often each experimental factor shows a statistically significant effect. The goal is to establish whether a factor reliably alters evaluation outcomes. For example, the choice of summarization algorithm was statistically significant for all metrics and across all reference summaries. This indicates that changing the summarizer systematically changes report quality, making it a primary factor that would be varied during our next stage of optimization.
- We aggregated factorial contribution tables that quantify how much variance each factor explains when both main effects and pairwise interactions are considered. For instance, the role instruction does not always appear dominant in the ANOVA tables, yet its interaction with the summarization algorithm or with the LLM can account for a substantial portion of performance variability. This indicates that the effect of role is conditional: it shapes outcomes indirectly by modifying how other components operate.
- To translate statistical influence into actionable design choices, we conducted direct level-vs.-level comparisons (Figure 9). For each factor, we compared competing levels under otherwise identical experimental conditions (same ABM, LLM, evidence format, prompt configuration, and replication). We then counted how often each level produced the best evaluation score.
4.3. Optimization Stage: Anthropic’s Claude Opus 4.6 and Google’s Gemini 3.1 Pro Preview
5. Discussion
5.1. Key Contributions
5.2. Practical Implications
5.3. Limitations and Future Work
6. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| ABM | Agent-Based Model |
| LLM | Large Language Model |
| M&S | Modeling and Simulation |
| ODD | Overview, Design concepts, Details |
References
- Zellner, M.L.; Milz, D.; Lyons, L.; Hoch, C.; Radinsky, J. Finding the balance between simplicity and realism in participatory modeling for environmental planning. Environ. Model. Softw. 2022, 157, 105481. [Google Scholar] [CrossRef] [Scilit]
- Grigoryan, G.; Collins, A.J. Feature Importance for Uncertainty Quantification in Agent-Based Modeling. In Proceedings of the 2023 Winter Simulation Conference (WSC); IEEE: New York, NY, USA, 2023; pp. 233–242. [Google Scholar]
- Jalali, M.S.; Beaulieu, E. Strengthening a weak link: Transparency of causal loop diagrams—Current state and recommendations. Syst. Dyn. Rev. 2024, 40, e1753. [Google Scholar] [CrossRef] [Scilit]
- Allender, S.; Owen, B.; Kuhlberg, J.; Lowe, J.; Nagorcka-Smith, P.; Whelan, J.; Bell, C. A community based systems diagram of obesity causes. PLoS ONE 2015, 10, e0129683. [Google Scholar] [CrossRef] [Scilit]
- Siokou, C.; Morgan, R.; Shiell, A. Group model building: A participatory approach to understanding and acting on systems. Public Health Res. Pract. 2014, 25, e2511404. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jack, A. Foresight report on obesity–Author’s reply. Lancet 2007, 370, 1755. [Google Scholar] [CrossRef] [Scilit]
- van Veen, B.L.; Roland Ortt, J. Simplification errors in predictive models. Futur. Foresight Sci. 2024, 6, e184. [Google Scholar] [CrossRef] [Scilit]
- Swarup, S. Adequacy: What makes a simulation good enough? In Proceedings of the 2019 Spring Simulation Conference (SpringSim); IEEE: New York, NY, USA, 2019; pp. 1–12. [Google Scholar]
- Uleman, J.F.; Stronks, K.; Rutter, H.; Arah, O.A.; Rod, N.H. Mapping complex public health problems with causal loop diagrams. Int. J. Epidemiol. 2024, 53, dyae091. [Google Scholar] [CrossRef] [Scilit]
- Hedelin, B.; Gray, S.; Woehlke, S.; BenDor, T.K.; Singer, A.; Jordan, R.; Zellner, M.; Giabbanelli, P.; Glynn, P.; Jenni, K.; et al. What’s left before participatory modeling can fully support real-world environmental planning processes: A case study review. Environ. Model. Softw. 2021, 143, 105073. [Google Scholar] [CrossRef] [Scilit]
- Wei, Y.; Knoeferle, P. Causal inference: Relating language to event representations and events in the world. Front. Psychol. 2023, 14, 1172928. [Google Scholar] [CrossRef] [Scilit]
- Shrestha, A.; Mielke, K.; Nguyen, T.A.; Giabbanelli, P.J. Automatically explaining a model: Using deep neural networks to generate text from causal maps. In Proceedings of the 2022 Winter Simulation Conference (WSC); IEEE: New York, NY, USA, 2022; pp. 2629–2640. [Google Scholar]
- Phatak, A.; Mago, V.K.; Agrawal, A.; Inbasekaran, A.; Giabbanelli, P.J. Narrating causal graphs with large language models. In Proceedings of the Hawaii International Conference on System Sciences 2024 (HICSS-57), Honolulu, HI, USA, 3–6 January 2024. [Google Scholar]
- Gandee, T.J.; Giabbanelli, P.J. Combining Natural Language Generation and Graph Algorithms to Explain Causal Maps Through Meaningful Paragraphs. In Proceedings of the International Conference on Conceptual Modeling; Springer: Berlin/Heidelberg, Germany, 2024; pp. 359–376. [Google Scholar]
- Giabbanelli, P.J. GPT-based models meet simulation: How to efficiently use large-scale pre-trained language models across simulation tasks. In Proceedings of the 2023 Winter Simulation Conference (WSC); IEEE: New York, NY, USA, 2023; pp. 2920–2931. [Google Scholar]
- Schoenberg, W. Building and Learning With Models Using AI. Syst. Dyn. Rev. 2026, 42, e70019. [Google Scholar] [CrossRef] [Scilit]
- Hosseinichimeh, N.; Majumdar, A.; Williams, R.; Ghaffarzadegan, N. From text to map: A system dynamics bot for constructing causal loop diagrams. Syst. Dyn. Rev. 2024, 40, e1782. [Google Scholar] [CrossRef] [Scilit]
- Veldhuis, G.; van Wermeskerken, F.; Visker, O.; Steinmann, P.; Deuten, S.; van Waas, R. From Data to Model Structure: A Generative Algorithm to Develop System Dynamics Models. SSRN 2024. Available online: https://ssrn.com/abstract=4921086 (accessed on 1 April 2026).
- Giabbanelli, P.J.; Gandee, T.J.; Agrawal, A.; Hosseinichimeh, N. Benchmarking and Assessing Transformations Between Text and Causal Maps via Large Language Models. Appl. Ontol. 2025, 20, 125–134. [Google Scholar] [CrossRef] [Scilit]
- Verduzco, J.C.; Holbrook, E.; Strachan, A. GPT-4 as an interface between researchers and computational software: Improving usability and reproducibility. arXiv 2023, arXiv:2310.11458. [Google Scholar] [CrossRef] [Scilit]
- Méndez, G.; Gervás, P. Using ChatGPT for story sifting in narrative generation. In Proceedings of the 14th International Conference on Computational Creativity, Waterloo, ON, Canada, 19–23 June 2023. [Google Scholar]
- Aoki, N.; Mori, N.; OKada, M. Analysis of LLM-based narrative generation using the agent-based simulation. In Proceedings of the 2023 15th International Congress on Advanced Applied Informatics Winter (IIAI-AAI-Winter); IEEE: New York, NY, USA, 2023; pp. 284–289. [Google Scholar]
- Achter, S.; Borit, M.; Cottineau, C.; Meyer, M.; Polhill, J.G.; Radchuk, V. How to conduct more systematic reviews of agent-based models and foster theory development-Taking stock and looking ahead. Environ. Model. Softw. 2024, 173, 105867. [Google Scholar] [CrossRef] [Scilit]
- DeAngelis, D.L.; Diaz, S.G. Decision-making in agent-based modeling: A current review and future prospectus. Front. Ecol. Evol. 2019, 6, 237. [Google Scholar] [CrossRef] [Scilit]
- Bianchi, F.; Squazzoni, F. Agent-based models in sociology. Wiley Interdiscip. Rev. Comput. Stat. 2015, 7, 284–306. [Google Scholar] [CrossRef] [Scilit]
- Qiu, L.; Phang, R. Agent-Based Modeling in Political Decision Making. In Oxford Research Encyclopedia of Politics; Oxford Academic: Oxford, UK, 2020. [Google Scholar]
- Lee, J.S.; Filatova, T.; Ligmann-Zielinska, A.; Hassani-Mahmooei, B.; Stonedahl, F.; Lorscheid, I.; Voinov, A.; Polhill, J.G.; Sun, Z.; Parker, D.C. The complexities of agent-based modeling output analysis. J. Artif. Soc. Soc. Simul. 2015, 18, 4. [Google Scholar] [CrossRef] [Scilit]
- An, L.; Grimm, V.; Turner, B.L., II. Meeting grand challenges in agent-based models. J. Artif. Soc. Soc. Simul. 2020, 23, 13. [Google Scholar] [CrossRef] [Scilit]
- Papadopoulou, L.; McEntaggart, K.; Etienne, J. Communicating Scientific Uncertainty in Advice Provision to Decision-Makers: Review of Approaches and Recommendations for UK Statutory Nature Conservation Bodies; JNCC Report No. 617; Joint Nature Conservation Committee: Peterborough, UK, 2018. Available online: https://data.jncc.gov.uk/data/7ade8ea1-c616-4bc1-ac0d-b8c2fb9d1d6e/JNCC-Report-617-FINAL-WEB.pdf (accessed on 11 March 2026).
- Sun, Z.; Lorscheid, I.; Millington, J.D.; Lauf, S.; Magliocca, N.R.; Groeneveld, J.; Balbi, S.; Nolzen, H.; Müller, B.; Schulze, J.; et al. Simple or complicated agent-based models? A complicated issue. Environ. Model. Softw. 2016, 86, 56–67. [Google Scholar] [CrossRef] [Scilit]
- Giabbanelli, P.J.; Daumas, C.; Flandre, N.Y.; Pitkar, A.; Vazquez-Estrada, J. Promoting empathy in decision-making by turning agent-based models into stories using large-language models. J. Simul. 2025, 1–21. [Google Scholar] [CrossRef] [Scilit]
- Grimm, V.; Railsback, S.F.; Vincenot, C.E.; Berger, U.; Gallagher, C.; DeAngelis, D.L.; Edmonds, B.; Ge, J.; Giske, J.; Groeneveld, J.; et al. The ODD protocol for describing agent-based and other simulation models: A second update to improve clarity, replication, and structural realism. J. Artif. Soc. Soc. Simul. 2020, 23, 7. [Google Scholar] [CrossRef] [Scilit]
- Müller, B.; Bohn, F.; Dreßler, G.; Groeneveld, J.; Klassert, C.; Martin, R.; Schlüter, M.; Schulze, J.; Weise, H.; Schwarz, N. Describing human decisions in agent-based models–ODD+ D, an extension of the ODD protocol. Environ. Model. Softw. 2013, 48, 37–48. [Google Scholar] [CrossRef] [Scilit]
- Laatabi, A.; Marilleau, N.; Nguyen-Huu, T.; Hbid, H.; Ait Babram, M. ODD+ 2D: An ODD based protocol for mapping data to empirical ABMs. J. Artif. Soc. Soc. Simul. 2018, 21, 9. [Google Scholar] [CrossRef] [Scilit]
- Achter, S.; Borit, M.; Chattoe-Brown, E.; Siebers, P.O. RAT-RS: A reporting standard for improving the documentation of data use in agent-based modelling. Int. J. Soc. Res. Methodol. 2022, 25, 517–540. [Google Scholar] [CrossRef] [Scilit]
- Monks, T.; Currie, C.S.; Onggo, B.S.; Robinson, S.; Kunc, M.; Taylor, S.J. Strengthening the reporting of empirical simulation studies: Introducing the STRESS guidelines. J. Simul. 2019, 13, 55–67. [Google Scholar] [CrossRef] [Scilit]
- Giabbanelli, P.J.; Tison, B.; Keith, J. The application of modeling and simulation to public health: Assessing the quality of agent-based models for obesity. Simul. Model. Pract. Theory 2021, 108, 102268. [Google Scholar] [CrossRef] [Scilit]
- Hall, A.; Virrantaus, K. Visualizing the workings of agent-based models: Diagrams as a tool for communication and knowledge acquisition. Comput. Environ. Urban Syst. 2016, 58, 1–11. [Google Scholar] [CrossRef] [Scilit]
- Szangolies, L.; Rohwäder, M.S.; Ahmed, H.; Jahanmiri, F.; Wagner, A.; Souto-Veiga, R.; Grimm, V.; Gallagher, C. Visual ODD: A Standardised Visualisation Illustrating the Narrative of Agent-Based Models. J. Artif. Soc. Soc. Simul. 2024, 27, 1. [Google Scholar] [CrossRef] [Scilit]
- Dorin, A.; Geard, N. The practice of agent-based model visualization. Artif. Life 2014, 20, 271–289. [Google Scholar] [CrossRef] [Scilit]
- Kornhauser, D.; Wilensky, U.; Rand, W. Design guidelines for agent based model visualization. J. Artif. Soc. Soc. Simul. 2009, 12, 1. [Google Scholar]
- Vandin, A.; Giachini, D.; Lamperti, F.; Chiaromonte, F. Automated and distributed statistical analysis of economic agent-based models. J. Econ. Dyn. Control 2022, 143, 104458. [Google Scholar] [CrossRef] [Scilit]
- Hooten, M.; Wikle, C.; Schwob, M. Statistical implementations of agent-based demographic models. Int. Stat. Rev. 2020, 88, 441–461. [Google Scholar] [CrossRef] [Scilit]
- Gürcan, Ö. Llm-augmented agent-based modelling for social simulations: Challenges and opportunities. In HHAI 2024: Hybrid Human AI Systems for the Social Good; IOS Press: Amsterdam, The Netherlands, 2024; pp. 134–144. [Google Scholar]
- Ferraro, A.; Galli, A.; La Gatta, V.; Postiglione, M.; Orlando, G.M.; Russo, D.; Riccio, G.; Romano, A.; Moscato, V. Agent-Based Modelling Meets Generative AI in Social Network Simulations. In Proceedings of the International Conference on Advances in Social Networks Analysis and Mining; Springer: Berlin/Heidelberg, Germany, 2024; pp. 155–170. [Google Scholar]
- Zhong, S.; Japkowicz, N.; Giabbanelli, P. Do we Still Need People? Comparing Human and LLM Personas in Political Modeling and Simulation. In Proceedings of the 2025 ACM/IEEE 28th International Conference on Model Driven Engineering Languages and Systems Companion (MODELS-C); IEEE: New York, NY, USA, 2025; pp. 512–521. [Google Scholar]
- Ghaffarzadegan, N.; Majumdar, A.; Williams, R.; Hosseinichimeh, N. Generative agent-based modeling: An introduction and tutorial. Syst. Dyn. Rev. 2024, 40, e1761. [Google Scholar] [CrossRef] [Scilit]
- Larooij, M.; Törnberg, P. Validation is the central challenge for generative social simulation: A critical review of LLMs in agent-based modeling. Artif. Intell. Rev. 2025, 59, 15. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Acharya, D.B.; Kuppan, K.; Divya, B. Agentic AI: Autonomous Intelligence for Complex Goals–A Comprehensive Survey. IEEE Access 2025, 13, 18912–18936. [Google Scholar] [CrossRef] [Scilit]
- Padilla, J.J.; Shuttleworth, D.; O’Brien, K. Agent-based model characterization using natural language processing. In Proceedings of the 2019 Winter Simulation Conference (WSC); IEEE: New York, NY, USA, 2019; pp. 560–571. [Google Scholar]
- Shuttleworth, D.; Padilla, J.J. Towards semi-automatic model specification. In Proceedings of the 2021 Winter Simulation Conference (WSC); IEEE: New York, NY, USA, 2021; pp. 1–12. [Google Scholar]
- Khatami, S. AI-Enhanced ABM Development: Facilitating Agent-Based Modeling Using Artificial Intelligence. Ph.D. Thesis, Norwegian University of Science and Technology, Trondheim, Norway, 2025. [Google Scholar]
- Davis, C.W.; Jetter, A.J.; Giabbanelli, P.J. Towards an Automatic Construction of Simulation Scenarios: A Systematic Review. In Proceedings of the 2023 Winter Simulation Conference (WSC); IEEE: New York, NY, USA, 2023; pp. 2494–2505. [Google Scholar]
- Martínez, J.; Llinas, B.; Botello, J.G.; Padilla, J.J.; Frydenlund, E. Enhancing GPT-3.5’s Proficiency in Netlogo Through Few-Shot Prompting and Retrieval-Augmented Generation. In Proceedings of the 2024 Winter Simulation Conference (WSC); IEEE: New York, NY, USA, 2024; pp. 666–677. [Google Scholar]
- Frydenlund, E.; Martínez, J.; Padilla, J.J.; Palacio, K.; Shuttleworth, D. Modeler in a box: How can large language models aid in the simulation modeling process? Simulation 2024, 100, 727–749. [Google Scholar] [CrossRef] [Scilit]
- Akhavan, A.; Jalali, M.S. Generative AI and simulation modeling: How should you (not) use large language models like ChatGPT. Syst. Dyn. Rev. 2024, 40, e1773. [Google Scholar] [CrossRef] [Scilit]
- Jackson, I.; Rolf, B. Do Natural Language Processing models understand simulations? Application of GPT-3 to translate simulation source code to English. IFAC-Pap. 2023, 56, 221–226. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Yu, P.S.; Zhang, J. A systematic survey of text summarization: From statistical methods to large language models. ACM Comput. Surv. 2025, 57, 277. [Google Scholar] [CrossRef] [Scilit]
- Koh, H.Y.; Ju, J.; Liu, M.; Pan, S. An empirical survey on long document summarization: Datasets, models, and metrics. ACM Comput. Surv. 2022, 55, 154. [Google Scholar] [CrossRef] [Scilit]
- Miller, D. Leveraging BERT for extractive text summarization on lectures. arXiv 2019, arXiv:1906.04165. [Google Scholar] [CrossRef] [Scilit]
- Gandee, T.J.; Giabbanelli, P.J. Faithful Narratives from Complex Conceptual Models: Should Modelers or Large Language Models Simplify Causal Maps? Mach. Learn. Knowl. Extr. 2025, 7, 116. [Google Scholar] [CrossRef] [Scilit]
- Beltagy, I.; Peters, M.E.; Cohan, A. Longformer: The long-document transformer. arXiv 2020, arXiv:2004.05150. [Google Scholar] [CrossRef] [Scilit]
- Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; Liu, P.J. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 2020, 21, 140. [Google Scholar]
- Xia, T.C.; Bertini, F.; Montesi, D. Large Language Models Evaluation for PubMed Extractive Summarisation. ACM Trans. Comput. Healthc. 2026, 7, 4. [Google Scholar] [CrossRef] [Scilit]
- Gunnu, V.; Shah, S.; Minukuri, A.; Gopu, J. Text Summarization. In Practical Solutions for Modern NLP Challenges: Mastering LLMs and SLMs for Real-World NLP in Cloud and Open-Source; Springer: Berlin/Heidelberg, Germany, 2026; pp. 281–323. [Google Scholar]
- August, T.; Lo, K.; Smith, N.A.; Reinecke, K. Know Your Audience: The benefits and pitfalls of generating plain language summaries beyond the “general” audience. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA, 11–16 May 2024; pp. 1–26. [Google Scholar]
- Giabbanelli, P.J.; Agrawal, A. Towards Personalized Explanations for Health Simulations: A Mixed-Methods Framework for Stakeholder-Centric Summarization. In Proceedings of the AAAI Symposium Series; AAAI: Washington, DC, USA, 2025; Volume 7, pp. 506–515. [Google Scholar]
- National Academies of Sciences, Engineering, and Medicine. Executive Summary. In Modeling Human and Organizational Behavior: Application to Military Simulations; The National Academies Press: Washington, DC, USA, 1998. [Google Scholar] [CrossRef] [Scilit]
- National Academies of Sciences, Engineering, and Medicine. Executive Summary. In Behavioral Modeling and Simulation: From Individuals to Societies; The National Academies Press: Washington, DC, USA, 2008. [Google Scholar] [CrossRef] [Scilit]
- Kho, V. McKinsey Executive Summaries: Short but Powerful. Writ. Univ. Beyond J. First-Year Stud. Writ. UTM 2022, 2, 87–95. [Google Scholar]
- Harvard Kennedy School. How to Write an Executive Summary. Writing Workshop with Lauren Brodsky, Communications Program, Harvard Kennedy School, Cambridge, MA, USA, 25 February 2026; Available online: https://www.hks.harvard.edu/events/how-write-executive-summary-0 (accessed on 11 March 2026).
- Naismith, B.; Mulcaire, P.; Burstein, J. Automated evaluation of written discourse coherence using GPT-4. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023), Toronto, ON, Canada, 13 July 2023; pp. 394–403. [Google Scholar]
- Moore, S.; Nguyen, H.A.; Chen, T.; Stamper, J. Assessing the quality of multiple-choice questions using gpt-4 and rule-based methods. In Proceedings of the European Conference on Technology Enhanced Learning; Springer: Berlin/Heidelberg, Germany, 2023; pp. 229–245. [Google Scholar]
- Fischer, T.; Remus, S.; Biemann, C. Measuring faithfulness of abstractive summaries. In Proceedings of the 18th Conference on Natural Language Processing (KONVENS 2022), Potsdam, Germany, 12–15 September 2022; pp. 63–73. [Google Scholar]
- Gatt, A.; Krahmer, E. Survey of the state of the art in natural language generation: Core tasks, applications and evaluation. J. Artif. Intell. Res. 2018, 61, 65–170. [Google Scholar] [CrossRef] [Scilit]
- Li, M.; Gao, Q.; Yu, T. Kappa statistic considerations in evaluating inter-rater reliability between two raters: Which, when and context matters. BMC Cancer 2023, 23, 799. [Google Scholar] [CrossRef] [Scilit]
- Raghavan, P.; Fosler-Lussier, E.; Lai, A.M. Inter-annotator reliability of medical events, coreferences and temporal relations in clinical narratives by annotators with varying levels of clinical expertise. In Proceedings of the AMIA Annual Symposium, Chicago, IL, USA, 3–7 November 2012; p. 1366. [Google Scholar]
- van Schaik, T.A.; Pugh, B. A field guide to automatic evaluation of llm-generated summaries. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, Washington, DC, USA, 14–18 July 2024; pp. 2832–2836. [Google Scholar]
- Papineni, K.; Roukos, S.; Ward, T.; Zhu, W.J. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, Philadelphia, PA, USA, 6–12 July 2002; pp. 311–318. [Google Scholar]
- Chin-Yew, L. Rouge: A package for automatic evaluation of summaries. In Proceedings of the Workshop on Text Summarization Branches Out, Barcelona, Spain, 25–26 July 2004; pp. 74–81. [Google Scholar]
- Myla, S.D.; Saini, E.R.; Kapoor, E.N. Auto Text Summarization in Natural Language Processing. In Proceedings of the 2024 2nd International Conference on Intelligent Data Communication Technologies and Internet of Things (IDCIoT); IEEE: New York, NY, USA, 2024; pp. 1258–1267. [Google Scholar]
- Caparrós-Laiz, C.; García-Díaz, J.A.; Valencia-García, R. Evaluating extractive automatic text summarization techniques in spanish. In Proceedings of the International Conference on Technologies and Innovation; Springer: Berlin/Heidelberg, Germany, 2021; pp. 79–92. [Google Scholar]
- Roy, S.S.; Mercer, R.E. Generating extractive and abstractive summaries in parallel from scientific articles incorporating citing statements. In Proceedings of the 4th New Frontiers in Summarization Workshop, Singapore, 6–10 December 2023; pp. 75–86. [Google Scholar]
- Guo, Y.; Qiu, W.; Wang, Y.; Cohen, T. Automated lay language summarization of biomedical scientific reviews. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2021; Volume 35, pp. 160–168. [Google Scholar]
- Cooperman, S.R.; Brandão, R.A. Investigating the proficiency of an AI tool in summarizing foot and ankle literature: A quantitative, qualitative and accuracy analysis. Foot Ankle Surg. Tech. Rep. Cases 2024, 4, 100384. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Xu, G.; Ren, M. LLM-Generated Natural Language Meets Scaling Laws: New Explorations and Data Augmentation Methods. arXiv 2024, arXiv:2407.00322. [Google Scholar] [CrossRef] [Scilit]
- Chen, B.; Zhang, Z.; Langrené, N.; Zhu, S. Unleashing the potential of prompt engineering for large language models. Patterns 2025, 6, 101260. [Google Scholar] [CrossRef] [Scilit]
- Schmidt, D.C.; Spencer-Smith, J.; Fu, Q.; White, J. Towards a catalog of prompt patterns to enhance the discipline of prompt engineering. ACM SIGAda Ada Lett. 2024, 43, 43–51. [Google Scholar] [CrossRef] [Scilit]
- Braun, M.; Greve, M.; Kegel, F.; Kolbe, L.M.; Beyer, P.E. Can (A)I have a word with you? A taxonomy on the design dimensions of AI prompts. In Proceedings of the 57th Annual Hawaii International Conference on System Sciences, HICSS 2024. Hawaii International Conference on System Sciences (HICSS), Honolulu, HI, USA, 3–6 January 2024; pp. 559–568. [Google Scholar]
- Memmert, L.; Cvetkovic, I.; Bittner, E. The more is not the merrier: Effects of prompt engineering on the quality of ideas generated by gpt-3. In Proceedings of the 57th Hawaii International Conference on System Sciences, Honolulu, HI, USA, 3–6 January 2024. [Google Scholar]
- Amini, R.; Norouzi, S.S.; Hitzler, P.; Amini, R. Towards complex ontology alignment using large language models. In Proceedings of the International Knowledge Graph and Semantic Web Conference; Springer: Berlin/Heidelberg, Germany, 2024; pp. 17–31. [Google Scholar]
- Gutschmidt, A.; Nast, B. Assessing Model Quality Using Large Language Models. In Proceedings of the IFIP Working Conference on The Practice of Enterprise Modeling; Springer: Berlin/Heidelberg, Germany, 2024; pp. 105–122. [Google Scholar]
- Saeedizade, M.J.; Blomqvist, E. Navigating ontology development with large language models. In Proceedings of the European Semantic Web Conference; Springer: Berlin/Heidelberg, Germany, 2024; pp. 143–161. [Google Scholar]
- Zhao, Y.; Vetter, N.; Aryan, K. Using large language models for ontoclean-based ontology refinement. arXiv 2024, arXiv:2403.15864. [Google Scholar]
- Sanchez, S.M.; Sanchez, P.J.; Wan, H. Work smarter, not harder: A tutorial on designing and conducting simulation experiments. In Proceedings of the 2020 Winter Simulation Conference (WSC); IEEE: New York, NY, USA, 2020; pp. 1128–1142. [Google Scholar]
- Sanchez, S.M. Robust design: Seeking the best of all possible worlds. In Proceedings of the 2000 Winter Simulation Conference Proceedings (Cat. No. 00ch37165); IEEE: New York, NY, USA, 2000; Volume 1, pp. 69–76. [Google Scholar]
- Stahl, M.; Biermann, L.; Nehring, A.; Wachsmuth, H. Exploring LLM prompting strategies for joint essay scoring and feedback generation. In Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024); Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 283–298. [Google Scholar]
- Cuskley, C.; Woods, R.; Flaherty, M. The limitations of large language models for understanding human language and cognition. Open Mind 2024, 8, 1058–1083. [Google Scholar] [CrossRef] [Scilit]
- Abaskohi, A.; Ramesh, A.V.; Nanisetty, S.; Goel, C.; Vazquez, D.; Pal, C.; Gella, S.; Carenini, G.; Laradji, I.H. AgentAda: Skill-Adaptive Data Analytics for Tailored Insight Discovery. arXiv 2025, arXiv:2504.07421. [Google Scholar]
- Pérez, A.S.; Boukhary, A.; Papotti, P.; Lozano, L.C.; Elwood, A. An LLM-based approach for insight generation in data analysis. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 562–582. [Google Scholar]
- Liu, N.F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; Liang, P. Lost in the middle: How language models use long contexts. Trans. Assoc. Comput. Linguist. 2024, 12, 157–173. [Google Scholar] [CrossRef] [Scilit]
- Li, P.; He, Y.; Yashar, D.; Cui, W.; Ge, S.; Zhang, H.; Rifinski Fainman, D.; Zhang, D.; Chaudhuri, S. Table-gpt: Table fine-tuned gpt for diverse table tasks. Proc. Acm Manag. Data 2024, 2, 176. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Li, Y.; Wang, J.; Sun, B.; Ma, W.; Sun, P.; Zhang, M. Large language models as evaluators for recommendation explanations. In Proceedings of the 18th ACM Conference on Recommender Systems, Bari, Italy, 14–18 October 2024; pp. 33–42. [Google Scholar]
- Ni, J.; Pu, J.; Yang, Z.; Zhou, K.; Wang, H.; Xiao, X.; Wang, D.; Li, X.; Luo, J.; Hu, C. From large to super-tiny: End-to-end optimization for cost-efficient LLMs. arXiv 2025, arXiv:2504.13471. [Google Scholar]
- Chen, L.; Zaharia, M.; Zou, J. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv 2024, arXiv:2305.05176. [Google Scholar]
- Yue, M.; Zhao, J.; Zhang, M.; Du, L.; Yao, Z. Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning. In Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Kopels, M.C.; Ullah, I.I. Modeling post-Pleistocene megafauna extinctions as complex social-ecological systems. Quat. Res. 2024, 121, 1–14. [Google Scholar] [CrossRef] [Scilit]
- Gibson, M.; Pereira, J.P.; Slade, R.; Rogelj, J. Agent-based modelling of future dairy and plant-based milk consumption for UK climate targets. J. Artif. Soc. Soc. Simul. 2022, 25, 3. [Google Scholar] [CrossRef] [Scilit]
- Dressler, G.; Groeneveld, J.; Buchmann, C.M.; Guo, C.; Hase, N.; Thober, J.; Frank, K.; Müller, B. Implications of behavioral change for the resilience of pastoral systems—Lessons from an agent-based model. Ecol. Complex. 2019, 40, 100710. [Google Scholar] [CrossRef] [Scilit]
- Arena. Arena AI Leaderboards. Consulted Leaderboards for Text and Vision Models. Available online: https://arena.ai/leaderboard (accessed on 4 March 2026).
- Google DeepMind. Gemini 3.1 Pro Model Card. 2026. Available online: https://deepmind.google/models/model-cards/gemini-3-1-pro/ (accessed on 11 March 2026).
- Artificial Analysis. Gemini 3.1 Pro Preview: The New Leader in AI. 2026. Available online: https://artificialanalysis.ai/articles/gemini-3-1-pro-preview-new-leader-in-ai (accessed on 11 March 2026).
- Moonshot AI. Kimi-K2.5. Hugging Face Model Card. 2026. Available online: https://huggingface.co/moonshotai/Kimi-K2.5 (accessed on 17 March 2026).
- Anthropic. Introducing Claude Opus 4.6. Product Announcement. 2026. Available online: https://www.anthropic.com/news/claude-opus-4-6 (accessed on 11 March 2026).
- Qwen Team. Qwen3.5-27B. Hugging Face Model Card. 2026. Available online: https://huggingface.co/Qwen/Qwen3.5-27B (accessed on 11 March 2026).
- Cheng, S.; Giabbanelli, P.J.; Kuang, Z. Identifying the Building Blocks of Social Simulation Models: A Qualitative Analysis using Open-Source Codes in NetLogo. In Proceedings of the 2023 Annual Modeling and Simulation Conference (ANNSIM); IEEE: New York, NY, USA, 2023; pp. 306–317. [Google Scholar]
- Freund, A.J.; Giabbanelli, P.J. Are we modeling the evidence or our own biases? A comparison of conceptual models created from reports. In Proceedings of the 2021 Annual Modeling and Simulation Conference (ANNSIM); IEEE: New York, NY, USA, 2021; pp. 1–12. [Google Scholar]











| Feature | Purpose | Method | Library | Output |
|---|---|---|---|---|
| Descriptive statistics | Summarize global properties of the signal | Direct statistical computation | NumPy/Python | Mean, median, variance, standard deviation, min/max values and indices, start/end value, counts of valid/dropped samples |
| Peaks (local maxima) | Identify salient positive events | Peak detection using neighbor comparison and prominence estimation | scipy.signal | Indices and values of peaks, with prominence |
| Valleys (local minima) | Identify salient negative events | Peak detection applied to the negated signal | scipy.signal | Indices and values of valleys, with prominence |
| Inflection points | Detect changes in curvature of the signal | Savitzky–Golay filtering and second-derivative sign analysis | scipy.signal | Indices where the second derivative changes sign |
| Local trends | Determine monotonic behavior in short segments | Rolling Mann–Kendall trend test with Sen’s slope estimator | pymannkendall | Start/end, Trend label, Kendall’s , Sen’s slope, p-value, significance flag |
| Change points | Detect structural regime shifts in the signal | PELT change-point detection algorithm | ruptures | Indices marking boundaries between statistical regimes |
| Oscillatory regimes | Identify periodic behavior and dominant cycles | Continuous Wavelet Transform (Morlet wavelet) | pywt (PyWavelets) | Dominant oscillation period and indices where oscillation occurs |
| Part | Use | Python Library |
|---|---|---|
| Data Extraction | Run simulation; export repeated runs | pynetlogo (0.5.2), pandas (2.2.3) |
| Extract parameters, docs, and code from NetLogo files | re, json, xml.etree.ElementTree (3.12.2) | |
| Report Generation | Plot repeated-simulation data | pandas (2.2.3), matplotlib (3.10.5) |
| Build statistical table evidence | scipy (1.15.2), PyWavelets (1.8.0) | |
| Generate explanations | openai (2.15.0) | |
| Convert images to base64 | base64 (3.12.2) | |
| Combine prompt features and sweep settings | itertools (3.12.2) | |
| Summarize source text | transformers (4.49.0), torch (2.10.0), sentencepiece (0.2.1), bert-extractive-summarizer (0.10.1) | |
| Evaluation | Text/reference scoring | nltk (3.8.1), rouge_score (0.1.2), textstat (0.7.7) |
| DoE/ANOVA analysis | pandas (2.2.3), scipy (1.15.2), statsmodels (0.14.2) |
| Reference Family | Variable/Metric | BLEU | METEOR | R-1 | R-2 | R-L | Reading Ease |
|---|---|---|---|---|---|---|---|
| author | Agent-based model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| author | Summarization algorithm | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| author | Simulation evidence | 0.05 | 0.71 | <0.01 | <0.01 | <0.01 | <0.01 |
| author | LLM | <0.01 | <0.01 | 0.02 | <0.01 | <0.01 | <0.01 |
| author | Use of roles | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | 0.74 |
| author | Generating insights | 0.68 | 0.48 | 0.02 | <0.01 | <0.01 | <0.01 |
| author | Providing examples | 0.43 | 0.05 | 0.11 | 0.18 | 0.65 | <0.01 |
| long LLM | Agent-based model | <0.01 | 0.74 | <0.01 | <0.01 | <0.01 | <0.01 |
| long LLM | Summarization algorithm | N/A | N/A | N/A | N/A | N/A | N/A |
| long LLM | Simulation evidence | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| long LLM | LLM | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| long LLM | Use of roles | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | 0.94 |
| long LLM | Generating insights | 0.01 | 0.07 | <0.01 | <0.01 | <0.01 | <0.01 |
| long LLM | Providing examples | 0.84 | <0.01 | 0.27 | 0.56 | 0.80 | <0.01 |
| short LLM | Agent-based model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| short LLM | Summarization algorithm | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| short LLM | Simulation evidence | 0.84 | 0.64 | <0.01 | <0.01 | <0.01 | <0.01 |
| short LLM | LLM | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| short LLM | Use of roles | <0.01 | 0.22 | <0.01 | <0.01 | <0.01 | 0.74 |
| short LLM | Generating insights | 0.58 | 0.05 | 0.80 | <0.01 | <0.01 | <0.01 |
| short LLM | Providing examples | 0.72 | 0.06 | 0.02 | 0.51 | 0.15 | <0.01 |
| modeler | Agent-based model | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| modeler | Summarization algorithm | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 | <0.01 |
| modeler | Simulation evidence | 0.08 | 0.30 | 0.16 | <0.01 | <0.01 | <0.01 |
| modeler | LLM | <0.01 | <0.01 | <0.01 | 0.52 | 0.08 | <0.01 |
| modeler | Use of roles | <0.01 | 0.77 | 0.02 | <0.01 | <0.01 | 0.74 |
| modeler | Generating insights | 0.91 | 0.43 | 0.13 | <0.01 | <0.01 | <0.01 |
| modeler | Providing examples | 0.90 | 0.20 | 0.28 | 0.55 | 0.90 | <0.01 |
| Factor | Retained Level (s) | Discarded Level (s) | Rationale |
|---|---|---|---|
| Evidence | plot | Table, plot + table | Dominates lexical metrics across runs; most consistent winner |
| Example | True | False | Systematically improves BLEU and ROUGE metrics |
| Insights | False | True | Turning insights off improves most lexical metrics and readability |
| Role | False | True | Role instructions reduce performance in most comparisons |
| Summarizer | BERT, BART, T5 | longformer_ext | BERT wins overlap metrics; T5 wins readability; BART wins R-L |
| LLM | ABM | Prompt | Summarizer | #errors | Notes |
|---|---|---|---|---|---|
| Claude | Fauna | Plot, Examples | BERT | 5 | Lacks transitions between trend descriptions. Abrupt transitions from describing one graph to another, which confuses the reader regarding which metric is described. Overestimates some values or ticks. The report describes variable per variable rather than chronology. The report is overall very accurate but omits details. |
| Grazing | BART | 1 | Confuses some values set to certain metrics. There is more consistency in the way metrics are introduced or the way transitions are made. Sentences make sense and the reader does not become confused regarding which metric is being described. It is better self-contained even though it misses some information. A bit too concise. | ||
| Gemini | Milk | T5 | 2 | Very short report, missing a lot of information. Misses some numbers and lacks clarity on the exact metrics described. Poor transitions. | |
| Fauna | BERT | 5 | Many numbers missing. Many parameter values wrongly reported. Overestimates some values. | ||
| Qwen | Grazing | Plot + Table | BART | 0 | Very short and not informative enough. Omits a vast amount of information. Its formulations and trend descriptions do not fully support understanding. |
| Milk | Table, Insights, role | LongFormer | 1 | Juxtaposition of sentences which makes it really hard to understand and contextualize which metric is being described. Poor coherence and narrative. The report is very short and not very informative. | |
| Kimi | Fauna | Plot + Table, Insights, role | BERT | 7 | Some numbers missing. Hallucinated the number of simulations. Describes the trend of a metric without ever mentioning that metric, thus confusing the reader in thinking it still talks about the previous metric. Abrupt transitions. |
| Grazing | Table, Examples, insights, role | T5 | 4 | Imprecise on the specific metrics described. Does not present the parameters. Omits information. Describes the trend of a metric without ever mentioning that metric, thus confusing the reader by thinking it is still talks about the previous metric. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Flandre, N.Y.; Giabbanelli, P.J. Distilling the Complexity of Agent-Based Simulations into Textual Explanations via Large Language Models. Big Data Cogn. Comput. 2026, 10, 121. https://doi.org/10.3390/bdcc10040121
Flandre NY, Giabbanelli PJ. Distilling the Complexity of Agent-Based Simulations into Textual Explanations via Large Language Models. Big Data and Cognitive Computing. 2026; 10(4):121. https://doi.org/10.3390/bdcc10040121
Chicago/Turabian StyleFlandre, Noé Y., and Philippe J. Giabbanelli. 2026. "Distilling the Complexity of Agent-Based Simulations into Textual Explanations via Large Language Models" Big Data and Cognitive Computing 10, no. 4: 121. https://doi.org/10.3390/bdcc10040121
APA StyleFlandre, N. Y., & Giabbanelli, P. J. (2026). Distilling the Complexity of Agent-Based Simulations into Textual Explanations via Large Language Models. Big Data and Cognitive Computing, 10(4), 121. https://doi.org/10.3390/bdcc10040121

