How the Choice of LLM and Prompt Engineering Affects Chatbot Effectiveness
Abstract
1. Introduction
- A comprehensive analysis of the performance of various large language models (LLMs) on the Rasa Pro platform, highlighting the effectiveness of smaller models such as Gemini-1.5-Flash-8B and Gemma2-9B-IT.
- Demonstration of the significant impact of prompt engineering techniques, such as using structured formats like YAML and JSON, on the accuracy and efficiency of chatbot responses.
- Presentation of practical insights for chatbot designers, emphasizing the importance of model selection and prompt construction in optimizing chatbot performance.
2. Related Work
3. Research Methodology
3.1. Quantitative Methods
3.2. Qualitative Methods
- The use of smaller language models (LLMs) can lead to achieving comparable results in terms of accuracy for user intent recognition and the selection of appropriate conversation flows compared to larger models [11].
- Transforming the bot description within the prompt, such as providing information about the chatbot’s purpose, capabilities, and intended use cases, from plain text to a structured format (Markdown, YAML, JSON) can contribute to increasing the accuracy of responses generated by LLM models [19].
- Specifying expected outcomes within the prompt precisely should translate into greater accuracy of responses generated by LLM models [21].
3.3. Inter-Rater Agreement
- Availability and scalability: the chatbot system based on the Rasa Pro platform and the tested language models were configured to operate in a cloud environment, allowing for the easy scaling of computational resources and availability for other researchers.
- Openness and reproducibility: all system components were selected to be available for free, at least in a limited version.
- Working name: adopted in this study for easier identification.
- Cloud platform: where the model is available.
- Full model name: according to the provider’s naming convention.
- Number of parameters: characterizing the size and complexity of the model.
- Reference to the literature: allowing for a detailed review of the model description.
- Actions StartFlow(flow_name): If the given flow_name corresponded to an existing flow, the accuracy was 1; otherwise, it was 0.
- Actions Clarify(flow_name[1..n]): If at least one of the given flow_name was correct, the accuracy was calculated as the inverse of the number of provided flow names.
- The LLM response correctly identified the flow name, resulting in an accuracy of 100.00%.
- The LLM response did not match the correct flow name, resulting in an accuracy of 0.00%.
- The LLM response included the correct flow name purchase_phone_number among the provided options. Since one out of two provided flow names was correct, the accuracy was calculated as one half, resulting in an accuracy of 50.00%.
- The LLM response included the correct flow name payment_inquiry among the provided options. Since one out of three provided flow names was correct, the accuracy was calculated as one-third, resulting in an accuracy of 33.33%.
- The LLM response did not include the correct flow name, resulting in an accuracy of 0.00%.
4. Results and Analysis
5. Conclusions
Funding
Data Availability Statement
Conflicts of Interest
References
- Benaddi, L.; Ouaddi, C.; Khriss, I.; Ouchao, B. Analysis of Tools for the Development of Conversational Agents. Comput. Sci. Math. Forum 2023, 6, 5. [Google Scholar] [CrossRef] [Scilit]
- Dagkoulis, I.; Moussiades, L. A Comparative Evaluation of Chatbot Development Platforms. In Proceedings of the 26th Pan-Hellenic Conference on Informatics, Athens, Greece, 25–27 November 2022. [Google Scholar] [CrossRef] [Scilit]
- Introduction to Rasa Pro. 2025. Available online: https://rasa.com/docs/rasa-pro/ (accessed on 22 January 2025).
- Costa, L.A.L.F.d.; Melchiades, M.B.; Girelli, V.S.; Colombelli, F.; Araujo, D.A.d.; Rigo, S.J.; Ramos, G.d.O.; Costa, C.A.d.; Righi, R.d.R.; Barbosa, J.L.V. Advancing Chatbot Conversations: A Review of Knowledge Update Approaches. J. Braz. Comput. Soc. 2024, 30, 55–68. [Google Scholar] [CrossRef] [Scilit]
- Tamrakar, R.; Wani, N. Design and Development of CHATBOT: A Review. In Proceedings of the International Conference on “Latest Trends in Civil, Mechanical and Electrical Engineering”, Online, 12–13 April 2021. [Google Scholar]
- Brabra, H.; Baez, M.; Benatallah, B.; Gaaloul, W.; Bouguelia, S.; Zamanirad, S. Dialogue Management in Conversational Systems: A Review of Approaches, Challenges, and Opportunities. IEEE Trans. Cogn. Dev. Syst. 2022, 14, 783–798. [Google Scholar] [CrossRef] [Scilit]
- Matic, R.; Kabiljo, M.; Zivkovic, M.; Cabarkapa, M. Extensible Chatbot Architecture Using Metamodels of Natural Language Understanding. Electronics 2021, 10, 2300. [Google Scholar] [CrossRef] [Scilit]
- Sanchez Cuadrado, J.; Perez-Soler, S.; Guerra, E.; De Lara, J. Automating the Development of Task-oriented LLM-based Chatbots. In Proceedings of the 6th ACM Conference on Conversational User Interfaces, Luxembourg, 8–10 July 2024; CUI ’24. Association for Computing Machinery: New York, NY, USA, 2024; pp. 1–10. [Google Scholar] [CrossRef] [Scilit]
- Marvin, G.; Hellen, N.; Jjingo, D.; Nakatumba-Nabende, J. Prompt Engineering in Large Language Models. In Proceedings of the Data Intelligence and Cognitive Informatics, Tirunelveli, India, 27–28 June 2023; Jacob, I.J., Piramuthu, S., Falkowski-Gilski, P., Eds.; IEEE: Piscataway, NJ, USA, 2024; pp. 387–402. [Google Scholar] [CrossRef] [Scilit]
- Benram, G. Understanding the Cost of Large Language Models (LLMs). 2024. Available online: https://www.tensorops.ai/post/understanding-the-cost-of-large-language-models-llms (accessed on 20 January 2025).
- Bocklisch, T.; Werkmeister, T.; Varshneya, D.; Nichol, A. Task-Oriented Dialogue with In-Context Learning. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Nadeau, D.; Kroutikov, M.; McNeil, K.; Baribeau, S. Benchmarking Llama2, Mistral, Gemma and GPT for Factuality, Toxicity, Bias and Propensity for Hallucinations. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Yi, J.; Ye, R.; Chen, Q.; Zhu, B.; Chen, S.; Lian, D.; Sun, G.; Xie, X.; Wu, F. On the Vulnerability of Safety Alignment in Open-Access LLMs. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 11–16 August 2024; Ku, L.W., Martins, A., Srikumar, V., Eds.; IEEE: Piscataway, NJ, USA, 2024; pp. 9236–9260. [Google Scholar] [CrossRef] [Scilit]
- Zhao, S.; Tuan, L.A.; Fu, J.; Wen, J.; Luo, W. Exploring Clean Label Backdoor Attacks and Defense in Language Models. IEEE/ACM Trans. Audio Speech Lang. Process. 2024, 32, 3014–3024. [Google Scholar] [CrossRef] [Scilit]
- Niu, Z.; Ren, H.; Gao, X.; Hua, G.; Jin, R. Jailbreaking Attack against Multimodal Large Language Model. arXiv 2024, arXiv:2402.02309. [Google Scholar]
- Amujo, O.E.; Yang, S.J. Evaluating the Efficacy of Foundational Models: Advancing Benchmarking Practices to Enhance Fine-Tuning Decision-Making. arXiv 2024, arXiv:2407.11006. [Google Scholar]
- Gemma Team. Gemma: Open Models Based on Gemini Research and Technology. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.A.; Lacroix, T.; Roziere, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. LLaMA: Open and Efficient Foundation Language Models. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- He, J.; Rungta, M.; Koleczek, D.; Sekhon, A.; Wang, F.X.; Hasan, S. Does Prompt Formatting Have Any Impact on LLM Performance? arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Arora, G.; Jain, S.; Merugu, S. Intent Detection in the Age of LLMs. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Cao, B.; Cai, D.; Zhang, Z.; Zou, Y.; Lam, W. On the Worst Prompt Performance of Large Language Models. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Płaza, M.; Pawlik, Ł.; Deniziak, S. Call Transcription Methodology for Contact Center Systems. IEEE Access 2021, 9, 110975–110988. [Google Scholar] [CrossRef] [Scilit]
- Pawlik, L.; Plaza, M.; Deniziak, S.; Boksa, E. A method for improving bot effectiveness by recognising implicit customer intent in contact centre conversations. Speech Commun. 2022, 143, 33–45. [Google Scholar] [CrossRef] [Scilit]
- Rasa Inspector. 2025. Available online: https://rasa.com/docs/rasa-pro/production/inspect-assistant/ (accessed on 21 January 2025).
- Codespaces Documentation. Available online: https://docs.github.com/en/codespaces (accessed on 21 January 2025).
- Gemini API. Available online: https://ai.google.dev/gemini-api/docs (accessed on 21 January 2025).
- GroqCloud. Available online: https://groq.com/groqcloud/ (accessed on 21 January 2025).
- llama-models/models/llama3_2/MODEL_CARD.md at main · meta-llama/llama-models. Available online: https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md (accessed on 22 January 2025).
- llama-models/models/llama3_1/MODEL_CARD.md at main · meta-llama/llama-models. Available online: https://github.com/meta-llama/llama-models/blob/main/models/llama3_1/MODEL_CARD.md (accessed on 22 January 2025).
- Gemini Team Google. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- google/gemma-2-9b-it · Hugging Face. 2024. Available online: https://huggingface.co/google/gemma-2-9b-it (accessed on 22 January 2025).
- Gemini 2.0 Flash (Experimental)|Gemini API. Available online: https://ai.google.dev/gemini-api/docs/models/gemini-v2 (accessed on 22 January 2025).
- Banerjee, D.; Singh, P.; Avadhanam, A.; Srivastava, S. Benchmarking LLM powered Chatbots: Methods and Metrics. arXiv 2023. [Google Scholar] [CrossRef] [Scilit]
- Kossack, P.; Unger, H. Emotion-Aware Chatbots: Understanding, Reacting and Adapting to Human Emotions in Text Conversations. In Proceedings of the Advances in Real-Time and Autonomous Systems; Unger, H., Schaible, M., Eds.; Springer: Cham, Switzerland, 2024; pp. 158–175. [Google Scholar]
- Vishal, M.; Vishalakshi Prabhu, H. A Comprehensive Review of Conversational AI-Based Chatbots: Types, Applications, and Future Trends. In Internet of Things (IoT): Key Digital Trends Shaping the Future; Misra, R., Rajarajan, M., Veeravalli, B., Kesswani, N., Patel, A., Eds.; Springer: Singapore, 2023; pp. 293–303. [Google Scholar]
- Dam, S.K.; Hong, C.S.; Qiao, Y.; Zhang, C. A Complete Survey on LLM-based AI Chatbots. arXiv 2024, arXiv:2406.16937. [Google Scholar]
- Wake, N.; Kanehira, A.; Sasabuchi, K.; Takamatsu, J.; Ikeuchi, K. Bias in Emotion Recognition with ChatGPT. arXiv 2023, arXiv:2310.11753. [Google Scholar]




| Flow | Example Phrase |
|---|---|
| Add additional phone number | How can I add another phone to my plan? |
| Cancel device insurance | How do I cancel the insurance on my phone? |
| Payment inquiry | Can you explain the recent charges on my account? |
| Purchase internet service | What internet plans do you offer? |
| Purchase phone number | What’s the process for getting a new phone number with you? |
| Report service issues | My internet is acting up. Can you fix it? |
| Request device insurance | I’m interested in getting insurance for my new phone. |
| Request device repair | I need help repairing my broken phone. |
| Top up account | How can I increase my data usage? |
| Transfer phone number | Can you help me move my phone number to you? |
| Name | Platform | Model | Parameters | Ref. |
|---|---|---|---|---|
| Llama-1B | Groq Cloud | llama-3.2-1b-preview | 1.23 billion | [28] |
| Llama-3B | Groq Cloud | llama-3.2-3b-preview | 3.21 billion | [28] |
| Llama-8B | Groq Cloud | llama-3.1-8b-instant | 8 billion | [29] |
| Gemini1.5F-8B | Google Cloud | gemini-1.5-flash-8b | 8 billion | [30] |
| Gemma-9B | Groq Cloud | gemma2-9b-it | 9.24 billion | [31] |
| Llama-70B | Groq Cloud | llama-3.1-70b-versatile | 70 billion | [29] |
| Gemini1.5F | Google Cloud | gemini-1.5-flash | Very large 1 | [30] |
| Gemini2.0F | Google Cloud | gemini-2.0-flash-exp | Very large 1 | [32] |
| Prompt Template | Provider | Model | Configuration |
|---|---|---|---|
| Default | Groq | Llama-3B | - name: SingleStepLLMCommandGenerator llm: provider: groq model: llama-3.2-1b-preview |
| YAML | Groq | Gemma-9B | - name: SingleStepLLMCommandGenerator prompt_template: data/yaml.jinja2 llm: provider: groq model: gemma2-9b-it |
| JSON | Gemini | Gemini1.5F | - name: SingleStepLLMCommandGenerator prompt_template: data/json.jinja2 llm: provider: gemini model: gemini-1.5-flash |
| Prompt Section | Prompt Content |
|---|---|
| Flows | These are the flows that can be started: purchase_phone_number: Assist users in purchasing a new phone number. cancel_device_insurance: Help users cancel their device insurance. |
| Phrase | The user just said "Please cancel my device insurance". |
| Actions | These are your available actions: * Starting another flow, described by "StartFlow(flow_name)" * Clarifying which flow should be started. An example would be Clarify(list_contacts, add_contact) |
| Command | Your action list: |
| Flow | LLM Response | Accuracy [%] |
|---|---|---|
| Purchase internet service | StarfFlow(purchase_internet_service) | 100.00 |
| Purchase phone number | StarfFlow(purchase_internet_service) | 0.00 |
| Purchase phone number | Clarify(purchase_internet_service, purchase_phone_number) | 50.00 |
| Payment inquiry | Clarify(payment_inquiry, top_up_account purchase_phone_number) | 33.33 |
| Transfer phone number | Clarify(purchase_internet_service, purchase_phone_number) | 0.00 |
| Format | Content |
|---|---|
| Plain text | These are the flows that can be started: purchase_phone_number: Assist users in purchasing a new phone number. These are your available actions: * Starting another flow, described by StartFlow(flow_name) |
| Markdown | ### Flows * **purchase_phone_number:** Assist users in purchasing a new phone number. ### Actions * **StartFlow(flow_name):** Starting another flow. |
| YAML | flows: - flow_name: "purchase_phone_number" flow_description: "Assist users in purchasing a new phone number." actions: - action_name: StartFlow(flow_name) action_description: Starting another flow. |
| JSON | "flows": [{"flow_name": "purchase_phone_number", "flow_description": "Assist users in purchasing a new phone number."}] "actions": [{"action_name": "StartFlow(flow_name)", "action_description": "Starting another flow"}] |
| LLM | Plain Text | Markdown | YAML | JSON | Best Format | Best Improv. |
|---|---|---|---|---|---|---|
| Llama-1B | 7.46 | 6.82 | 12.71 | 5.32 | YAML | 5.25 |
| Llama-3B | 32.33 | 38.68 | 40.83 | 46.37 | JSON | 14.04 |
| Llama-8B | 68.99 | 66.82 | 68.42 | 65.53 | Plain text | −0.57 |
| Gemini1.5F-8B | 64.65 | 81.52 | 81.40 | 86.27 | JSON | 21.62 |
| Gemma-9B | 69.11 | 75.37 | 76.08 | 77.86 | JSON | 6.26 |
| Llama-70B | 84.03 | 83.50 | 82.89 | 89.64 | JSON | 5.61 |
| Gemini1.5F | 90.25 | 88.63 | 91.15 | 88.74 | YAML | 0.90 |
| Gemini2.0F | 83.27 | 85.38 | 87.77 | 89.92 | YAML | 4.50 |
| LLM | Plain text | Markdown | YAML | JSON | ||||
|---|---|---|---|---|---|---|---|---|
| Precise | Improv. | Precise | Improv. | Precise | Improv. | Precise | Improv. | |
| Llama-1B | 10.50 | +3.04 | 13.50 | +6.68 | 7.98 | −4.73 | 11.15 | +5.83 |
| Llama-3B | 46.81 | +14.48 | 45.79 | +7.11 | 51.90 | +11.07 | 51.73 | +5.36 |
| Llama-8B | 69.16 | +0.17 | 73.17 | +6.35 | 71.59 | +3.17 | 69.16 | +3.63 |
| Gemini1.5F-8B | 67.34 | +2.69 | 85.20 | +3.68 | 85.40 | +4.00 | 88.10 | +1.83 |
| Gemma-9B | 67.90 | −1.21 | 79.77 | +4.40 | 83.28 | +7.20 | 77.07 | −0.79 |
| Llama-70B | 82.71 | −1.32 | 84.22 | +0.72 | 83.32 | +0.43 | 88.72 | −0.92 |
| Gemini1.5F | 88.49 | −1.76 | 88.74 | +0.11 | 89.37 | −1.78 | 90.65 | +1.91 |
| Gemini2.0F | 84.50 | +1.23 | 86.65 | +1.27 | 86.98 | −0.79 | 88.10 | −1.82 |
| LLM | Best Prompt | Plain Text/Concise | Best Prompt | Best Improv. |
|---|---|---|---|---|
| Llama-1B | Markdown/Precise | 7.46 | 13.50 | +6.04 |
| Llama-3B | YAML/Precise | 32.33 | 51.90 | +19.57 |
| Llama-8B | Markdown/Precise | 68.99 | 73.17 | +4.18 |
| Gemini1.5F-8B | JSON/Precise | 64.65 | 88.10 | +23.45 |
| Gemma-9B | YAML/Precise | 69.11 | 83.28 | +14.17 |
| Llama-70B | JSON/Concise | 84.03 | 89.64 | +5.61 |
| Gemini1.5F | YAML/Concise | 90.25 | 91.15 | +0.90 |
| Gemini2.0F | JSON/Concise | 83.27 | 89.92 | +6.65 |
| Finding | Description |
|---|---|
| Smaller models’ performance | Smaller models (e.g., Gemini1.5F-8B, Gemma-9B) can achieve results close to larger models with optimized prompts. |
| Impact of prompt format | Structured formats (YAML, JSON) significantly improve response accuracy for many models. |
| Command precision | Precise commands improve response quality for smaller models but have less impact on larger models. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Pawlik, L. How the Choice of LLM and Prompt Engineering Affects Chatbot Effectiveness. Electronics 2025, 14, 888. https://doi.org/10.3390/electronics14050888
Pawlik L. How the Choice of LLM and Prompt Engineering Affects Chatbot Effectiveness. Electronics. 2025; 14(5):888. https://doi.org/10.3390/electronics14050888
Chicago/Turabian StylePawlik, Lukasz. 2025. "How the Choice of LLM and Prompt Engineering Affects Chatbot Effectiveness" Electronics 14, no. 5: 888. https://doi.org/10.3390/electronics14050888
APA StylePawlik, L. (2025). How the Choice of LLM and Prompt Engineering Affects Chatbot Effectiveness. Electronics, 14(5), 888. https://doi.org/10.3390/electronics14050888

