Better Prompts, Better Usefulness: A Systematic Review and Experimental Evaluation of Structured Prompting Techniques in Large Language Models
Abstract
1. Introduction
2. Related Work
2.1. Generative AI and Large Language Models (LLMs)
2.2. GenAI in Business Contexts
2.3. Prompt Engineering
Prompt Enhancement Techniques
3. Materials and Methods
3.1. Systematic Literature Review (SLR)
3.1.1. Search Strategy
3.1.2. Inclusion/Exclusion Criteria
- Peer-reviewed journal articles indexed in Scopus and/or Web of Science;
- Studies published in English;
- Studies in which prompting techniques for large language models represent the primary focus of the investigation;
- Studies that explicitly conceptualize prompting as a methodological element, by defining, analyzing, evaluating, or comparing prompting approaches (e.g., zero-shot, few-shot, chain-of-thought, and their combinations).
- Non-peer-reviewed or non-academic publications;
- Studies in which prompting is only mentioned incidentally or used implicitly without methodological discussion;
- Studies focused exclusively on model architecture, training procedures or benchmark optimization without substantive analysis of prompting techniques.
3.1.3. Screening Method
3.1.4. Data Extraction
3.1.5. Quality Assessment
3.2. Experimental Design
3.2.1. Identification of Business Use Cases
3.2.2. Selection of Prompting Techniques
- Zero-shot prompting;
- One-shot prompting;
- Few-shot prompting;
- Chain-of-thought prompting;
- One-shot prompting combined with Chain-of-Thought;
- Few-shot prompting combined with Chain-of-Thought.
3.2.3. Prompt Construction and Output Generation
3.2.4. Data Collection and Sampling
3.2.5. Measurements
3.2.6. Conceptual Model
4. Results of Systematic Literature Review
4.1. Baseline Prompting Techniques
Zero-Shot Prompting
4.2. Task Alignment Prompting Techniques
4.2.1. One-Shot Prompting
4.2.2. Few-Shot Prompting
4.2.3. Role Prompting
4.3. Output Transparency Prompting Techniques
4.3.1. Chain-of-Thought (CoT) Prompting
4.3.2. Self-Consistency Prompting
4.3.3. Chain-of-Verification (CoV) Prompting
5. Results of Empirical Evaluation
5.1. Sample Characteristics
5.2. Measurements Reliability and Data Preparation
5.3. Descriptive Statistics of Key Variables
5.4. Hypothesis Testing
5.4.1. Effects of Prompting Techniques on Perceived Output Usefulness (H1)
5.4.2. Estimated Marginal Means
5.4.3. Pairwise Comparisons
5.5. Moderating Role of Generative AI Usage (H2)
Interaction Model: Testing Moderation by Usage Frequency
5.6. Summary of Hypothesis Testing
6. Discussion
7. Managerial Contributions
8. Theoretical Implications
9. Limitations and Future Research
10. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| ML | Machine Learning |
| DGM | Deep Generative Model |
| GenAI | Generative Artificial Intelligence |
| LLM | Large Language Model |
| PRISMA | Preferred Reporting Items for Systematic Reviews and Meta-Analyses |
| SLR | Systematic Literature Review |
| KPI | Key Performance Indicator |
| CoT | Chain-of-Thought |
| CoV | Chain-of-Verification |
| IV | Independent Variable |
| DV | Dependent Variable |
| LMEM | Linear Mixed-Effects Model |
| EMM | Estimated Marginal Means |
| PECS | Prompt Engineering Competence Scale |
| OS | One-Shot Prompting |
| FS | Few-Shot Prompting |
Appendix A. Prompting Configurations Used in the Experimental Study
Appendix A.1. Analysis Task
| Technique | Prompt | Output |
|---|---|---|
| Zero-shot | Explain the following marketing KPIs in a clear and professional way and classify them into meaningful categories for a quarterly performance review: Churn Rate, Conversion Rate, Customer Acquisition Cost (CAC), Customer Lifetime Value (CLV), Return on Investment (ROI), Click-Through Rate (CTR). | Churn Rate measures how many customers stop using a product or service. Conversion Rate shows the percentage of users who complete a desired action. CAC indicates how much it costs to acquire a new customer. CLV estimates the total revenue a customer is expected to generate over time. ROI compares the financial return of an initiative to its cost. CTR measures how many users click on an ad relative to the number of impressions. Categories:
|
| One-shot | Explain the following marketing KPIs and classify them into meaningful categories for a quarterly performance review. Example of expected style: “This metric helps leadership understand whether our current initiatives are creating meaningful business impact and where strategic adjustments may be required.” KPIs: Churn Rate, Conversion Rate, Customer Acquisition Cost (CAC), Customer Lifetime Value (CLV), Return on Investment (ROI), Click-Through Rate (CTR). | Churn Rate indicates how many customers discontinue their relationship with the company, highlighting potential retention issues. Conversion Rate reflects how effectively users move from interest to action, informing where strategic adjustments may be needed. CAC shows the investment required to acquire each new customer, helping evaluate whether acquisition efforts are sustainable. CLV estimates the long-term value of each customer, supporting decisions on where to prioritize resources. ROI assesses whether an initiative generates meaningful business impact relative to its cost. CTR reveals how effectively our ads capture user interest. Categories:
|
| Few-shot | Explain the following marketing KPIs and classify them into meaningful categories for a quarterly performance review. Follow the examples below: # Example 1: “This metric highlights whether our current strategy is generating sustainable growth and indicates where resource allocation may need to shift in the next quarter.” # Example 2 “This metric shows how efficiently our campaigns convert attention into action, helping identify which channels require optimisation to improve performance.” # Example 3 “A sudden change in this metric often signals emerging risks in customer behaviour and prompts a deeper review of retention and engagement patterns.” KPIs: Churn Rate, Conversion Rate, Customer Acquisition Cost (CAC), Customer Lifetime Value (CLV), Return on Investment (ROI), Click-Through Rate (CTR). | Churn Rate highlights potential risks in customer retention and signals when engagement patterns may require deeper investigation. Conversion Rate shows how efficiently we convert interest into action, helping identify which parts of the funnel need optimisation. CAC indicates whether our acquisition strategy is sustainable and where spending may need to be adjusted. CLV provides insight into long-term revenue potential and helps determine where to focus investment. ROI shows whether our initiatives are generating sufficient impact to justify continued resource allocation. CTR reflects how effectively our messaging captures attention and whether creative adjustments are needed. Categories:
|
| Chain-of-Thought | Explain the following marketing KPIs and classify them into meaningful categories for a quarterly performance review. Before answering, think through the steps: what each KPI measures, which part of the funnel it belongs to and why it matters for quarterly evaluation. Provide only the final explanation. KPIs: Churn Rate, Conversion Rate, Customer Acquisition Cost (CAC), Customer Lifetime Value (CLV), Return on Investment (ROI), Click-Through Rate (CTR). | Churn Rate measures the percentage of customers who discontinue their relationship with the company and helps assess retention performance. Conversion Rate shows how effectively users complete a desired action, indicating funnel efficiency. CAC reflects the average cost of acquiring a new customer and is essential for evaluating acquisition scalability. CLV estimates the long-term value of each customer, supporting investment decisions. ROI evaluates whether an initiative generates more value than it costs. CTR measures how many users engage with an ad relative to impressions, indicating the relevance of campaign messaging. Categories:
|
| One-shot + Chain-of-Thought | Explain the following KPIs and classify them into meaningful categories for a quarterly performance review. Example of expected style: “This metric helps leadership understand whether our current initiatives are creating meaningful business impact and where strategic adjustments may be required.” Think step by step about what each KPI measures and where it fits in the marketing funnel, but provide only the final explanation. KPIs: Churn Rate, Conversion Rate, Customer Acquisition Cost (CAC), Customer Lifetime Value (CLV), Return on Investment (ROI), Click-Through Rate (CTR). | Churn Rate shows the percentage of customers who stop using our service and helps identify retention challenges. Conversion Rate measures how effectively users complete a targeted action, indicating funnel performance. CAC reflects the average cost of acquiring a new customer and helps assess acquisition efficiency. CLV estimates the long-term value of each customer and supports strategic investment decisions. ROI compares the net value created by an initiative to its cost. CTR measures how many users click on an ad relative to impressions, indicating the relevance of our messaging. Categories:
|
| Few-shot + Chain-of-Thought | Explain the following KPIs and classify them into meaningful categories for a quarterly performance review. Follow the examples below. # Example 1: “This metric highlights whether our current strategy is generating sustainable growth and indicates where resource allocation may need to shift in the next quarter.” # Example 2: “This metric shows how efficiently our campaigns convert attention into action, helping identify which channels require optimisation to improve performance.” # Example 3: “A sudden change in this metric often signals emerging risks in customer behaviour and prompts a deeper review of retention and engagement patterns.” Think step by step about measurement, funnel stage, and strategic relevance, but provide only the final explanation. KPIs: Churn Rate, Conversion Rate, Customer Acquisition Cost (CAC), Customer Lifetime Value (CLV), Return on Investment (ROI), Click-Through Rate (CTR). | Churn Rate indicates how many customers we lose within a given period and highlights emerging retention risks. Conversion Rate shows how effectively we turn interest into action, helping identify which parts of the funnel require optimisation. CAC measures the total cost of acquiring each new customer and is essential for evaluating acquisition sustainability. CLV estimates the long-term revenue potential of each customer and supports decisions on where to focus investment. ROI evaluates whether an initiative generates sufficient financial return relative to its cost. CTR reflects how engaging and relevant our ads are to the target audience. Categories:
|
Appendix A.2. Synthesis Task
| Technique | Prompt | Output |
|---|---|---|
| Zero-shot | Please summarise the source text. | The EU has increased investments in the energy transition, but progress remains too slow to meet 2030 climate goals. Member States face delays in permitting, infrastructure limitations, and challenges in attracting private capital. Rising electricity demand also requires major grid expansion and digital integration. Experts argue that coordinated strategies, targeted incentives, and stable regulations are needed to support innovation and competitiveness. |
| One-shot | Summarize the source text in a clear and concise way. Example of desired style: “Provide a short summary that highlights the main goals, challenges and implications.” | Goals: Accelerate renewable energy deployment and reduce fossil-fuel dependence. Challenges: Slow permitting, infrastructure gaps, limited private investment, and rising electricity demand. Implications: Without coordinated strategies, incentives, and regulatory stability, the EU risks missing 2030 climate targets and weakening its economic competitiveness. |
| Few-shot | Summarise source text following the style shown in the examples. # Example 1: “Provide a structured summary that identifies the main objectives and the key obstacles.” # Example 2: “Include a brief explanation of why the topic is relevant for policy or strategic decision-making | The EU is increasing investments in renewable energy, yet current efforts fall short of what is needed to meet 2030 climate objectives. Key obstacles include slow permitting procedures, insufficient infrastructure, and difficulties in mobilizing private capital. Growing electricity demand further intensifies the need for grid expansion and digital integration. For policymakers, the findings highlight the importance of coordinated investment strategies, targeted incentives, and a stable regulatory environment to sustain innovation and maintain competitiveness. |
| Chain-of-Thought | First identify the key ideas in the text. Then summarise the source text in a clear and concise way based on those ideas. | Key ideas:
Although the EU is investing more in renewable energy, progress is too slow to meet 2030 climate targets. Structural barriers such as slow permitting, inadequate infrastructure, and limited private investment hinder advancement. With electricity demand rising, experts call for coordinated strategies, targeted incentives, and stable regulations to support innovation and competitiveness. |
| One-shot + Chain-of-Thought | Summarise the source text in a clear and concise way. Example of desired style: “Provide a short summary that highlights the main goals, challenges, and implications.”. Before summarising, identify the key ideas in the text | Key ideas:
|
| Few-shot + Chain-of-Thought | Summarise the source text following the style shown in the examples. # Example 1: “Provide a structured summary that identifies the main objectives and the key obstacles.” # Example 2: “Include a brief explanation of why the topic is relevant for policy or strategic decision-making.” Before summarising, identify the key ideas in the text. | Key ideas:
|
Appendix B. Data Extraction Tables (SLR)
Appendix B.1. Baseline Prompting Techniques
| Authors | Technique | Task/Domain | Evidence Type | Reported Effects |
|---|---|---|---|---|
| Ghimire, P. et al. [30] | Zero-shot | Cost estimation in construction | Human evaluation | Incomplete responses with an average confidence score of 1.91 (64%). Lack of predefined priority logic, reducing reliability and indicating necessity of structured prompting strategies |
| Bharti, P.K et al. [31] | Zero-shot | Peer reviews assessment | Human evaluation, recall, precision, f3 | The model’s classification was tentative and sentiment-biased. justification was partial and lacked confidence, leading to moderate recall and lower precision. |
| Tony et al. [12] | Zero-shot | Code generation/software security | Vulnerability count; security rule violations | Zero-shot prompting produces insecure code with a high number of vulnerabilities, offering limited control over security properties. |
| Razumovskaia et al. [37] | Zero-shot | Multilingual NLU (Intent Detection, NER, NLI) | F1, Accuracy | Zero-shot in-context learning shows very low performance across multilingual tasks, especially for low-resource languages, often close to baseline. |
| Indira Kumar et al. [36] | Zero-shot | Harmful meme text classification (cyberbullying, sarcasm, sentiment, emotion, harmfulness) | Accuracy, Precision, Recall, F1 | Zero-shot prompting enables multi-task classification without labeled data, but performance varies widely across tasks and models, especially for nuanced categories. It is efficient in handling multiple tasks simultaneously, but its reliability is influence by model and quality of output |
| Giannilias et al. [32] | Zero-shot | Cybersecurity text classification (dark web hacker forum posts) | Accuracy, Precision, Recall, F1 | Zero-shot prompting enables classification of unstructured hacker forum posts but shows uneven performance across categories and high variance between models. |
| Jaradat et al. [63] | Zero-shot | Traffic crash analysis (classification & IE) | Accuracy, F1, Jaccard | Zero-shot enables immediate inference but underperforms compared to few-shot and fine-tuning in complex multimodal tasks. |
| Zhang et al. [34] | Zero-shot | Early sepsis diagnosis from structured clinical data | Accuracy, Precision, Recall, F1 | Standard prompts yield acceptable recall but lower accuracy and F1, with limited interpretability and less structured diagnostic reasoning. |
| Yin & Guo [64] | Zero-shot | Financial forecasting (earnings direction prediction from numerical financial statements) | Accuracy, Precision, Recall, F1 | Direct-answer prompting yields stable but lower predictive performance and provides no transparent reasoning for financial decision-making. |
| Shafikuzzaman et al. [65] | Zero-shot | Software requirements classification (FR vs. NFR, NFR subclasses, security) | Precision, Recall, F1 (macro) | Zero-shot LLMs (GPT-4o) achieve strong performance without training data, sometimes approaching FS models. |
| Byra et al. [45] | Zero-shot | Medical image classification (chest X-rays: pneumonia vs. normal; breast US: benign vs. malignant) | Accuracy, AUC, ICC | Zero-shot classification using GPT-4-generated shape and texture descriptors achieves reasonable AUC (≈0.72–0.76) but shows high variability depending on descriptor quality. |
| Šuster et al. [44] | Zero-shot | Risk-of-bias assessment in clinical trials (RoB2, systematic reviews) | Macro F1, Accuracy | Zero-shot prompting fails to capture complex methodological reasoning required for RoB2 assessment, yielding F1 scores close to random or majority baselines. |
| Lee, A.V.Y. et al. [66] | Zero-shot | Education; commonsense reasoning; knowledge creation discourse | Human qualitative comparison (length, depth, idea quality) | Zero-shot prompting produces reasonable answers but does not sufficiently support idea elaboration or improvement in student discourse. |
| Zhen et al. [33] | Zero-shot | Traffic crash severity classification (road safety) | Macro F1, Macro Accuracy | Plain zero-shot prompting struggles with class imbalance and fails to reliably detect fatal crashes, especially for smaller models. |
| Shah et al. [42] | Zero-shot | Clinical QA, summarization, decision support | Not applicable (review) | Zero-shot prompting can be effective for simple clinical tasks but is prone to hallucinations and misapplication of knowledge in complex medical contexts. |
| Meshkin et al. [38] | Zero-shot | Regulatory NLP; PK drug–drug interaction (DDI) sentence classification | Precision, Recall, F1-score, Specificity | Several open-source LLMs (notably Flan-T5) achieve F1 ≈0.88–0.89 in zero-shot settings, matching or exceeding a BioBERT model trained on >20 k sentences. |
| Meshkin et al. [38] | Zero-shot | Regulatory NLP; identification of intrinsic factors affecting drug exposure | Precision (≈78.5%) | Zero-shot prompting enables large-scale extraction of clinically relevant intrinsic factors from >700 k FDA label sentences without fine-tuning. |
| Ono et al. [67] | Zero-shot | Histopathology image classification (tau lesions: AP, NP, TA) | Accuracy, Sensitivity, Specificity | Zero-shot GPT-4V accurately recognises staining and tissue type but fails to reliably identify specific tau lesions, achieving only ~40% accuracy. |
| Viswanathan et al. [39] | Zero-shot (instruction-only) | Text clustering (entity canonicalization, intent clustering, tweet clustering) | Macro F1, Micro F1, Pairwise F1, Accuracy, NMI | Even instruction-only prompts (no demonstrations) enable LLMs to inject task-specific structure into text representations, improving clustering quality. |
| Lee et al. [47] | Zero-shot | Automatic scoring of student-written science explanations | Accuracy, Precision, Recall, F1, QWK | Zero-shot prompting provides a feasible but limited baseline for automatic scoring, particularly struggling with minority proficiency categories. |
| Choi, J. [68] | Zero-shot | Relevance evaluation in information retrieval | Cohen’s kappa | Zero-shot prompts consistently outperform few-shot prompts, suggesting that GPT models leverage their pretrained knowledge more effectively without in-context examples. |
| Hebenstreit, K. et al. [69] | Zero-shot | Multi-domain multiple-choice question answering | Krippendorff’s alpha; accuracy | Direct prompting without explicit reasoning instructions consistently underperforms reasoning-based prompts across all evaluated models and datasets. |
| Miao, J. et al. [35] | Zero-shot | Medical QA and general reasoning | Accuracy (reported in prior literature), qualitative review | Zero-shot prompting enables broad task generalization but often lacks the specificity and accuracy required for nuanced clinical decision-making in nephrology. |
| Filienko, D. et al. [58] | Zero-shot | Problem-Solving Therapy dialogue (symptom assessment, goal setting) | Human Likert ratings; empathy metrics (ER, IP, EX) | Zero-shot prompting with explicit task instructions improves structure but fails to reliably capture implicit therapeutic skills required for PST, leading to lower overall quality than other techniques. |
| Thanasi-Boçe, M & Hoxa [62] | Zero-shot | entrepreneurship-learning, | Empirical (mixed-methods) | Zero-shot prompting is a technique where the prompt does not provide any prior information about the task that the LLM is supposed to perform. |
| Ayad, S; Alsayoud, F [70] | Zero-shot | Business process modeling | Empirical (experimental/applied) | Prompt without examples used to generate BPM elements from textual input |
| Carlson, NA; Burbano, V [48] | Zero-shot | Data annotation/text classification | Empirical (mixed-methods) | Direct task instruction without examples |
| Abou El Karam et al. [49] | Zero-shot | Change management/employee attitude assessment | Empirical (experimental) | Classification of open-ended employee responses without examples |
| Wang, J. [9] | Zero-shot | Business communication writing (negative message) | Empirical (experimental) | Direct prompt without additional guidance or interaction |
| Jovic, M et al. [71] | Zero-shot | Written feedback generation | Empirical (experimental) | Minimal instruction without contextual or iterative guidance |
| Jovic, M; et al. [71] | Zero-shot | Written feedback generation | Empirical (experimental) | Baseline AI feedback without structured context |
| Tocev, T.; Atanasovski, A. [72] | Zero-shot | IFRS advisory/financial reporting | Empirical (experimental) | Single-pass query without examples or stepwise decomposition |
| Tocev, T.; Atanasovski, A. [72] | Zero-shot | IFRS advisory/financial reporting | Empirical (experimental) | Direct prompt describing the case once |
Appendix B.2. Task Alignment Prompting Techniques
| Authors | Technique | Task/Domain | Evidence Type | Reported Effects |
|---|---|---|---|---|
| Golnari et al. [40] | One-shot | Medical | One-shot offers an improvement in respect to zero-shot | |
| Giannilias et al. [32] | One-shot | Cybersecurity text classification (dark web hacker forum posts) | Accuracy, Precision, Recall, F1 | Providing a single representative example per class improves classification accuracy by reducing ambiguity and clarifying label boundaries. |
| Agbareia et al. [41] | One-shot | Retinal disease classification from OCT images (AMD, DR, CSR, MH, Normal) | Accuracy (%), CI | Single-shot prompting allows basic multimodal diagnosis but shows limited accuracy, especially for complex retinal pathologies. |
| Shah et al. [42] | One-shot | Clinical decision support; diagnostic illustration | Not applicable (review) | Providing a single exemplar helps constrain model outputs and improves task understanding in narrowly defined clinical tasks. |
| Razumovskaia et al. [37] | Few-shot | Multilingual NLU (Intent Detection, NER, NLI) | F1, Accuracy | Few-shot prompting improves performance over zero-shot but remains significantly weaker than supervised approaches, particularly in cross-lingual settings. |
| Byra et al. [45] | Few-shot | Medical image classification (chest X-rays; breast US) | Accuracy, AUC | Few-shot descriptor selection (n ≥ 10) substantially improves accuracy (up to 0.81–0.83) by removing poorly performing text descriptors, without retraining the VLM. |
| Viswanathan et al. [39] | Few-shot | Text clustering | Macro F1, Micro F1, Pairwise F1, Accuracy, NMI | Few-shot prompting for key phrase expansion is the most effective LLM integration strategy, outperforming classical and neural baselines on most datasets. |
| Lee et al. [47] | Few-shot | Automatic scoring in education | Accuracy, Precision, Recall, F1, QWK | Few-shot prompting improves scoring accuracy by constraining output structure and aligning model behavior with human scoring patterns. |
| Jiang, Y. et al. [73] | Few-shot | Out-of-scope intent classification in task-oriented dialogue | AU-IOC, in-scope accuracy, OOS recall | Prompt-based learning that incorporates natural-language intent descriptions achieves higher AU-IOC scores across all datasets, especially in extremely low-data regimes (1–5 shots). |
| Qi, X. et al. [43] | Few-shot | Construction of Knowledge Graphs in the domain of equipment O&M | Precision, Recall, F1 | 10% increase in Recall, Few-shot learning helps LLMs better understand and adapt to task requirements, thereby substantially enhancing their generalization capabilities. |
| Bharti, P.K. et al. [31] | Few-shot | Peer reviews assessment | human evaluation, recall, precision, f2 | By leveraging example-driven analogical reasoning, this approach improved classification accuracy. It correctly identified the trivial review’s superficiality and the exhaustive review’s depth. However, the generated explanations were less detailed, covering partial aspects of the reviews. |
| Tony et al. [12] | Few-shot | Code generation/software security | Vulnerability count; security rule violations | Few-shot prompting reduces vulnerabilities compared to zero-shot, but effectiveness depends on vulnerability type and examples provided. |
| Golnari et al. [40] | Few-shot | Medical | - | The performance of the baseline prompt improves by 13.3% for detecting paroxysmal events using a two-shot approach as compared to zero-shot. |
| Du et al. [74] | Few-shot | power equipment defect grading | Accuracy | Few-shot examples improve performance over zero-shot but show limited gains in domain-specific reasoning. |
| Indira Kumar et al. [36] | Few-shot | Harmful meme text classification (cyberbullying, sarcasm, sentiment, emotion, harmfulness) | Accuracy, Precision, Recall, F1 | Few-shot prompting improves F1 and accuracy for several tasks when compared to zero-shot, particularly cyberbullying detection, though gains are uneven and sensitive to example choice. |
| Jaradat et al. [63] | Few-shot | Traffic crash analysis (severity, driver fault, actions) | Accuracy, F1, Jaccard | Few-shot prompting with GPT-4.5 achieves exceptional performance in classifications with accuracy reaching 98.9% and precision as high as 100%. |
| Shafikuzzaman et al. [65] | Few-shot | Software requirements classification (FR/NFR, NFR subclasses, security) | Precision, Recall, F1 (macro) | Few-shot prompting reliably improves classification performance across all tasks. |
| Agbareia et al. [41] | Few-shot | Retinal disease classification from OCT images | Accuracy (%), CI, p-values | Few-shot prompting with reference images substantially improves diagnostic accuracy across most retinal conditions, with gains up to +64% in some classes. |
| Šuster et al. [44] | Few-shot | Risk-of-bias assessment in clinical trials | Macro F1, Accuracy | Few-shot prompting with justification exemplars does not substantially improve performance, suggesting that in-context examples are insufficient for complex evidence appraisal tasks. |
| Lee, A.V.Y. et al. [66] | Few-shot | Education; commonsense reasoning; knowledge creation discourse | Human qualitative comparison | Few-shot prompting improves response relevance and structure by providing exemplars aligned with human expectations. |
| Zhen et al. [33] | Few-shot | Traffic crash severity classification | Macro F1, Macro Accuracy | Few-shot prompting improves performance mainly for smaller models (LLaMA3-8B), but introduces trade-offs across severity categories. |
| Shah et al. [42] | Few-shot | Clinical QA, medical education, classification | Not applicable (review) | Few-shot prompting improves performance by clarifying task structure and reducing ambiguity, particularly in domain-specific medical queries. |
| Meshkin et al. [38] | Few-shot | Regulatory NLP; PK-DDI classification | Precision, Recall, F1-score | Few-shot prompting provides modest gains for some models but does not consistently outperform zero-shot learning in imbalanced regulatory datasets. |
| Ono, Dickson & Koga [67] | Few-shot | Histopathology image classification | Accuracy, Sensitivity, Specificity | Few-shot prompting with a small number of reference images improves diagnostic accuracy to ~50–60%, but performance remains unstable. |
| Viswanathan et al. [39] | Few-shot | Semi-supervised text clustering | Macro F1, Pairwise F1 | LLMs can act as effective pseudo-oracles for pairwise constraints, approaching human-supervised clustering performance at lower cost. |
| Gu, X. et al. [61] | Few-shot | Sentiment classification | Accuracy, F1 | Few-shot AGCVT prompting achieves the highest accuracy while preserving interpretability and low data requirements. |
| Choi, J. [68] | Few-shot | Relevance evaluation in information retrieval | Cohen’s kappa | Few-shot prompting introduces variability and bias that can reduce alignment with human relevance judgments, even when increasing the number of examples. |
| Wan, T.; Chen, Z. [46] | Few-shot | Feedback generation for physics conceptual questions | Usefulness ratings | Providing a small number of example response–feedback pairs enables GPT to generate feedback that students perceive as more useful and more responsive to their reasoning. |
| Miao, J. et al. [35] | Few-shot | Medical QA and clinical reasoning | Qualitative synthesis of prior studies | Few-shot prompting improves task alignment and reduces hallucinations, but performance gains tend to plateau after a small number of examples. |
| Filienko, D. et al. [58] | Few-shot | Problem-Solving Therapy dialogue | Human Likert ratings; empathy metrics | Few-shot prompting with clinician-curated examples enables the model to internalize implicit therapeutic norms, resulting in more coherent, personalized, and protocol-adherent PST dialogues. |
| Ayad, S; Alsayoud, F [70] | Few-shot | Business process modeling | Empirical (experimental/applied) | Prompt includes examples to guide BPM model generation |
| Carlson, NA; Burbano, V [48] | Few-shot | Data annotation/text classification | Empirical (mixed-methods) | Including examples in prompt |
| Wang, J. [9] | Few-shot | Business communication writing | Empirical (experimental) | Providing genre-based outlines or examples |
| Tocev, T.; Atanasovski, A. [72] | Few-shot | IFRS advisory/financial reporting | Empirical (experimental) | Prompt includes worked IFRS examples before the target case |
| Carlson, NA; Burbano, V [48] | Role | Data annotation/text classification | Empirical (mixed-methods) | Assigning an identity to focus responses on relevant domain knowledge |
| Karam et al. [49] | Role | Change management/decision support | Empirical (experimental) | Implicit expert framing via allies strategy categories |
| Wang, J. [9] | Role | Business communication writing | Empirical (experimental) | Assigning ChatGPT an expert professional role |
| Tocev, T.; Atanasovski, A. [72] | Role | IFRS advisory/professional decision support | Empirical (experimental) | Assigning GPT-4 the role of an IFRS expert |
| Mao, W. et al. [50] | Role | Recommendation systems | Empirical (experimental) | Assigning expert or recommender roles to the LLM |
| Cheung, K.S. [75] | Role | Professional decision support | Conceptual (viewpoint) | Framing the LLM as a professional valuer |
Appendix B.3. Output Transparency Prompting Techniques
| Authors | Technique | Task/Domain | Evidence Type | Reported Effects |
|---|---|---|---|---|
| Chen et al. [51] | Chain-of-Thought | Vulnerability detection | F1, Accuracy, Recall | CoT-based prompting significantly improves both vulnerability detection performance and interpretability; removing CoT leads to a clear drop in F1. |
| Qi et al. [43] | Chain-of-Thought | Construction of Knowledge Graphs in the domain of equipment O&M | Precision, Recall, F1 | slight 2% increase in metrics. The dataset primarily consists of short texts, where the advantages of CoT—typically more pronounced with longer texts—are not fully realized. |
| Ghimire et al. [30] | Chain-of-Thought | Cost estimation in construction | human evaluation | 2.52 (84%) av confidence score marks 20% improvement from ZS. This improvement confirms that structured, step-by-step reasoning enhances both response reliability and alignment with estimation logic, highlighting improvements in response quality and module adherence. |
| Bharti et al. [31] | Chain-of-Thought | Peer reviews assessment | human evaluation, recall, precision, f1 | Most detailed and interpretable outputs and provided well structured, evidence based justifications |
| Köksal, A.; Alatan, A.A. [60] | Chain-of-Thought | small scale remote sensing | recall, precision, f1, param, rs | Model with CoT reasoning in its training captions achieved more than 1.5× military-related answers, while the precision drops by 3%, due to increasing false positives in C2. The intermediate reasoning in the CoT training likely taught the model what clues to look for. This aligns with observations that CoT can make models better at justifying and thereby correctly executing a task. Thus, incorporating reasoning-focused data is beneficial for fine-tuning multimodal models in this context. |
| Darwiyanto et al. [52] | Chain-of-Thought | automated program repair | plausible patches | The designed CoT prompting structure has been generally shown to improve the ability of LLMs to generate solutions for APR tasks. |
| Qiao et al. [76] | Chain-of-Thought | Medical | Accuracy; Recall; BLEU-1; reasoning-quality metrics | Embedding structured Chain-of-Thought during training improves reasoning coherence and interpretability compared to unstructured CoT. |
| Tony et al. [12] | Chain-of-Thought (CoT) | Code generation/software security | Vulnerability count; human evaluation | Chain-of-Thought improves reasoning transparency and helps the model avoid obvious insecure patterns during code generation. |
| Jeon at al [55] | Chain-of-Thought | Medical | Cohen’s d effect sizes calculated against this Control condition (baseline prompt) | Most consistent performance (M = 65.61%, SD = 17.17, 95% CI [60.68, 70.54]). stable performance compared to Control (d = −0.301 to 0.308, lowest in logic and facts tasks and highest in Chinese tasks) |
| Du et al. [74] | Chain-of-Thought | power equipment defect grading | human eval., accuracy, ecs | Improves raw grading accuracy but also generates explanations that are both practically useful and theoretically grounded. The high Trustworthiness scores underscore the value of grounding model inferences in curated domain knowledge, while the expert suggestions point the way toward future enhancements that integrate richer contextual features and uncertainty quantification into the reasoning pipeline. |
| Teng et al. [54] | Chain-of-Thought | Depression detection (clinical text analysis) | CCC, MAE, Accuracy | Chain-of-Thought prompting improves both severity estimation and transparency of clinical reasoning. |
| Chen et al. [56] | Chain-of-Thought | Long-document abstractive summarization (scientific, biomedical, legal, governmental texts) | ROUGE-1/2/L (F1), BLEU, BERTScore, FactCC, human evaluation | Chain-of-Thought prompting improves factual consistency, structural coherence, and content coverage in long-document summarization by enforcing step-by-step reasoning before summary generation. |
| Chen & Li [57] | Chain-of-Thought | Instruction following; length-controlled text generation | Acc%, Vlt%, Avg. response length, ROUGE, BLEU | Chain-of-Thought prompting significantly reduces length violations and improves reasoning coherence, enabling better compliance with strict length constraints. By introducing and analyzing COT, we gain deeper insights into model reasoning, enhancing output accuracy and enabling finer control over output length. |
| Zhang et al. [34] | Chain-of-Thought | Early sepsis diagnosis from structured clinical data | Accuracy, Precision, Recall, F1 | Chain-of-Thought prompting significantly improves F1 score (≈+7–8 pp vs. ML baselines) and enhances clinical interpretability by simulating expert reasoning without task-specific training. |
| Zhang et al. [77] | Chain-of-Thought | Reasoning tasks across NLP, multimodal tasks, and language agents | Not applicable (survey) | Across a wide range of studies, Chain-of-Thought prompting consistently enhances multi-step reasoning, interpretability, and task generalization, especially in large-scale LLMs (>10B parameters). |
| Yin & Guo [64] | Chain-of-Thought | Financial forecasting and portfolio optimization (quantitative finance) | Accuracy, Precision, Recall, F1; Sharpe Ratio; Alpha; Max Drawdown | Chain-of-Thought prompting enables step-by-step numerical reasoning that outperforms human analysts in earnings direction prediction and supports profitable, risk-controlled investment strategies. CoT enhanced model achieved superior accuracy, demonstrating LLM’s capacity to automate financial reasoning and reducing reliance on specialized expertise |
| Hlaing et al. [78] | Chain-of-Thought | Dental licensing exam QA (prosthodontics, Korean language) | Accuracy (%), confidence intervals, reliability (Cronbach’s α) | Models with native or elicited Chain-of-Thought reasoning achieve passing-level accuracy comparable to human averages, outperforming language-optimized models in a non-Indo-European language. |
| Park et al. [79] | Chain-of-Thought | Radiology report classification (TRC) | AUROC, Accuracy, Precision, Recall, F1 | CoT improves interpretability and reasoning structure but underperforms ICL in difficult categories when used alone. |
| Bukhary et al. [59] | Chain-of-Thought | Visual defect detection in AV software requirements | Precision, Recall, F1-score, Correctness score | Chain-of-Thought prompting improves reasoning quality and explanation completeness but may introduce over-analysis and hallucinated defects, especially in safety-critical visuals. |
| Liao et al. [80] | Chain-of-Thought | Traffic scene understanding & motion forecasting (autonomous driving) | BERTScore (Precision, Recall, F1) | CoT prompting enables LLMs to generate structured, human-like reasoning over traffic scenes, significantly enhancing semantic understanding without updating model weights. |
| Zhen et al. [33] | Chain-of-Thought | Traffic crash severity analysis & inference | Macro F1, Macro Accuracy | CoT prompting enables step-by-step reasoning over crash causes, improving overall inference quality and interpretability. |
| Shah et al. [42] | Chain-of-Thought | Clinical reasoning, decision support, education | Not applicable (review) | Chain-of-Thought prompting enhances multi-step clinical reasoning and interpretability, making LLM outputs more aligned with clinician expectations. |
| Lee et al. [47] | Chain-of-Thought | Automatic scoring | Accuracy, F1 | Zero-shot CoT alone does not substantially improve scoring accuracy, indicating that generic reasoning is insufficient for rubric-based grading. |
| Gu et al. [61] | Chain-of-Thought | Sentiment classification | Accuracy, F1 | Chain-of-thought improves reasoning transparency but manual construction is costly and non-scalable. |
| Yang, G. et al. [53] | Chain-of-Thought | Code generation | Pass@1, CoT-Pass@1 | Providing explicit reasoning steps enables lightweight language models to better understand control flow and constraints, resulting in significantly higher code correctness. |
| Hebenstreit et al. [69] | Chain-of-Thought | Multi-domain multiple-choice question answering | Krippendorff’s alpha; accuracy | The standard zero-shot CoT trigger “Let’s think step by step” improves performance consistently across models and datasets, demonstrating strong generalizability. |
| Miao et al. [35] | Chain-of-Thought | Clinical diagnosis and treatment planning in nephrology | Diagnostic accuracy, qualitative clinical evaluation | Chain-of-thought prompting enables step-by-step clinical reasoning that mirrors physician decision-making, leading to more accurate diagnoses and clearer justification of medical decisions. |
| Miao et al. [35] | Chain-of-Thought | Clinical decision support and auditing | Explainability analysis | By externalizing reasoning steps, chain-of-thought prompting enhances explainability, enabling error tracing, auditing, and greater trust in high-stakes medical decisions. |
| Filienko et al. [58] | Chain-of-Thought | Problem-Solving Therapy dialogue | Human Likert ratings; empathy metrics | Zero-shot chain-of-thought prompting enhances exploratory and reflective responses but can reduce actionability and protocol adherence in dialogue-based therapeutic tasks. |
| Thanasi-Boçe, M; Hoxha, J [62] | Chain-of-Thought | entrepreneurship-learning, critical thinking, problem solving | Empirical (mixed-methods) | It involves providing a prompt that presents a sequence of intermediated reasoning instructions or steps to complete a task |
| Ayad, S; Alsayoud, F [70] | Chain-of-Thought | Business process modeling | Empirical (experimental/applied) | Prompt requests step-by-step reasoning before generating BPM models |
| Carlson, NA; Burbano, V [48] | Chain-of-Thought | Data annotation/text classification | Empirical (mixed-methods) | Explicit step-by-step reasoning process |
| Karam et al. [49] | Chain-of-Thought | Organizational analysis | Empirical (experimental) | Stepwise reasoning via explicit class descriptions |
| Tocev, T.; Atanasovski, A. [72] | Chain-of-Thought | IFRS advisory/financial reporting | Empirical (experimental) | Sequential decomposition into subquestions with stepwise reasoning |
| Mao, W. et al. [50] | Chain-of-Thought | Personalized recommendation | Empirical (experimental) | Step-by-step reasoning to infer user preferences |
| Cheung, K.S. [75] | Chain-of-Thought | Property valuation reporting | Conceptual (viewpoint) | Structured step-by-step prompt aligned with RICS “Red Book” valuation standards |
| Jeon at al. [55] | Self-Consistency | Medical | Empirical | Moderate variability (d = −0.623 to 0.362), lowest in logic global facts and highest in causal judgement |
| Thanasi-Boçe, M; Hoxha, J [62] | Chain of verification | Transactional decision-making | Empirical (mixed-methods) | CoV will first generate an initial response, and after, it will create verification questions to fact-check its response. |
| Carlson, NA; Burbano, V [48] | Self-Consistency | Data annotation/text classification | Empirical (mixed-methods) | Multiple independent attempts at same task |
Appendix C. Hypothesis Testing
Appendix C.1. H1 Testing (Perceived Output Usefulness)
| Construct | Items | Cronbach’s α |
|---|---|---|
| Perceived Output Usefulness | 5 | 0.947 |
| Predictor | Estimate | SE | df | t | p |
|---|---|---|---|---|---|
| Zero-Shot (Intercept) | 3.285 | 0.053 | 691 | 61.45 | <0.001 *** |
| CoT | 0.843 | 0.067 | 1249 | 12.65 | <0.001 *** |
| Few-shot | 0.823 | 0.067 | 398 | 12.27 | <0.001 *** |
| Few-shot + CoT | 1.522 | 0.070 | 398 | 22.74 | <0.001 *** |
| One-shot | 0.741 | 0.068 | 399 | 11.03 | <0.001 *** |
| One-shot + CoT | 1.553 | 0.067 | 383 | 22.88 | <0.001 *** |
| Task Type (Synthesis) | −0.033 | 0.037 | 311 | −0.89 | 0.36 |
| Effect | Num df | Den df | F | p |
|---|---|---|---|---|
| Technique | 5 | 304 | 144.64 | <0.001 *** |
| Task Type | 1 | 304 | 0.81 | 0.370 |
| Interaction | 5 | 304 | 0.99 | 0.427 |
| Technique | Mean | SE | 95% CI Lower | 95% CI Upper |
|---|---|---|---|---|
| Zero-shot | 3.27 | 0.051 | 3.17 | 3.37 |
| CoT | 4.11 | 0.051 | 4.01 | 4.21 |
| Few-shot | 4.09 | 0.051 | 3.99 | 4.19 |
| Few-shot + CoT | 4.79 | 0.050 | 4.69 | 4.89 |
| One-shot | 4.01 | 0.051 | 3.91 | 4.11 |
| One-shot + CoT | 4.82 | 0.051 | 4.72 | 4.92 |
Appendix C.2. H2 Testing (Perceived Prompt Usefulness)
| Construct | Items | Cronbach’s α |
|---|---|---|
| Perceived Prompt Usefulness | 4 | 0.994 |
| Predictor | Estimate | SE | df | t | p |
|---|---|---|---|---|---|
| Zero-Shot (Intercept) | 3.078 | 0.111 | 135 | 27.65 | <0.001 *** |
| CoT | 1.051 | 0.066 | 376 | 15.83 | <0.001 *** |
| Few-shot | 1.089 | 0.067 | 377 | 16.15 | <0.001 *** |
| Few-shot + CoT | 1.778 | 0.069 | 377 | 26.61 | <0.001 *** |
| One-shot | 0.948 | 0.069 | 377 | 14.18 | <0.001 *** |
| One-shot + CoT | 1.876 | 0.068 | 380 | 27.71 | <0.001 *** |
| Frequency of Use | −0.037 | 0.026 | 103 | −1.39 | 0.169 |
| Random Effect | Variance | SD |
|---|---|---|
| Respondent (Intercept) | 0.038 | 0.196 |
| Residual | 0.141 | 0.376 |
| Predictor | Estimate | SE | df | t | p |
|---|---|---|---|---|---|
| (Intercept) | 3.154 | 0.190 | 379.80 | 16.63 | <0.001 *** |
| CoT | 1.053 | 0.263 | 358.90 | 4.01 | <0.001 *** |
| Few-shot | 0.852 | 0.246 | 364.00 | 3.46 | 0.001 *** |
| Few-shot + CoT | 1.905 | 0.257 | 380.50 | 7.42 | <0.001 *** |
| One-shot | 0.816 | 0.259 | 349.70 | 3.15 | 0.002 ** |
| One-shot + CoT | 1.667 | 0.262 | 373.60 | 6.36 | <0.001 *** |
| Frequency of Use | −0.057 | 0.049 | 379.90 | −1.16 | 0.245 |
| CoT × Frequency | 0.000 | 0.007 | 361.10 | 0.00 | 0.999 |
| Few-shot × Frequency | 0.062 | 0.064 | 366.40 | 0.97 | 0.331 |
| Few-shot + CoT × Frequency | −0.036 | 0.068 | 380.90 | −0.54 | 0.593 |
| One-shot × Frequency | 0.035 | 0.066 | 354.80 | 0.53 | 0.596 |
| One-shot + CoT × Frequency | 0.056 | 0.067 | 372.90 | 0.82 | 0.421 |
| Random Effect | Variance | SD |
|---|---|---|
| Respondent (Intercept) | 0.038 | 0.194 |
| Residual | 0.142 | 0.377 |
| Model Comparison | Δχ2 | df | p |
|---|---|---|---|
| Main effects vs. Interaction | 3.245 | 5 | 0.662 |
References
- Bankins, S.; Ocampo, A.C.; Marrone, M.; Restubog, S.L.D.; Woo, S.E. A multilevel review of artificial intelligence in organizations: Implications for organizational behavior research and practice. J. Organ. Behav. 2024, 45, 159–182. [Google Scholar] [CrossRef] [Scilit]
- Brynjolfsson, E.; Li, D.; Raymond, L.R. Generative AI at Work. Q. J. Econ. 2025, 140, 889–942. [Google Scholar] [CrossRef] [Scilit]
- Simkute, A.; Tankelevitch, L.; Kewenig, V.; Scott, A.E.; Sellen, A.; Rintel, S. Ironies of generative AI: Understanding and mitigating productivity loss in human-AI interaction. Int. J. Hum. Comput. Interact. 2025, 41, 2898–2919. [Google Scholar] [CrossRef] [Scilit]
- Sikha, V.K.; Siramgari, D.; Korada, L. Mastering Prompt Engineering: Optimizing Interaction with Generative AI Agents. J. Eng. Appl. Sci. Technol. 2023, 5, 1–8. [Google Scholar] [CrossRef] [Scilit]
- Bozkurt, A. Tell Me Your Prompts and I Will Make Them True: The Alchemy of Prompt Engineering and Generative AI. Open Prax. 2024, 16, 111–118. [Google Scholar] [CrossRef] [Scilit]
- Alostad, H. Large Language Models as Kuwaiti Annotators. Big Data Cogn. Comput. 2025, 9, 33. [Google Scholar] [CrossRef] [Scilit]
- Wardle, G.; Sušnjak, T. Image First or Text First? Optimising the Sequencing of Modalities in Large Language Model Prompting and Reasoning Tasks. Big Data Cogn. Comput. 2025, 9, 149. [Google Scholar] [CrossRef] [Scilit]
- Nugroho, H.S.; Shaferi, I. Analysis of Prompt Engineering Effectiveness in Stock Recommendation by ChatGPT: An Experimental Study in the Indonesian Market. Int. Conf. Sustain. Econ. Manag. Account. Proc. (ICSEMA) 2025, 1, 1755–1763. [Google Scholar] [CrossRef] [Scilit]
- Wang, J. Improving ChatGPT’s Competency in Generating Effective Business Communication Messages: Integrating Rhetorical Genre Analysis into Prompting Techniques. J. Tech. Writ. Commun. 2024, 54, 369–395. [Google Scholar] [CrossRef] [Scilit]
- Elnashar, A.; White, J.; Schmidt, D.C. Enhancing structured data generation with GPT-4o evaluating prompt efficiency across prompt styles. Front. Artif. Intell. 2025, 8, 1558938. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bae, J.; Kwon, S.; Myeong, S. Enhancing software code vulnerability detection using GPT-4o and Claude-3.5 Sonnet: A study on prompt engineering techniques. Electronics 2024, 13, 2657. [Google Scholar] [CrossRef] [Scilit]
- Tony, C.; Díaz Ferreyra, N.E.; Mutas, M.; Dhif, S.; Scandariato, R. Prompting techniques for secure code generation: A systematic investigation. ACM Trans. Softw. Eng. Methodol. 2025, 34, 1–53. [Google Scholar] [CrossRef] [Scilit]
- Son, M.; Lee, S. Advancing multimodal large language models: Optimizing prompt engineering strategies for enhanced performance. Appl. Sci. 2025, 15, 3992. [Google Scholar] [CrossRef] [Scilit]
- Panneer, S.V. Prompt Engineering for Conversational AI Systems: A Systematic Review of Techniques and Applications. Int. J. Sci. Res. Comput. Sci. Eng. Inf. Technol. 2025, 11, 733–741. [Google Scholar] [CrossRef] [Scilit]
- Chowdhury, S.; Budhwar, P.; Dey, P.K.; Joel-Edgar, S.; Abadie, A. AI-employee collaboration and business performance: Integrating knowledge-based view, socio-technical systems and organisational socialisation framework. J. Bus. Res. 2022, 144, 31–49. [Google Scholar] [CrossRef] [Scilit]
- Janiesch, C.; Zschech, P.; Heinrich, K. Machine learning and deep learning. Electron. Mark. 2021, 31, 685–695. [Google Scholar] [CrossRef] [Scilit]
- Banh, L.; Strobel, G. Generative artificial intelligence. Electron. Mark. 2023, 33, 63. [Google Scholar] [CrossRef] [Scilit]
- Naveed, H.; Khan, A.U.; Qiu, S.; Saqib, M.; Anwar, S.; Usman, M.; Akhtar, N.; Barnes, N.; Mian, A. A Comprehensive Overview of Large Language Models. ACM Trans. Intell. Syst. Technol. 2025, 16, 1–72. [Google Scholar] [CrossRef] [Scilit]
- Raiaan, M.A.K.; Mukta, M.S.H.; Fatema, K.; Fahad, N.M.; Sakib, S.; Mim, M.M.J.; Ahmad, J.; Ali, M.E.; Azam, S. A Review on Large Language Models: Architectures, Applications, Taxonomies, Open Issues and Challenges. IEEE Access 2024, 12, 26839–26874. [Google Scholar] [CrossRef] [Scilit]
- Castelvecchi, D. Can we open the black box of AI? Nature 2016, 538, 20–23. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fui-Hoon Nah, F.; Zheng, R.; Cai, J.; Siau, K.; Chen, L. Generative AI and ChatGPT: Applications, challenges, and AI-human collaboration. J. Inf. Technol. Case Appl. Res. 2023, 25, 277–304. [Google Scholar] [CrossRef] [Scilit]
- Feuerriegel, S.; Hartmann, J.; Janiesch, C.; Zschech, P. Generative AI. Bus. Inf. Syst. Eng. 2024, 66, 111–126. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.Y.; Zheng, Z.; Zhang, F.; Feng, J.C.; Fu, Y.Y.; Zhai, J.D.; He, B.-S.; Zhang, X.; Du, X.Y. A comprehensive taxonomy of prompt engineering techniques for large language models. Front. Comput. Sci. 2026, 20, 2003601. [Google Scholar] [CrossRef] [Scilit]
- Liu, X.; Wang, J.; Yuan, X.; Sun, J.; Dong, G.; Di, P.; Wang, W.; Wang, D. Prompting frameworks for large language models: A survey. ACM Comput. Surv. 2026, 58, 1–38. [Google Scholar] [CrossRef] [Scilit]
- Polo, F.M.; Xu, R.; Weber, L.; Silva, M.; Bhardwaj, O.; Choshen, L.; de Oliveira, A.F.M.; Sun, Y.; Yurochkin, M. Efficient multi-prompt evaluation of LLMs. In Proceedings of the 38th International Conference on Neural Information Processing Systems (NeurIPS ’24); Curran Associates Inc.: Red Hook, NY, USA, 2024; pp. 22483–22512. [Google Scholar] [CrossRef] [Scilit]
- Gozzi, M.; Di Maio, F. Comparative Analysis of Prompt Strategies for Large Language Models: Single-Task vs. Multitask Prompts. Electronics 2024, 13, 4712. [Google Scholar] [CrossRef] [Scilit]
- Vaira, L.A.; Lechien, J.R.; Abbate, V.; Gabriele, G.; Frosolini, A.; De Vito, A.; Maniaci, A.; Mayo-Yáñez, M.; Boscolo-Rizzo, P.; Saibene, A.M.; et al. Enhancing AI chatbot responses in health care: The SMART prompt structure in head and neck surgery. OTO Open 2025, 9, e70075. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Geroimenko, V. The Essential Guide to Prompt Engineering: Key Principles, Techniques, Challenges, and Security Risks; SpringerBriefs in Computer Science; Springer Nature: Cham, Switzerland, 2025. [Google Scholar] [CrossRef] [Scilit]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. J. Clin. Epidemiol. 2021, 134, 178–189. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ghimire, P.; Kim, K.; Stentz, T.; Roy, T. Modular Chain-of-Thought (CoT) for LLM-Based Conceptual Construction Cost Estimation. Buildings 2026, 16, 396. [Google Scholar] [CrossRef] [Scilit]
- Bharti, P.K.; Panchal, M.; Dalal, V.; Agarwal, M.; Ekbal, A. Not all peer reviews are significant: A dataset of exhaustive vs. trivial scientific peer reviews leveraging chain-of-thought reasoning. Scientometrics 2026, 131, 2261–2301. [Google Scholar] [CrossRef] [Scilit]
- Giannilias, T.; Papadakis, A.; Nikolaou, N.; Zahariadis, T. Classification of Hacker’s Posts Based on Zero-Shot, Few-Shot, and Fine-Tuned LLMs in Environments with Constrained Resources. Future Internet 2025, 17, 207. [Google Scholar] [CrossRef] [Scilit]
- Zhen, H.; Shi, Y.; Huang, Y.; Yang, J.J.; Liu, N. Leveraging Large Language Models with Chain-of-Thought and Prompt Engineering for Traffic Crash Severity Analysis and Inference. Computers 2024, 13, 232. [Google Scholar] [CrossRef] [Scilit]
- Zhang, W.; Wu, M.; Zhou, L.; Shao, M.; Wang, C.; Wang, Y. A sepsis diagnosis method based on Chain-of-Thought reasoning using Large Language Models. Biocybern. Biomed. Eng. 2025, 45, 269–277. [Google Scholar] [CrossRef] [Scilit]
- Miao, J.; Thongprayoon, C.; Suppadungsuk, S.; Krisanapan, P.; Radhakrishnan, Y.; Cheungpasitporn, W. Chain of Thought Utilization in Large Language Models and Application in Nephrology. Medicina 2024, 60, 148. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Indira Kumar, A.K.; Sthanusubramoniani, G.; Gupta, D.; Nair, A.R.; Alotaibi, Y.A.; Zakariah, M. Multi-task detection of harmful content in code-mixed meme captions using large language models with zero-shot, few-shot, and fine-tuning approaches. Egypt. Inform. J. 2025, 30, 100683. [Google Scholar] [CrossRef] [Scilit]
- Razumovskaia, E.; Vulić, I.; Korhonen, A. Analyzing and Adapting Large Language Models for Few-Shot Multilingual NLU: Are We There Yet? Trans. Assoc. Comput. Linguist. 2025, 13, 1096–1120. [Google Scholar] [CrossRef] [Scilit]
- Meshkin, H.; Zirkle, J.; Arabidarrehdor, G.; Chaturbedi, A.; Chakravartula, S.; Mann, J.; Thrasher, B.; Li, Z. Harnessing large language models’ zero-shot and few-shot learning capabilities for regulatory research. Brief. Bioinform. 2024, 25, bbae354. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Viswanathan, V.; Gashteovski, K.; Lawrence, C.; Wu, T.; Neubig, G. Large language models enable few-shot clustering. Trans. Assoc. Comput. Linguist. 2024, 12, 321–333. [Google Scholar] [CrossRef] [Scilit]
- Golnari, P.; Prantzalos, K.; Hood, V.; Meskis, M.A.; Isom, L.L.; Wilcox, K.; Parent, J.M.; Lal, D.; Lhatoo, S.D.; Goodkin, H.P.; et al. Ontology accelerates few-shot learning capability of large language model: A study in extraction of drug efficacy in a rare pediatric epilepsy. Int. J. Med. Inf. 2025, 201, 105942. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Agbareia, R.; Omar, M.; Zloto, O.; Glicksberg, B.S.; Nadkarni, G.N.; Klang, E. Multimodal LLMs for retinal disease diagnosis via OCT: Few-shot versus single-shot learning. Ther. Adv. Ophthalmol. 2025, 17, 25158414251340569. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shah, K.; Xu, A.Y.; Sharma, Y.; Daher, M.; McDonald, C.; Diebo, B.G.; Daniels, A.H. Large Language Model Prompting Techniques for Advancement in Clinical Medicine. J. Clin. Med. 2024, 13, 5101. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Qi, X.; Yang, B.; Wang, S.; Zhang, Z.; Zhang, Y.; Du, K. Few-Shot and Chain-of-Thought Prompting for Equipment Maintenance Knowledge Graph Construction Via Large Language Models. SSRN 2025, 335, 115266. [Google Scholar] [CrossRef] [Scilit]
- Šuster, S.; Baldwin, T.; Verspoor, K. Zero- and few-shot prompting of generative large language models provides weak assessment of risk of bias in clinical trials. Res. Synth. Methods 2024, 15, 988–1000. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Byra, M.; Rachmadi, M.F.; Skibbe, H. Few-shot medical image classification with simple shape and texture text descriptors using vision-language models. Bull. Pol. Acad. Sci. Tech. Sci. 2025, 73, 153838. [Google Scholar] [CrossRef] [Scilit]
- Wan, T.; Chen, Z. Exploring generative AI assisted feedback writing for students’ written responses to a physics conceptual question with prompt engineering and few-shot learning. Phys. Rev. Phys. Educ. Res. 2024, 20, 010152. [Google Scholar] [CrossRef] [Scilit]
- Lee, G.-G.; Latif, E.; Wu, X.; Liu, N.; Zhai, X. Applying large language models and chain-of-thought for automatic scoring. Comput. Educ. Artif. Intell. 2024, 6, 100213. [Google Scholar] [CrossRef] [Scilit]
- Carlson, N.A.; Burbano, V. The use of LLMs to annotate data in management research: Foundational guidelines and warnings. Strateg. Manag. J. 2026, 47, 699–725. [Google Scholar] [CrossRef] [Scilit]
- Karam, B.A.E.; Fissaa, T.; Marghoubi, R. AI-Powered Assessment of Resistance to Change in the Context of Digital Transformation. Int. J. Adv. Comput. Sci. Appl. 2025, 16, 546. [Google Scholar] [CrossRef] [Scilit]
- Mao, W.; Wu, J.; Chen, W.; Gao, C.; Wang, X.; He, X. Reinforced prompt personalization for recommendation with large language models. ACM Trans. Inf. Syst. 2025, 43, 1–27. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.; Huang, Y.; Chen, X.; Shen, P.; Yun, L. GPTVD: Vulnerability detection and analysis method based on LLM’s chain of thoughts. Autom. Softw. Eng. 2026, 33, 3. [Google Scholar] [CrossRef] [Scilit]
- Darwiyanto, E.; Gusnaen, R.A.; Nurtantyana, R. A comparative study of large language models with chain-of-thought prompting for automated program repair. IAES Int. J. Artif. Intell. 2025, 14, 4579–4589. [Google Scholar] [CrossRef] [Scilit]
- Yang, G.; Zhou, Y.; Chen, X.; Zhang, X.; Zhuo, T.Y.; Chen, T. Chain-of-thought in neural code generation: From and for lightweight language models. IEEE Trans. Softw. Eng. 2024, 50, 2437–2457. [Google Scholar] [CrossRef] [Scilit]
- Teng, S.; Liu, J.; Jain, R.K.; Chai, S.; Hou, R.; Tateyama, T.; Lin, L.; Chen, Y.W. Enhancing depression detection with chain-of-thought prompting: From emotion to reasoning using large language models. In Proceedings of the 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Copenhagen, Denmark, 14–18 July 2025; pp. 1–6. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jeon, S.; Kim, H.-G. A comparative evaluation of chain-of-thought-based prompt engineering techniques for medical question answering. Comput. Biol. Med. 2025, 196, 110614. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, X.; Chen, Z.; Cheng, S. CoTHSSum: Structured long-document summarization via chain-of-thought reasoning and hierarchical segmentation. J. King Saud. Univ. Comput. Inf. Sci. 2025, 37, 40. [Google Scholar] [CrossRef] [Scilit]
- Chen, P.; Li, Z. Length Instruction Fine-Tuning with Chain-of-Thought (LIFT-COT): Enhancing Length Control and Reasoning in Edge-Deployed Large Language Models. Electronics 2025, 14, 1662. [Google Scholar] [CrossRef] [Scilit]
- Filienko, D.; Wang, Y.; Jazmi, C.E.; Xie, S.; Cohen, T.; De Cock, M.; Yuwen, W. Toward large language models as a therapeutic tool: Comparing prompting techniques to improve GPT-delivered problem-solving therapy. AMIA Annu. Symp. Proc. 2025, 2024, 417–426. [Google Scholar] [PubMed]
- Bukhary, N.; Ahmad, M.; Rashad, K.; Rai, S.; Shapsough, S.; Kaddoura, Y.; Dghaym, D.; Zualkernan, I. Few-Shot Evaluation of Vision Language Models for Detecting Visual Defects in Autonomous Vehicle Software Requirement Specifications. IEEE Access 2025, 13, 117914–117942. [Google Scholar] [CrossRef] [Scilit]
- Köksal, A.; Alatan, A.A. SAMChat: Introducing Chain-of-Thought Reasoning and GRPO to a Multimodal Small Language Model for Small-Scale Remote Sensing. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2026, 19, 795–804. [Google Scholar] [CrossRef] [Scilit]
- Gu, X.; Chen, X.; Lu, P.; Li, Z.; Du, Y.; Li, X. AGCVT-prompt for sentiment classification: Automatically generating chain of thought and verbalizer in prompt learning. Eng. Appl. Artif. Intell. 2024, 132, 107907. [Google Scholar] [CrossRef] [Scilit]
- Thanasi-Boçe, M.; Hoxha, J. From ideas to ventures: Building entrepreneurship knowledge with LLM, prompt engineering, and conversational agents. Educ. Inf. Technol. 2024, 29, 24309–24365. [Google Scholar] [CrossRef] [Scilit]
- Jaradat, S.; Elhenawy, M.; Nayak, R.; Paz, A.; Ashqar, H.I.; Glaser, S. Multimodal Data Fusion for Tabular and Textual Data: Zero-Shot, Few-Shot, and Fine-Tuning of Generative Pre-Trained Transformer Models. AI 2025, 6, 72. [Google Scholar] [CrossRef] [Scilit]
- Yin, M.; Guo, M. Complex forecasting and investment strategy optimization via chain-of-thought of large language models. Expert Syst. Appl. 2026, 298, 129913. [Google Scholar] [CrossRef] [Scilit]
- Shafikuzzaman, M.; Islam, M.R.; Zaman, S.; Ma, A.; Sifat, A.I. On the Effectiveness of Zero-Shot and Few-Shot Pretrained Language Models for Software Requirement Classification. IEEE Access 2025, 13, 159439–159453. [Google Scholar] [CrossRef] [Scilit]
- Lee, A.V.Y.; Teo, C.L.; Tan, S.C. Prompt Engineering for Knowledge Creation: Using Chain-of-Thought to Support Students’ Improvable Ideas. AI 2024, 5, 1446–1461. [Google Scholar] [CrossRef] [Scilit]
- Ono, D.; Dickson, D.W.; Koga, S. Evaluating the efficacy of few-shot learning for GPT-4Vision in neurodegenerative disease histopathology: A comparative analysis with convolutional neural network model. Neuropathol. Appl. Neurobiol. 2024, 50, e12997. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Choi, J. Binary or Graded, Few-Shot or Zero-Shot: Prompt Design for GPTs in Relevance Evaluation. Adv. Artif. Intell. Mach. Learn. 2024, 4, 2687–2702. [Google Scholar] [CrossRef] [Scilit]
- Hebenstreit, K.; Praas, R.; Kiesewetter, L.P.; Samwald, M. A comparison of chain-of-thought reasoning strategies across datasets and models. PeerJ Comput. Sci. 2024, 10, e1999. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ayad, S.; Alsayoud, F. Prompt engineering techniques for semantic enhancement in business process models. Bus. Process Manag. J. 2024, 30, 2611–2641. [Google Scholar] [CrossRef] [Scilit]
- Jovic, M.; Papakonstantinidis, S.; Kirkpatrick, R. From red ink to algorithms: Investigating the use of large language models in academic writing feedback. Lang. Test. Asia 2025, 15, 59. [Google Scholar] [CrossRef] [Scilit]
- Tocev, T.; Atanasovski, A. AI Chatbot as IFRS Advisory Tool: GPT-4 Experimental Design. Intell. Syst. Account. Financ. Manag. 2026, 33, e70031. [Google Scholar] [CrossRef] [Scilit]
- Jiang, Y.; De Raedt, M.; Deleu, J.; Demeester, T.; Develder, C. Few-shot out-of-scope intent classification: Analyzing the robustness of prompt-based learning. Appl. Intell. 2024, 54, 1474–1496. [Google Scholar] [CrossRef] [Scilit]
- Du, J.; Li, B.; Chen, Z.; Shen, L.; Liu, P.; Ran, Z. Knowledge-Augmented Zero-Shot Method for Power Equipment Defect Grading with Chain-of-Thought LLMs. Electronics 2025, 14, 3101. [Google Scholar] [CrossRef] [Scilit]
- Cheung, K.S. Real Estate Insights Unleashing the potential of ChatGPT in property valuation reports: The ‘Red Book’ compliance Chain-of-thought (CoT) prompt engineering. J. Prop. Investig. Financ. 2024, 42, 200–206. [Google Scholar] [CrossRef] [Scilit]
- Qiao, J.; Li, S.; Liu, J.; Yu, H.; Xiao, Y.; Yu, H.; Zheng, Y. Med-SCoT: Structured chain-of-thought reasoning and evaluation for enhancing interpretability in medical visual question answering. Comput. Med. Imaging Graph. 2025, 126, 102659. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, Z.; Yao, Y.; Zhang, A.; Tang, X.; Ma, X.; He, Z.; Wang, Y.; Gerstein, M.; Wang, R.; Zhao, H.; et al. Igniting Language Intelligence: The Hitchhiker’s Guide from Chain-of-Thought Reasoning to Language Agents. ACM Comput. Surv. 2025, 57, 1–39. [Google Scholar] [CrossRef] [Scilit]
- Hlaing, N.H.M.M.; Park, K.; Hahn, S.; Lee, S.Y.; Yeo, I.-S.L.; Lee, J.-H. Chain-of-Thought reasoning versus linguistic optimization for artificial intelligence models on the prosthodontics section of a dental licensing examination. J. Prosthet. Dent. 2026, 135, 394.e1–394.e8. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Park, J.; Sim, W.S.; Yu, J.Y.; Park, Y.R.; Lee, Y.H. Evaluation of Context-Aware Prompting Techniques for Classification of Tumor Response Categories in Radiology Reports Using Large Language Model. J. Imaging Inform. Med. 2025, 39, 2782–2793. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liao, H.; Kong, H.; Wang, B.; Wang, C.; Wang, K.Y.; He, Z.; Xu, C.; Li, Z. CoT-Drive: Efficient motion forecasting for autonomous driving with LLMs and chain-of-thought prompting. IEEE Trans. Artif. Intell. 2026, 7, 625–641. [Google Scholar] [CrossRef] [Scilit]
- Mayring, P. Qualitative Content Analysis: Theoretical Background and Procedures. In Advances in Mathematics Education; Bikner-Ahsbahs, A., Knipping, C., Presmeg, N., Eds.; Springer: Dordrecht, The Netherlands, 2015; pp. 365–380. [Google Scholar] [CrossRef] [Scilit]
- Hsieh, H.-F.; Shannon, S.E. Three Approaches to Qualitative Content Analysis. Qual. Health Res. 2005, 15, 1277–1288. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Stratton, S.J. Population Research: Convenience Sampling Strategies. Prehospital Disaster Med. 2021, 36, 373–374. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jager, J.; Putnick, D.L.; Bornstein, M.H., II. More Than Just Convenient: The Scientific Merits of Homogeneous Convenience Samples. Monogr. Soc. Res. Child Dev. 2017, 82, 13–30. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abeysinghe, B.; Circi, R. The challenges of evaluating LLM applications: An analysis of automated, human, and LLM-based approaches. In Proceedings of the First Workshop Large Language Models for Evaluation in Information Retrieval, Washington, DC, USA, 18 July 2024; Available online: https://ceur-ws.org/Vol-3752/paper1.pdf (accessed on 17 June 2026).
- Huang, Y.; Ma, T.; Yang, K.; Zhang, Z. FinSent-DistillQ: A distilled large language model with chain-of-thought fine-tuning for financial sentiment analysis. J. Intell. Inf. Syst. 2026, 64, 735–771. [Google Scholar] [CrossRef] [Scilit]
- Gibreel, O.; Arpaci, I. Development and validation of the prompt engineering competence scale (PECS). Inf. Dev. 2025, 02666669251336455. [Google Scholar] [CrossRef] [Scilit]
- Garg, A.; Nisumba Soodhani, K.; Rajendran, R. Enhancing data analysis and programming skills through structured prompt training: The impact of generative AI in engineering education. Comput. Educ. Artif. Intell. 2025, 8, 100380. [Google Scholar] [CrossRef] [Scilit]
- Tassoti, S. Assessment of Students Use of Generative Artificial Intelligence: Prompting Strategies and Prompt Engineering in Chemistry Education. J. Chem. Educ. 2024, 101, 2475–2482. [Google Scholar] [CrossRef] [Scilit]
- Vilakati, S. Prompt engineering for accurate statistical reasoning with large language models in medical research. Front. Artif. Intell. 2025, 8, 1658316. [Google Scholar] [CrossRef] [Scilit] [PubMed]









| Phase | Query |
|---|---|
| Exploratory | (“prompt engineering” OR “prompt design” OR “prompt enhancement” OR “prompting technique*” OR “prompt strategy” OR “prompt optimisation”) AND (“generative AI” OR “large language model*” OR “LLM” OR “GPT” OR “ChatGPT” OR “foundation model*” OR “text generation”) AND (“marketing” OR “business application” OR “business”) |
| Final | TITLE (“prompting techniques” OR “zero-shot” OR “one-shot” OR “few-shot” OR “chain-of-thought” OR “role prompt” OR “self-consistency” OR “chain-of-verification”) AND TITLE-ABS-KEY (“large language model*” OR “LLM*” OR “generative AI”) AND TITLE-ABS-KEY (“quality” OR “relevance” OR “usefulness” OR “effectiveness” OR “efficiency” OR “accuracy” OR “clarity” OR “reliability” OR “decision-making”) |
| Business Use Case | Prompt (Raw) |
|---|---|
| Review/Text Enhancement | Can you please correct this email that I need to send to my boss? It needs to have a formal and professional tone of voice. |
| Analysis | Can you please tell me more about these financial KPIs? |
| Synthesis | Can you please summarize this market research article for me? I’d like to save time... give me the main context and outcomes. |
| General Information Retrieval | Please tell me what time is in Sao Paulo now. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Cantini, A.; De Mauro, A. Better Prompts, Better Usefulness: A Systematic Review and Experimental Evaluation of Structured Prompting Techniques in Large Language Models. Big Data Cogn. Comput. 2026, 10, 224. https://doi.org/10.3390/bdcc10070224
Cantini A, De Mauro A. Better Prompts, Better Usefulness: A Systematic Review and Experimental Evaluation of Structured Prompting Techniques in Large Language Models. Big Data and Cognitive Computing. 2026; 10(7):224. https://doi.org/10.3390/bdcc10070224
Chicago/Turabian StyleCantini, Alessia, and Andrea De Mauro. 2026. "Better Prompts, Better Usefulness: A Systematic Review and Experimental Evaluation of Structured Prompting Techniques in Large Language Models" Big Data and Cognitive Computing 10, no. 7: 224. https://doi.org/10.3390/bdcc10070224
APA StyleCantini, A., & De Mauro, A. (2026). Better Prompts, Better Usefulness: A Systematic Review and Experimental Evaluation of Structured Prompting Techniques in Large Language Models. Big Data and Cognitive Computing, 10(7), 224. https://doi.org/10.3390/bdcc10070224
