Artificial Intelligence in Statistics Education: Leveraging LLMs for Analysis and Learning
Abstract
1. Introduction
2. Large Language Models: An Overview
3. Materials and Methods
3.1. Data and Applications
3.1.1. Lady Tasting Tea Dataset: Fisher’s Exact Test
- : The lady’s predictions are random (no ability)
- : The lady’s predictions are not random (some ability)
3.1.2. Titanic Dataset: Logistic Regression
- 1.
- PassengerId: A unique identifier assigned to each passenger.
- 2.
- Survived: An indicator if the passenger survived or not (0 = No, 1 = Yes).
- 3.
- Pclass: The class of the ticket purchased by the passenger (1 = first, 2 = second, 3 = third). This can also be an indicator of the passenger’s socio-economic status.
- 4.
- Name: The name of the passenger.
- 5.
- Sex: The sex of the passenger.
- 6.
- Age: The age of the passenger.
- 7.
- SibSp: The number of siblings or spouses of the passenger aboard the ship.
- 8.
- Parch: The number of parents or children of the passenger aboard the ship.
- 9.
- Ticket: The ticket number of the passenger.
- 10.
- Fare: The fare paid by the passenger for the journey.
- 11.
- Cabin: The cabin number of the passenger.
- 12.
- Embarked: The port of embarkation of the passenger (C = Cherbourg; Q = Queenstown; S = Southampton).
3.1.3. Iris Dataset: Plotting Clusters
- K-means clustering: The algorithm begins by randomly assigning each observation to a predefined number of clusters K. It then iteratively calculates the centroid (the centroid is the vector of the averages of the p variables for the observations in the k-th cluster) for each cluster and reassigns each observation to the cluster with the nearest centroid (proximity is usually measured by the quadratic Euclidean distance: ). This process repeats until the cluster assignments stabilize and no longer change.
- Hierarchical clustering: Builds a hierarchy of clusters, most commonly through an agglomerative (bottom-up) approach where individual observations are successively merged into larger clusters based on their proximity. This process continues until all observations are merged into a single cluster (conversely, a divisive (top-down) method starts with a single large cluster and recursively splits it until each observation forms its own cluster).
3.1.4. Better Life Index 2024 Dataset: Principal Component Analysis
3.1.5. Extrovert vs. Introvert Behavior Dataset: Supervised Machine Learning
- 1.
- Time_spent_alone: Number of hours an individual typically spends alone daily, ranging from 0 to 11 h.
- 2.
- Stage_fear: Experience of stage fright (Yes) or not (No).
- 3.
- Social_event_attendance: Frequency of attending social events, on a scale from 0 to 10.
- 4.
- Going_outside: Frequency of going outside, with a range from 0 to 7.
- 5.
- Drained_after_socializing: Feeling drained after socializing (Yes/No).
- 6.
- Friends_circle_size: Number of close friends, ranging from 0 to 15.
- 7.
- Post_frequency: Frequency of posting on social media, on a scale from 0 to 10.
- 8.
- Personality: Target variable for classification, designating individuals as either an Extrovert or an Introvert.
3.2. Experimental Design and Evaluation Criteria
- Accuracy: Correctness of code and statistical analysis, accuracy of outputs, ability to handle edge cases or unusual inputs, range of tasks supported, proper use of libraries, and the model’s understanding of the prompt.
- Readability: Code clarity, inclusion of helpful comments or documentation, and adherence to style guidelines (this includes verifying whether the code aligns with commonly accepted R programming conventions and best practices).
4. Results
4.1. Case Studies
4.1.1. Results for the Lady Tasting Tea Dataset
- Question to LLMs: A lady declares that by tasting a cup of tea made with milk, she can discriminate whether the milk or tea infusion was first added to the cup. We will consider the problem of designing an experiment through which this assertion can be tested. It consists of mixing eight cups of tea, four in one way and four in the other, and presenting them to the subject for judgment in random order. The subject has been told that the test will consist, namely, that she will be asked to taste eight cups, four of each kind. The lady in question answered correctly in six out of eight trials. The results are the following: (1) tea poured first: three correct answers out of four cups; (2) milk poured first: three correct answers out of four cups.With these results, what should be concluded about the lady’s ability to discriminate whether milk or tea was poured first? Propose a method to assess this statistically and the R code to run the analyses. Provide the response formatted in Latex, including the R code.
- GPT-3.5 response: To assess the lady’s ability to discriminate whether milk or tea was poured first, we can use a binomial test. The null hypothesis would be that the lady’s ability to discriminate is no better than random chance (50% correct), and the alternative hypothesis would be that she can discriminate with better accuracy. To perform the statistical analysis in R, we can use the binom.test() function. Here’s the R code to run the analyses:
| Listing 1. R code generated by GPT-3.5 for the Lady Tasting Tea experiment. |
|
- GPT-4o response: The experimental data can be analyzed using Fisher’s Exact Test. We first construct a 2 × 2 contingency table as follows:
| Correct Guess | Wrong Guess | |
|---|---|---|
| Milk First | 3 | 1 |
| Tea First | 3 | 1 |
| Listing 2. R code generated by GPT-4o for the Lady Tasting Tea experiment. |
|
- Gemini response: The lady’s claim of being able to discriminate between milk-first and tea-first tea can be tested statistically.We can use a binomial test to assess if the lady’s success rate (6 out of 8) deviates significantly from what we would expect by chance alone (i.e., guessing). The binomial test assumes:
- Only two outcomes are possible (correct or incorrect) for each trial (cup).
- Trials are independent (the outcome of one cup doesn’t influence another).
- The probability of success (correct guess) remains constant across trials.
- Success = Correctly identifying milk-first or tea-first tea.
- Trials (n) = 8 cups of tea.
- Expected probability of success (p) = 0.5 (assuming pure chance, guessing).
| Listing 3. R code generated by Gemini for the Lady Tasting Tea experiment. |
|
- Comment to LLMs’ responses on the “Lady Tasting Tea” dataset: GPT-3.5 and GPT-4o propose different statistical tests for analyzing the results. GPT-3.5 recommends a binomial test, which is inappropriate in this context. The proposed R code performed two separate binomial tests against a p = 0.5 alternative, rather than a unified test for the two tea-milk pouring mixes. While a binomial test could assess if guessing was better than chance, our objective is to verify whether the pouring order (milk first or tea first) influences the lady’s correct guesses. For this, the Fisher exact test proposed by GPT-4o is the more fitting choice, as it considers more the success rates in both conditions, making it a more powerful test than the binomial test, which only considers the overall success rate. However, despite the correct table provided in the text by GPT-4, its implementation in the R code is wrong. As additional remarks, GPT-3.5 misuses the “%” character for R comments and includes LaTeX syntax errors such as “50%” instead of “50 \%”. Despite these issues, both models produce functioning R code. Gemini, like GPT-3.5, selects a binomial test, which is less appropriate for this experimental design. However, it appears more precise in its model explanation and result interpretation than GPT.
4.1.2. Results for the Titanic Dataset
- Question to LLMs: I have the R datatable “titanic_data” referring to the shipwreck of a large cruise ship. The data contains information about the passengers and whether they survived the shipwreck. In particular, the fields of the data table are: PassengerId: A unique identifier assigned to each passenger. Survived: An indicator if the passenger survived or not (0 = No, 1 = Yes). Pclass: The class of the ticket purchased by the passenger (1 = first class, 2 = second class, 3 = third class). This can also be an indicator of the passenger’s socio-economic status. Name: The name of the passenger. Sex: The sex of the passenger. Age: The age of the passenger. SibSp: The number of siblings or spouses of the passenger aboard the ship. Parch: The number of parents or children of the passenger aboard the ship. Ticket: The ticket number of the passenger. Fare: The fare paid by the passenger for the journey. Cabin: The cabin number of the passenger. Embarked: The port of embarkation of the passenger (C = Cherbourg; Q = Queenstown; S = Southampton). Suggest which could be relevant fields to include in a statistical model to predict the passenger probability of survival, propose the statistical method you think is most opportune given the information I provided, and display the R code to estimate the model.
- GPT-3.5 response: To predict the passengers’ probability of survival, we can consider a logistic regression model using the following relevant fields from the titanic_data data table:
- Survived (target variable): This is the indicator variable that represents whether the passenger survived or not. It will be the dependent variable in the statistical model.
- Pclass: The class of the ticket purchased by the passenger can be an important indicator of socioeconomic status and potential survival probabilities.
- Sex: Gender may have played a significant role in the survival rates during the shipwreck.
- Age: Age could be a relevant factor as it might have affected the priority given to certain age groups during the rescue efforts.
- SibSp: The number of siblings or spouses aboard the ship might have influenced the chances of survival.
- Parch: The number of parents or children aboard the ship could have affected the survival rates.
- Fare: The fare paid by the passenger might be an indicator of socio-economic status and could potentially impact survival probabilities.
- Embarked: The port of embarkation might be associated with different socio-economic backgrounds and could be a relevant factor in the model.
| Listing 4. R code generated by GPT-3.5 for logistic regression on the Titanic dataset. |
|
- GPT-4o response: Given the variables in the “titanic_data” data table, the following fields could be pertinent in predicting a passenger’s survival probability:
- Pclass: The ticket class might be an indicator of the passenger’s socio-economic status.
- Sex: Gender might have played a role in survival rates due to the “women and children first” protocol.
- Age: Similar to gender, age might have played a role in survival rates.
- SibSp and Parch: The number of family members could have influenced the survival rate.
- Fare: This might correlate with “Pclass” and socio-economic status.
- Embarked: The port of embarkation might be a proxy for socio-economic status.
| Listing 5. R code generated by GPT-4o for logistic regression on the Titanic dataset. |
|
- Gemini response: Here are the fields from your “titanic” data table that could be relevant for predicting passenger survival probability:
- Survived (Target Variable): This is the variable we want to predict (0 = No, 1 = Yes).
- Pclass: Ticket class can be an indicator of socio-economic status, potentially impacting access to lifeboats.
- Sex: Men and women may have had different survival rates due to social norms and lifeboat capacity.
- Age: Age could influence survival chances due to physical capabilities.
- SibSp: Number of siblings/spouses aboard might influence survival if they acted together.
- Parch: Number of parents/children aboard could impact survival decisions based on family responsibility.
- Fare: Higher fare might indicate access to better cabins closer to lifeboats.
- Embarked: Embarkation port might influence proximity to lifeboats when the evacuation started.
| Listing 6. R code generated by Gemini for logistic regression on the Titanic dataset. |
|
- 1.
- Loads the dplyr library for data manipulation (if not already loaded).
- 2.
- Loads the “titanic” data assuming it’s a data frame.
- 3.
- Selects the relevant features for the model.
- 4.
- Converts categorical variables like Embarked to factors for the model.
- 5.
- Builds the logistic regression model using the glm function with the following arguments:
- formula: Specifies the relationship between the target variable (Survived) and the independent variables.
- data: The data frame containing the model features.
- family = binomial: Indicates a binary classification problem.
- 6.
- Prints the summary of the model, which shows the coefficients, standard errors, p-values, and other information to evaluate the model’s performance.
- Data exploration: Analyze relationships between variables and identify potential outliers.
- Feature engineering: Create new features based on existing ones (e.g., combining SibSp and Parch).
- Variable selection: Use techniques like feature importance scores to identify the most impactful features.
- Model evaluation: Use metrics like accuracy, precision, recall, and AUC-ROC to assess the model’s performance.
- Comment to LLMs’ responses on the “Titanic” dataset: All three models—GPT-3.5, GPT-4o, and Gemini—correctly identify logistic regression as the appropriate method for predicting survival based on passenger characteristics. They also consistently select relevant predictor variables, excluding non-informative fields such as PassengerId, Name, Ticket, Cabin. The R code generated by the models is largely similar and syntactically correct. Both GPT-4 and Gemini provide valuable qualitative insights beyond just code. Notably, GPT-4 suggests handling missing data through imputation or exclusion, though it does not offer methods for managing empty values specifically. Both models recommend converting categorical variables (e.g., Sex, Embarked) into factors. Gemini further enhances its educational value by providing additional useful explanations for model improvement.
4.1.3. Results for the Iris Dataset
- Question to LLMs: I have a dataset named ‘iris’. The iris dataset is a built-in dataset in R that contains measurements for 4 different attributes (in centimeters) of 50 flowers from 3 different species. Propose R code to visually describe data and identify possible clusters and a measure of the goodness of the clusters classification.
- GPT-3.5 response: To visually describe the iris dataset and identify possible clusters along with a measure of the goodness of the clusters classification, you can use the following R code. This code will perform various steps such as data visualization, clustering, and evaluating the clustering results.
| Listing 7. R code generated by GPT-3.5 for clustering analysis on the Iris dataset. |
|
- 1.
- Loading libraries: Necessary libraries are loaded for data manipulation, visualization, and clustering.
- 2.
- Basic data exploration: The summary and str functions are used to get a basic understanding of the dataset.
- 3.
- Data visualization: A pair plot and a PCA plot are created to visualize relationships between variables and species.
- 4.
- Clustering: K-means clustering is performed with 3 clusters, as the dataset consists of 3 species.
- 5.
- Cluster visualization: The results of clustering are visualized using a PCA plot.
- 6.
- Cluster evaluation: Silhouette analysis and Dunn Index are used to evaluate the goodness of clustering.
- GPT-4o response: To visually describe the iris dataset, identify possible clusters, and measure the goodness of the cluster classification, you can use a combination of data visualization, clustering techniques, and cluster evaluation metrics in R. Below is a step-by-step R code that covers these tasks:
| Listing 8. R code generated by GPT-4o for clustering analysis on the Iris dataset. |
|
- 1.
- Load Libraries: We load the required libraries for data manipulation, visualization, and clustering.
- 2.
- Visualize the Dataset: We use pair plots and scatter plots to visualize the relationships between different variables in the dataset.
- 3.
- Clustering with K-means: We perform k-means clustering to identify potential clusters in the data and visualize these clusters.
- 4.
- Evaluate Clustering: We use silhouette width to evaluate the goodness of the clustering, which measures how similar an object is to its own cluster compared to other clusters.
- 5.
- Compare with Actual Species: We compare the clusters identified by k-means with the actual species labels in the dataset to see how well the clustering corresponds to the known classifications.
- Gemini response: Here’s the R code to visually describe the data, identify possible clusters, and assess the goodness of the clusters for the iris dataset:
| Listing 9. R code generated by Gemini for clustering analysis on the Iris dataset. |
|
- 1.
- Loads necessary libraries (factoextra and cluster).
- 2.
- Loads the iris dataset.
- 3.
- Visualizes the data using pair plots.
- 4.
- Standardizes the data (optional but recommended for k-means).
- 5.
- Uses the elbow method to determine the optimal number of clusters based on the Within-Sum-of-Squares.
- 6.
- Performs k-means clustering with the chosen number of clusters.
- 7.
- Visualizes the data points colored by their assigned cluster.
- 8.
- Calculates and plots the silhouette width for each data point.
- 9.
- Calculates the average silhouette width, which indicates cluster cohesion. Higher values imply better separation between clusters.
- Comment to LLMs’ responses on the “Iris” dataset: The responses from both versions of GPT are effective and differ primarily in the depth of analysis. GPT-3.5 suggests using Principal Component Analysis (PCA) to improve cluster visualization and evaluates clustering quality using both the silhouette index—also employed by GPT-4o—and the Dunn Index. Silhouette index, developed by Rousseeuw in 1987 [45] ranges from −1 to 1 and measures how similar an observation is to its own cluster compared to other clusters. The Dunn index instead measures the ratio of the minimum inter-cluster distance to the maximum intra-cluster distance, with higher values indicating well-separated and compact clusters. However, in doing so, GPT-3.5 omits the need to load the fpc package, which is required for calculating the Dunn Index. GPT-4o provides similar visualizations (pair and scatter plots) and uses k-means clustering, but adds value by comparing the resulting clusters to the actual species labels, enabling a more informative assessment of classification accuracy. Gemini also applies k-means clustering but, unlike GPT, uses the elbow method to determine the number of clusters, which is conceptually appropriate (GPT immediately suggests using 3 clusters based on the prompt’s specification of three different iris species). This approach involves iterating k-means for different values of k, and each time calculating the sum of the squared distances between each centroid and the points in its cluster. A plot is created with k values on the x-axis and the sum of squared distances on the y-axis. The point where the curve forms an “elbow” indicates the optimal number of clusters. However, Gemini R code contains syntax and logical errors. Specifically, the kmeans function expects the second argument (centers) to be either a single number (specifying the number of clusters) or a set of initial cluster centers, instead the script passes a sequence from 1 to 10. Finally, Gemini incorrectly applies the silhouette function; this function expects a clustering vector that provides the cluster assignments, not the k-means result object, and a distance matrix rather than the scaled data. Thus, while the rationale is sound, the implementation of the code requires correction.
4.1.4. Results for the Better Life Index 2024 Dataset
- Question to LLMs: I have a .csv file referring to the Better Life Index 2024 of the OECD. The data contains information about the scores of 24 indicators relative to 11 topics for different countries. Specifically, the fields (indicators) of the data table are:
- Country: Name of each country.
- Dwellings without basic facilities: Indicator that refers to the percentage of the population living in a dwelling without an indoor flushing toilet for the sole use of their household.
- Housing expenditure: Indicator that considers the expenditure of households in housing and maintenance of the house, as defined in the SNA (P31CP040: Housing, water, electricity, gas, and other fuels; P31CP050: Furnishings, households’ equipment, and routine maintenance of the house).
- Rooms per person: Indicator that refers to the number of rooms (excluding kitchenette, scullery/utility room, bathroom, toilet, garage, consulting rooms, office, shop) in a dwelling divided by the number of persons living in the dwelling.
- Household net adjusted disposable income: The maximum amount that a household can afford to consume without having to reduce its assets or increase its liabilities.
- Household net wealth: Considers the total wealth: financial and non-financial assets, net of liabilities, held by private households resident in the country.
- Labour market insecurity: This indicator is defined in terms of the expected earnings loss, measured as the percentage of the previous earnings, associated with unemployment.
- Employment rate: The number of employed persons aged 15 to 64 over the population of the same age.
- Long-term unemployment rate: This indicator refers to the number of persons who have been unemployed for one year or more as a percentage of the labor force (the sum of employed and unemployed persons).
- Personal earnings: Refers to the average annual wages per full-time equivalent dependent employee.
- Quality of support network: A measure of perceived social network support. The indicator is based on the question: “If you were in trouble, do you have relatives or friends you can count on to help you whenever you need them, or not?” and considers the respondents who respond positively.
- Educational attainment: Considers the number of adults aged 25 to 64 holding at least an upper secondary degree over the population of the same age.
- Student skills: Students’ average score in reading, mathematics, and science as assessed by the OECD’s Programme for International Student Assessment.
- Years in education: The average duration of education in which a 5-year-old child can expect to enroll during his/her lifetime until the age of 39.
- Air pollution: The population-weighted average of annual concentrations of particulate matter less than 2.5 microns in diameter (PM2.5) in the air.
- Water quality: Captures people’s subjective appreciation of the environment where they live, in particular the quality of the water. It is based on the question: “In the city or area where you live, are you satisfied or dissatisfied with the quality of water?” and considers people who responded that they are satisfied.
- Stakeholder engagement for developing regulations: This indicator describes the extent to which formal stakeholder engagement is built into the development of primary laws and subordinate regulations.
- Voter turnout: The ratio between the number of individuals who cast a ballot during an election (whether this vote is valid or not) to the population registered to vote.
- Life expectancy: Measures how long on average people could expect to live based on the age-specific death rates currently prevailing.
- Self-reported health: Refers to the percentage of the population aged 15 years old and over who report good or better health.
- Feeling safe walking alone at night: Based on the question: “Do you feel safe walking alone at night in the city or area where you live?” and shows people declaring they feel safe.
- Homicide rate: Deaths due to assault.
- Employees working very long hours: Measures the proportion of dependent employed whose usual hours of work per week are 50 h or more.
- Time devoted to leisure and personal care: Measures the number of minutes (or hours) per day that, on average, full-time employed people spend on leisure and personal care activities.
- Life satisfaction: Considers people’s evaluation of their life as a whole. It is a weighted sum of different response categories based on people’s ratings of their current life relative to the best and worst possible lives for them on a scale from 0 to 10, using the Cantril Ladder (also known as the “Self-Anchoring Striving Scale”).
- GPT-3.5 response: To reduce the dimensionality of the data from the Better Life Index 2024, Principal Component Analysis (PCA) is a suitable statistical method. PCA helps in transforming the data into a set of linearly uncorrelated variables called principal components, which retain most of the variation present in the original dataset. This method is particularly useful when dealing with a large number of correlated variables. Here is the R code to perform PCA on the dataset:
| Listing 10. R code generated by GPT-3.5 for principal component analysis on the Better Life Index dataset. |
|
- 1.
- Loading Necessary Libraries: The readr and tidyverse libraries are used to load and manipulate the data.
- 2.
- Loading the Data: The read_csv function reads the CSV file containing the Better Life Index data.
- 3.
- Selecting Indicators: The select function is used to exclude the ‘Country’ column and keep only the indicators for PCA.
- 4.
- Standardizing the Data: The scale function standardizes the data to have a mean of 0 and a standard deviation of 1, which is essential for PCA.
- 5.
- Performing PCA: The prcomp function performs PCA on the standardized data.
- 6.
- Summary of PCA Result: The summary function provides a summary of the PCA result, including the proportion of variance explained by each principal component.
- 7.
- Scree Plot: The screeplot function generates a scree plot to visualize the variance explained by each principal component.
- 8.
- Biplot: The biplot function creates a biplot to visualize the first two principal components and the loadings.
- 9.
- Extracting and Printing Loadings: The loadings of the first two principal components are extracted and printed.
- GPT-4o response: To reduce the dimensionality of your data set, Principal Component Analysis (PCA) is a suitable method. PCA transforms the data into a set of linearly uncorrelated variables called principal components, which capture the most variance in the data. Here’s how you can perform PCA in R:
| Listing 11. R code generated by GPT-4o for principal component analysis on the Better Life Index dataset. |
|
- Gemini response: For this dataset, Principal Component Analysis (PCA) is a well-suited method for dimensionality reduction. PCA identifies a new set of features, called principal components (PCs), that capture the most variance in the original data.
- High Dimensionality: The dataset has 24 indicators, which can be challenging to visualize and analyze directly. PCA reduces the number of dimensions while retaining the most important information.
- Mixed Data Types: While most indicators are numerical, “Water quality” and “Feeling safe walking alone at night” are categorical. PCA can handle mixed data types to some extent. (Note: You might need to pre-process categorical data before applying PCA).
| Listing 12. R code generated by Gemini for principal component analysis on the Better Life Index dataset. |
|
- This code assumes your data is stored in a CSV file named “Better_Life_Index_2024.csv”. Adjust the file name if needed.
- Preprocessing categorical data might be necessary before applying PCA. You can explore methods like one-hot encoding or label encoding based on the nature of the data.
- Choosing the number of principal components (n_components) is a crucial step. You can use a scree plot or a predefined threshold based on the explained variance.
- Comment to LLMs’ responses on the “Better Life Index 2024” dataset: All three models correctly identify PCA as a suitable technique for dimensionality reduction, given the dataset’s high number of indicators. Both GPT and Gemini normalized the indicators, which were expressed in different units (e.g., dollars or number of people). Normalizing variables before performing PCA is essential, as otherwise, variables with the largest mean and variance would dominate the principal components. No version of GPT provides guidance on how to handle infinite or missing values. GPT codes suggest two graphs: a scree plot to display the variance explained by each component, and a biplot to display the first two principal components. GPT-4 uses the factoextra package to produce more visually appealing graphs. The code produced by Gemini outlines all the major steps of PCA, though some elements, particularly the visualization, are left as comments or require user input to complete.
4.1.5. Results for the Extrovert vs. Introvert Behavior Dataset
- Question to LLMs: I am working with a .csv file named “personality_dataset”, which contains 2900 records and 8 features related to social behavior and personality traits. The dataset includes the following variables:
- -
- Time_spent_alone: Hours spent alone daily (0–11).
- -
- Stage_fear: Presence of stage fright (Yes/No).
- -
- Social_event_attendance: Frequency of social events (0–10).
- -
- Going_outside: Frequency of going outside (0–7).
- -
- Drained_after_socializing: Feeling drained after socializing (Yes/No).
- -
- Friends_circle_size: Number of close friends (0–15).
- -
- Post_frequency: Frequency of social media posts (0–10).
- -
- Personality: Target variable indicating personality type (Extrovert/Introvert).
- Comment to LLMs’ responses on the “Extrovert vs. Introvert Behavior” dataset: To address the classification task, both GPT-3.5 and GPT-4o employ the tidymodels package along with necessary preprocessing libraries. They use Random Forests as their classification model and assume that the required packages are already installed and that the dataset is available in the working directory.
- 1.
- Incorrect use of make.names(): The argument unique = TRUE causes the target variable to have a unique value for each record. This option should be set to FALSE or omitted.
- 2.
- Unnecessary use of dummyVars(): This function is included in the code but leads to errors and is not required for the preprocessing task. It should be removed.
- 3.
- Incorrect handling of the target variable during AUC computation: The code first converts the target variable to numeric and then attempts to manually set its levels. Once a variable is numeric, setting levels is redundant and results in an error unless the variable is converted back to a factor.
4.2. Code Evaluation
5. Discussion and Conclusions
Limitations and Future Research
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| ML | Machine Learning |
| DL | Deep Learning |
| LLM | Large Language Models |
| GPT | Generative Pre-trained Transformer |
| BERT | Bidirectional Encoder Representations from Transformers |
| RLHF | Reinforcement Learning from Human Feedback |
| RM | Reward Model |
| NLP | Natural Language Processing |
| PCA | Principal Component Analysis |
| BLI | Better Life Index |
| ROC AUC | Receiver Operating Characteristic Area Under the Curve |
Appendix A
- GPT-3.5 response: Here’s a comprehensive R code solution using the tidymodels framework. This pipeline includes:
- Handling missing data
- Encoding categorical variables
- Normalizing numeric features
- Data splitting (70/30)
- Model training with tuning
- Model evaluation on the test set
| Listing A1. R code generated by GPT-3.5 for supervised machine learning. |
|
- GPT-4o response: To create a robust machine learning pipeline in R for predicting personality types from your “personality_dataset.csv”, we’ll follow these steps:
- Load data and necessary libraries
- Handle missing values
- Convert categorical variables to factors
- Split data into training and test sets (70/30)
- Preprocess using recipes (normalization, imputation, encoding)
- Train a classification model (e.g., Random Forest)
- Evaluate on test set
| Listing A2. R code generated by GPT-4o for supervised machine learning. |
|
- step_impute_∗ handles missing data robustly using median (for numeric) and mode (for categorical).
- step_dummy() ensures that categorical variables are suitable for the model.
- Random Forest is chosen for its balance of performance and robustness.
- You can easily swap in other models (e.g., XGBoost, logistic regression) by modifying the model block.
| Listing A3. R code generated by Gemini for supervised machine learning. |
|
References
- Moor, J. The Dartmouth College Artificial Intelligence Conference: The Next Fifty years. AI Mag. 2006, 27, 87–89. [Google Scholar]
- Grzybowski, A.; Pawlikowska-Łagód, K.; Lambert, W.C. A history of artificial intelligence. Clin. Dermatol. 2024, 42, 221–229. [Google Scholar] [CrossRef] [Scilit]
- Russell, S.J.; Norvig, P. Artificial Intelligence: A Modern Approach; Pearson: London, UK, 2016. [Google Scholar]
- McCorduck, P.; Cfe, C. Machines Who Think: A Personal Inquiry into the History and Prospects of Artificial Intelligence; AK Peters: Boca Raton, FL, USA; CRC Press: Boca Raton, FL, USA, 2004. [Google Scholar]
- Mitchell, T. Machine Learning; McGraw-Hill International Editions; McGraw-Hill: Columbus, OH, USA, 1997. [Google Scholar]
- LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I. Language models are unsupervised multitask learners. OpenAI Blog 2019, 1, 9. [Google Scholar]
- Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language Models are Few-Shot Learners. In Proceedings of the Advances in Neural Information Processing Systems; Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 1877–1901. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.u.; Polosukhin, I. Attention is All you Need. In Proceedings of the Advances in Neural Information Processing Systems; Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
- Kalla, D.; Smith, N.; Samaah, F.; Kuraku, S. Study and Analysis of chat GPT and its Impact on Different Fields of Study. Int. J. Innov. Sci. Res. Technol. 2023, 8, 827–833. [Google Scholar]
- Dai, W.; Lin, J.; Jin, H.; Li, T.; Tsai, Y.S.; Gašević, D.; Chen, G. Can large language models provide feedback to students? A case study on ChatGPT. In Proceedings of the 2023 IEEE International Conference on Advanced Learning Technologies (ICALT), 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 323–325. [Google Scholar]
- Ng, D.T.K.; Leung, J.K.L.; Chu, S.K.W.; Qiao, M.S. Conceptualizing AI literacy: An exploratory review. Comput. Educ. Artif. Intell. 2021, 2, 100041. [Google Scholar] [CrossRef] [Scilit]
- Colonna, L. Artificial Intelligence in Education (AIED): Towards More Effective Regulation. Eur. J. Risk Regul. 2025, 17, 161–181. [Google Scholar] [CrossRef] [Scilit]
- Zawacki-Richter, O.; Marín, V.I.; Bond, M.; Gouverneur, F. Systematic review of research on artificial intelligence applications in higher education—Where are the educators? Int. J. Educ. Technol. High. Educ. 2019, 16, 39. [Google Scholar] [CrossRef] [Scilit]
- Ifenthaler, D.; Majumdar, R.; Gorissen, P.; Judge, M.; Mishra, S.; Raffaghelli, J.; Shimada, A. Artificial intelligence in education: Implications for policymakers, researchers, and practitioners. Technol. Knowl. Learn. 2024, 29, 1693–1710. [Google Scholar] [CrossRef] [Scilit]
- Wongvorachan, T.; Srisuttiyakorn, S.; Sriklaub, K. Optimizing Learning: Predicting Research Competency via Statistical Proficiency. Trends High. Educ. 2024, 3, 540–559. [Google Scholar] [CrossRef] [Scilit]
- U.S. Bureau of Labor Statistics. Occupational Employment and Wage Statistics. Available online: https://www.bls.gov/oes/current/oes152041.htm (accessed on 15 March 2026).
- Macher, D.; Paechter, M.; Papousek, I.; Ruggeri, K.; Freudenthaler, H.H.; Arendasy, M. Statistics anxiety, state anxiety during an examination, and academic achievement. Br. J. Educ. Psychol. 2013, 83, 535–549. [Google Scholar] [CrossRef] [Scilit]
- McGrath, A.L. Content, affective, and behavioral challenges to learning: Students’ experiences learning statistics. Int. J. Scholarsh. Teach. Learn. 2014, 8, 6. [Google Scholar] [CrossRef] [Scilit]
- Lund, B.D.; Wang, T. Chatting about ChatGPT: How may AI and GPT impact academia and libraries? Libr. Tech News 2023, 40, 26–29. [Google Scholar] [CrossRef] [Scilit]
- Budzianowski, P.; Vulić, I. Hello, it’s GPT-2—How can I help you? Towards the use of pretrained language models for task-oriented dialogue systems. arXiv 2019, arXiv:1907.05774. [Google Scholar]
- Tamkin, A.; Brundage, M.; Clark, J.; Ganguli, D. Understanding the capabilities, limitations, and societal impact of large language models. arXiv 2021, arXiv:2102.02503. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Xia, C.S.; Wang, Y.; Zhang, L. Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. In Proceedings of the Advances in Neural Information Processing Systems; Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2023; Volume 36, pp. 21558–21572. [Google Scholar]
- Bucaioni, A.; Ekedahl, H.; Helander, V.; Nguyen, P.T. Programming with ChatGPT: How far can we go? Mach. Learn. Appl. 2024, 15, 100526. [Google Scholar] [CrossRef] [Scilit]
- Moussiades, L.; Zografos, G.; Papakostas, G. GPT-4 vs. GPT-3.5 as coding assistants. preprint 2024. Available online: https://www.researchsquare.com/article/rs-3920214/v1 (accessed on 15 March 2026).
- Heitz, L.B.; Chamas, J.; Scherb, C. Evaluation of the Programming Skills of Large Language Models. arXiv 2024, arXiv:2405.14388. [Google Scholar] [CrossRef] [Scilit]
- Elgedawy, R.; Sadik, J.; Dutta, S.; Gautam, A.; Georgiou, K.; Gholamrezae, F.; Ji, F.; Lim, K.; Liu, Q.; Ruoti, S. Ocassionally secure: A comparative analysis of code generation assistants. arXiv 2024, arXiv:2402.00689. [Google Scholar] [CrossRef] [Scilit]
- Hou, W.; Ji, Z. A systematic evaluation of large language models for generating programming code. arXiv 2024, arXiv:2403.00894. [Google Scholar] [CrossRef] [Scilit]
- Song, X.; Xie, K.; Lee, L.; Chen, R.; Clark, J.M.; He, H.; He, H.; Min, J.; Zhang, X.; Zheng, S.; et al. Performance Evaluation of Large Language Models in Statistical Programming. arXiv 2025, arXiv:2502.13117. [Google Scholar] [CrossRef] [Scilit]
- Beer, R.; Feix, A.; Guttzeit, T.; Muras, T.; Müller, V.; Rauscher, M.; Schäffler, F.; Löwe, W. Examination of Code generated by Large Language Models. arXiv 2024, arXiv:2408.16601. [Google Scholar] [CrossRef] [Scilit]
- Górecki, J. Pair programming with ChatGPT for sampling and estimation of copulas. Comput. Stat. 2024, 39, 3231–3261. [Google Scholar] [CrossRef] [Scilit]
- Buscemi, A. A comparative study of code generation using chatgpt 3.5 across 10 programming languages. arXiv 2023, arXiv:2308.04477. [Google Scholar] [CrossRef] [Scilit]
- Tian, H.; Lu, W.; Li, T.O.; Tang, X.; Cheung, S.C.; Klein, J.; Bissyandé, T.F. Is ChatGPT the ultimate programming assistant—How far is it? arXiv 2023, arXiv:2304.11938. [Google Scholar]
- Tucker, M.C.; Shaw, S.T.; Son, J.Y.; Stigler, J.W. Teaching Statistics and Data Analysis with R. J. Stat. Data Sci. Educ. 2023, 31, 18–32. [Google Scholar] [CrossRef] [Scilit]
- Baumer, B.; Cetinkaya-Rundel, M.; Bray, A.; Loi, L.; Horton, N.J. R Markdown: Integrating a reproducible analysis tool into introductory statistics. arXiv 2014, arXiv:1402.1894. [Google Scholar] [CrossRef] [Scilit]
- Nolan, D.; Lang, D.T. Data Science in R: A Case Studies Approach to Computational Reasoning and Problem Solving; CRC Press: Boca Raton, FL, USA, 2015. [Google Scholar]
- Mascaró, M.; Sacristán, A.I.; Rufino, M.M. For the love of statistics: Appreciating and learning to apply experimental analysis and statistics through computer programming activities. Teach. Math. Its Appl. Int. J. IMA 2016, 35, 74–87. [Google Scholar] [CrossRef] [Scilit]
- Fisher, R.A. The Design of Experiments; Oliver and Boyd: Edinburgh, UK, 1937; Volume 2. [Google Scholar]
- British Board of Trade. Report on the Loss of the ‘Titanic’ (S.S.); Allan Sutton Publishing: Gloucester, UK, 1990. [Google Scholar]
- Hastie, T.; Tibshirani, R.; Friedman, J.H.; Friedman, J.H. The Elements of Statistical Learning: Data Mining, Inference, and Prediction; Springer: Berlin, Germany, 2009; Volume 2. [Google Scholar]
- Fisher, R.A. The use of multiple measurements in taxonomic problems. Ann. Eugen. 1936, 7, 179–188. [Google Scholar] [CrossRef] [Scilit]
- Kaggle. Extrovert vs. Introvert Behavior Data. Available online: https://www.kaggle.com/datasets/rakeshkapilavai/extrovert-vs-introvert-behavior-data/data (accessed on 15 March 2026).
- Evtikhiev, M.; Bogomolov, E.; Sokolov, Y.; Bryksin, T. Out of the bleu: How should we assess quality of the code generation models? J. Syst. Softw. 2023, 203, 111741. [Google Scholar] [CrossRef] [Scilit]
- CRAN. lintr: A ‘Linter’ for R Code. Available online: https://cran.r-project.org/web/packages/lintr/index.html (accessed on 15 March 2026).
- Rousseeuw, P.J. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. J. Comput. Appl. Math. 1987, 20, 53–65. [Google Scholar] [CrossRef] [Scilit]
- HIX.AI. Chat Interface. Available online: https://hix.ai/chat (accessed on 15 March 2026).
- Zheng, Q.; Xia, X.; Zou, X.; Dong, Y.; Wang, S.; Xue, Y.; Wang, Z.; Shen, L.; Wang, A.; Li, Y.; et al. Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x. arXiv 2023, arXiv:2303.17568. [Google Scholar]
- Cheng, Y.; Sanders, M.; Le, A.T. Using LLM to Generate Questions and Answers for Statistics Education: Is It Adequate? In Artificial Intelligence in Education—Creating an Equitable, Creative, and Effective Learning Environment; IntechOpen: London, UK, 2026. [Google Scholar]
- Cheung, B.H.H.; Lau, G.K.K.; Wong, G.T.C.; Lee, E.Y.P.; Kulkarni, D.; Seow, C.S.; Wong, R.; Co, M.T.H. ChatGPT versus human in generating medical graduate exam multiple choice questions—A multinational prospective study (Hong Kong SAR, Singapore, Ireland, and the United Kingdom). PLoS ONE 2023, 18, e0290691. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Powell, W.; Courchesne, S. Opportunities and risks involved in using ChatGPT to create first grade science lesson plans. PLoS ONE 2024, 19, e0305337. [Google Scholar] [CrossRef] [Scilit]
- Lee, U.; Kim, Y.; Lee, S.; Park, J.; Mun, J.; Lee, E.; Kim, H.; Lim, C.; Yoo, Y.J. Can we Use GPT-4 as a Mathematics Evaluator in Education?: Exploring the Efficacy and Limitation of LLM-based Automatic Assessment System for Open-ended Mathematics Question. Int. J. Artif. Intell. Educ. 2024, 35, 1560–1596. [Google Scholar] [CrossRef] [Scilit]
- Rudolph, J.; Tan, S.; Tan, S. ChatGPT: Bullshit spewer or the end of traditional assessments in higher education? J. Appl. Learn. Teach. 2023, 6, 342–363. [Google Scholar] [CrossRef] [Scilit]
| Topics | Indicators |
|---|---|
| Housing | Dwellings without basic facilities |
| Rooms per person | |
| Housing expenditure | |
| Income | Household net adjusted disposable income |
| Household net financial wealth | |
| Jobs | Labour market insecurity |
| Employment rate | |
| Long term unemployment rate | |
| Personal earnings | |
| Education | Educational attainment |
| Students’ cognitive skills | |
| Expected years in education | |
| Environment | Air pollution |
| Satisfaction with water quality | |
| Civic Engagement | Stakeholder engagement for regulations |
| Voter turnout | |
| Health | Life expectancy at birth |
| Self-reported health status | |
| Safety | Feeling safe walking alone at night |
| Homicide rates | |
| Work-Life Balance | Employees working very long hours |
| Time devoted to leisure and personal care | |
| Community | Social network support |
| Life Satisfaction | Life satisfaction |
| Data | GPT-3.5 | GPT-4o | Gemini | |||
|---|---|---|---|---|---|---|
| Error | Style | Error | Style | Error | Style | |
| Lady Tasting Tea | 1 | 2 | 0 | 2 | 0 | 1 |
| Titanic | 0 | 1 | 0 | 2 | 0 | 2 |
| Iris | 0 | 3 | 0 | 2 | 0 | 6 |
| Better Life Index | 0 | 0 | 0 | 5 | 0 | 3 |
| Extrovert vs. Introvert | 0 | 1 | 0 | 6 | 0 | 30 |
| Data | GPT-3.5 | GPT-4o | Gemini | Comments | |||
|---|---|---|---|---|---|---|---|
| Acc. | Read. | Acc. | Read. | Acc. | Read. | ||
| Lady Tasting Tea | 3/5 | 4/5 | 5/5 | 5/5 | 4/5 | 5/5 | GPT-4o better understands the prompt and proposes Fisher’s test. |
| Titanic | 4/5 | 3/5 | 5/5 | 4/5 | 4/5 | 4/5 | GPT-3.5 lacks useful comments (e.g., on handling NAs). |
| Iris | 4/5 | 5/5 | 5/5 | 5/5 | 2/5 | 5/5 | Gemini provides an R code containing several errors. |
| Better Life Index | 4/5 | 4/5 | 4/5 | 4/5 | 3/5 | 4/5 | Gemini lacks supplementary code for visualizing results. |
| Extrovert vs. Introvert | 4/5 | 2/5 | 3/5 | 3/5 | 2/5 | 4/5 | GPT-3.5 more accurate, Gemini clearer. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
di Bella, E.; Preti, S. Artificial Intelligence in Statistics Education: Leveraging LLMs for Analysis and Learning. Trends High. Educ. 2026, 5, 39. https://doi.org/10.3390/higheredu5020039
di Bella E, Preti S. Artificial Intelligence in Statistics Education: Leveraging LLMs for Analysis and Learning. Trends in Higher Education. 2026; 5(2):39. https://doi.org/10.3390/higheredu5020039
Chicago/Turabian Styledi Bella, Enrico, and Sara Preti. 2026. "Artificial Intelligence in Statistics Education: Leveraging LLMs for Analysis and Learning" Trends in Higher Education 5, no. 2: 39. https://doi.org/10.3390/higheredu5020039
APA Styledi Bella, E., & Preti, S. (2026). Artificial Intelligence in Statistics Education: Leveraging LLMs for Analysis and Learning. Trends in Higher Education, 5(2), 39. https://doi.org/10.3390/higheredu5020039

