Abstract
OpenAI’s new flagship model, ChatGPT-4o, released on 13 May 2024, offers enhanced natural language understanding and more coherent responses. This paper investigates ChatGPT-4o’s capabilities in financial data analysis, including zero-shot prompting, time series analysis, risk and return analysis, and ARMA-GARCH estimation. ChatGPT-4o’s performance is generally comparable to traditional statistical software like Stata, though some errors and discrepancies arise due to differences in implementation. Despite these issues, our findings indicate that ChatGPT-4o has significant potential for real-world financial analysis. Integrating ChatGPT-4o into financial research and practice may lead to more efficient data processing, improved analytical capabilities, and better-informed investment decisions.
Keywords:
ChatGPT; large language models; artificial intelligence (AI); generative AI (GenAI); finance research; financial analysis; academia; data analysis; stock return JEL Classification:
G00; C89; O33
1. Introduction
Released on 13 May 2024, OpenAI’s new flagship model, ChatGPT-4o (with its “o” standing for “omni”), represents a significant leap forward in generative artificial intelligence (GenAI). Designed to process and reason across audio, vision, and text in real time, ChatGPT-4o marks a notable improvement over its previous models like GPT-3.5 and GPT-4. Its enhancements include expanded knowledge coverage, more coherent and contextually relevant responses, and advanced analytical capabilities, making it a versatile tool for complex tasks like financial analysis.
The release of ChatGPT-4o has garnered widespread attention in industry and the media. GPT-4o has superior text, vision, and audio processing capabilities, making it twice as fast and 50% cheaper than its predecessor, GPT-4 Turbo. The model has achieved high Elo scores, significantly outperforming previous models in tasks that require comprehension, speech recognition, translation, and visual perception.1 These qualities are particularly relevant to the finance sector, where vast amounts of complex data require timely and accurate analysis.
As with each new model release, the launch of ChatGPT-4o prompts a wave of quick tests from experts from different fields and industries to validate the model’s performance. The finance industry, in particular, may benefit from the model’s advanced data analysis capabilities, especially given that financial data are complex and dynamic. The versatility of ChatGPT-4o promises substantial improvements in financial research and practice.
However, human expertise remains indispensable. While ChatGPT-4o can assist in handling large datasets, executing computations, and performing financial analysis, it lacks crucial judgment for real-time decision making in volatile markets. Therefore, incorporating human oversight is essential to interpret outputs and validate results. This necessity is particularly evident in finance, where data such as stock returns, risk factors, and regression estimates can change rapidly and often require domain-specific knowledge and strategic insight.
Understanding what ChatGPT-4o excels at is crucial for determining its potential applications in financial analysis. The model’s enhanced natural language processing capabilities make it adept at interpreting financial reports, analyzing market trends, and assisting in decision making. Its ability to handle vast amounts of data and run complex statistical analyses is particularly valuable in finance. Despite the extensive literature on data analysis using previous versions of ChatGPT (see the literature review in Section 2), few studies have explored ChatGPT-4o’s expanded capabilities in this domain. This gap exists because the model is still new, and its full potential remains untapped. This paper aims to be one of the first to investigate the use of ChatGPT-4o for data analysis in finance, highlighting its strengths and identifying areas for further improvement.
In this paper, we comprehensively evaluate ChatGPT-4o’s performance in financial analysis, utilizing daily stock return data from CRSP and the Fama French Factors. We undertake various tests to assess the model’s capabilities, including zero-shot prompting, time series analysis, risk and return analysis, and ARMA-GARCH estimation. Our empirical results demonstrate ChatGPT-4o’s performance in financial analysis. The model excels in interpreting complex datasets and offering insights into AI-assisted financial analysis with human oversight. We also demonstrate that GenAI should be viewed as a complementary tool rather than a replacement for human expertise, particularly in areas that require real-time decision making and interpretation. Further research is needed on how GenAI can be effectively integrated with human expertise to enhance financial analysis.
In this empirical assessment, we compare ChatGPT-4o’s performance to results from Stata, a widely used statistical software platform in academics. These tests underscore the model’s potential to transform financial research. ChatGPT-4o’s ability to handle large datasets and generate comprehensive analyses efficiently enhances the accuracy and speed of data-driven decision-making processes. Furthermore, our comparative analyses reveal that ChatGPT-4o’s performance is comparable to traditional statistical software like Stata, with minor discrepancies primarily due to differences in implementation methods. This validation highlights the model’s reliability and practical utility in real-world financial scenarios. The implications of these findings are significant, suggesting that integrating ChatGPT-4o into financial research and practice can lead to more efficient data processing and improved analytical capabilities. While ChatGPT-4o demonstrates the ability to accurately identify key variables, perform robust statistical analyses, and generate financial metrics such as Sharpe ratios and market betas, human validation and oversight remain imperative, especially in interpreting market trends and unexpected market events. Ultimately, we find that ChatGPT-4o’s analytical efficiency has the potential to enhance data-driven financial processes, yet it remains far from being a substitute for human judgment.
Moreover, recent advancements in GenAI have significantly expanded the capabilities of financial analysis as major technology firms introduce increasingly sophisticated models. For instance, DeepSeek, a China-trained model released on 10 January 2025, reportedly competes with top-tier models at a fraction of the training cost. DeepSeek’s emergence highlights an industry-wide trend shifting from brute-force scaling to intelligent optimization, by balancing power and efficiency without demanding excessive compute resources. Likewise, Google’s Gemini, OpenAI’s o-series, and other AI frameworks from Microsoft, Meta, and Anthropic focus on interpretability, efficiency, and domain-specific expertise. These efforts reflect a broader shift in GenAI research: moving beyond mere text generation to produce reliable, context-aware analytical tools that complement human insight in critical decision-making processes.
The contributions of this paper are twofold. First, it provides a comprehensive, empirical evaluation of ChatGPT-4o’s capabilities in the context of financial analysis, filling a critical gap in the current literature. This evaluation thoroughly explains the model’s strengths and limitations in financial analysis. Second, it highlights the practical implications of integrating advanced GenAI models in finance, offering valuable insights for researchers and practitioners. These implications cover various aspects of financial operations, including market analysis, portfolio management, and risk assessment, demonstrating the model’s broad applicability. However, domain expertise and real-time human oversight remain crucial.
Looking ahead, the implications of ChatGPT-4o in future research in GenAI are profound. Its ability to integrate and analyze data across multiple modalities opens new avenues for interdisciplinary studies, combining insights from finance, economics, computer science, and other fields. The model’s advanced capabilities can drive innovation in financial research, leading to more accurate models, improved decision-making processes, and a deeper understanding of market dynamics. Furthermore, the model’s adaptability to different data types and contexts suggests potential for future enhancements and applications beyond the current scope of this study. This paper lays the groundwork for future studies, exploring the vast potential of ChatGPT-4o and setting the stage for its widespread adoption in finance and beyond.
2. Related Literature
Many early papers focused on the general applications of ChatGPT across broad disciplines. For instance, the literature discusses ChatGPT’s capacity in scientific writing (Alkaissi & McFarlane, 2023; González-Padilla, 2022; Hosseini et al., 2024) and potential applications in academia and industry (Dale, 2021; Alshater, 2022; Dai et al., 2023; Lin, 2023; Rahman & Watanobe, 2023; Schlosky et al., 2024; L. X. Liu et al., 2024). It addresses concerns raised by academia (da Silva, 2023; Frye, 2022; Gao et al., 2022; Khalil & Er, 2023; Stokel-Walker, 2023; Thorp, 2023), highlights the limitations of ChatGPT (Borji, 2023; X. Liu et al., 2023; Shen et al., 2023; Zhu et al., 2023; Rice et al., 2024), and imposes standards to eliminate ethical issues and bias in the implementation of ChatGPT (Anders, 2023; Liebrenz et al., 2023; Lund & Wang, 2023; Yeo-Teh & Tang, 2023). These studies emphasize the need for validation and human oversight to ensure accuracy and reliability in its outputs. While ChatGPT can significantly enhance efficiency and depth in research, there is a caution against over-reliance on AI-generated analyses.
Another stream of literature focuses on the importance of interdisciplinary collaboration to maximize the benefits of ChatGPT and similar technologies (e.g., Bang et al., 2023; van Dis et al., 2023; Zhu et al., 2023; Fahad et al., 2024; Rice et al., 2024). These works underline the potential for AI to drive innovation across various fields.
The adoption of ChatGPT in business disciplines has garnered significant attention, with numerous studies highlighting its transformative potential. Specifically, in finance, ChatGPT has been utilized for various applications, ranging from general financial analysis to more specific areas. Examples in general financial analysis include Alshater (2022), Dowling and Lucey (2023), Chen et al. (2023), Zaremba and Demir (2023), Feng et al. (2024), and more,2 who discuss the potential applications of ChatGPT in financial data processing and interpretation, modeling and simulation, hypothesis generation, automated report generation, and even drafting research papers. These studies demonstrate how ChatGPT can streamline and improve efficiency in these areas.
Ali and Aysan (2023) explore how ChatGPT’s capabilities can transform various financial services, including customer service automation, fraud detection, and personalized financial advice, highlighting the model’s ability to enhance efficiency and accuracy in financial operations. Bhatia et al. (2024) introduce FinTral, a family of GPT-4 level multimodal financial large language models, to handle a variety of financial data inputs, including text and numerical data, to improve the accuracy and utility of financial analysis and forecasting.
Specific financial areas also see growing interest in the application of ChatGPT. For instance, Jha et al. (2023) investigate its role in corporate finance, particularly in automating and enhancing decision-making processes. Smales (2023) examines the impact of ChatGPT on interpreting and responding to monetary policy announcements, revealing its potential to improve market efficiency. Li et al. (2024) study the model’s ability to identify optimistic biases in human financial analysts, offering insights into how AI can mitigate subjective biases in financial analysis.
Further applications include analyzing stock price movements (Lopez-Lira & Tang, 2024), evaluating stock-picking strategies (Pelster & Val, 2024), and conducting sentiment analysis (Fatouros et al., 2023). Additionally, Aldridge (2023) demonstrates the utility of ChatGPT in performing linear regression analysis, highlighting its versatility in handling complex statistical tasks within finance. Cao and Zhai (2023) analyze the impact of ChatGPT on financial research domains such as ESG, corporate culture, and Federal Reserve opinion analysis. Kim et al. (2024) discuss applying large language models in analyzing financial statements.
Beyond business disciplines, ChatGPT has shown remarkable promise in fields that require rigorous mathematical analysis. For example, Frieder et al. (2023) evaluate the mathematical reasoning abilities of ChatGPT versions and GPT-4 using datasets specifically curated for graduate-level mathematics. This study highlights strengths in basic mathematical tasks while pointing out limitations in more advanced, graduate-level problem solving. Korinek (2023) explores the application of large language models like ChatGPT in various domains and details how ChatGPT can assist with mathematical derivations, data analysis, coding, and other research tasks, significantly boosting productivity for researchers.
The recent literature investigates ChatGPT’s new features in interpreting visual information. Yang et al. (2023) investigate the model’s performance in integrating visual and textual information, highlighting its strengths in understanding and generating descriptions of images, recognizing objects, and interpreting visual data. The authors also identify areas where the model struggles, such as handling complex visual scenes and nuances in image context. Wang et al. (2024) evaluate the capabilities of ChatGPT in interpreting scientific figures in their article published in NPJ Precision Oncology. They highlight ChatGPT’s strengths in recognizing and explaining various plot types effectively, which can aid researchers in understanding complex data visualizations.
Overall, the literature highlights ChatGPT’s extensive applications across various disciplines, underscoring its potential to enhance business and scientific research through advanced mathematical analysis and data processing capabilities.
Our paper fits into the literature by comprehensively evaluating ChatGPT-4o’s capabilities, specifically in financial data analysis. By examining its performance in various tasks, such as zero-shot prompting, time series analysis, risk and return analysis, and ARMA-GARCH estimation, we fill a critical gap in the current literature and highlight the practical implications of integrating ChatGPT-4o into financial research and practice.
3. Data
The financial data for the thirty companies included in the Dow Jones Industrial Average (DJIA) are sourced from the Center for Research in Security Prices (CRSP). For each Dow 30 stock from 2019 to 2023, we collect its permno, date, ticker, comnam, permco, cusip, vol, ret, bid, ask, shrout, numtrd, and retx, along with the CRSP return (including all distributions) Value-Weighted Index (vwretd) and the S&P 500 Composite Index Return (sprtrn), at daily frequency. This comprehensive dataset ensures we capture a detailed and accurate picture of each company’s stock performance over the specified period. To calculate each stock’s Sharpe ratio and market beta, data on the daily risk-free rate and the Fama–French three factors are sourced from Kenneth French’s website.3 These variables are essential for robust financial analysis and understanding the stocks’ underlying risk and return dynamics.
Table 1 reports the structure of the data used in the empirical tests. Panel A lists the 30 components of the DJIA as of 15 May 2024, collected from CNBC.4 Panel B presents the data structure for the “Dow 30 Daily Returns” file, which consists of panel data encompassing 30 firms over five years. This structure allows for a detailed examination of stock performance over time and across different companies. Panel C details the data structure for the “FamaFrench_3Factors” file, noting that we divide the original data by 100 to align with the CRSP data in the analysis. This adjustment ensures consistency and comparability across different data sources, facilitating more accurate and meaningful analyses.
Table 1.
Data. This table reports the data structure we use in the empirical tests. Panel A contains 30 components from the Dow Jones Industrial Average Stocks List. Panel B contains the data structure for the “Dow 30 Daily Returns” File. Panel C contains the data structure for the “FamaFrench_3Factors” file.
4. Zero-Shot Prompting Analysis
Our analysis begins with zero-shot open-end prompting, which means that the prompt used to interact with ChatGPT-4o does not contain examples or demonstrations but asks it to generate results about the best approach. One advantage of zero-shot open-end prompting is that it may reveal unseen problems and provide unexpected solutions.
After uploading the “Dow 30 Daily Returns” File, illustrated in Panel B of Table 1, we ask ChatGPT-4o: “Act as a finance professor and statistician. Analyze this data for the daily returns of Dow 30 components from 2019 to 2023, pulled from the Center for Research in Security Prices (CRSP) via the Wharton Research Data Services (WRDS) platform. Perform analysis that can be used in academic papers. Illustrate the meaning and implications of the results”.
In Panel A of Table 2, ChatGPT-4o begins its response by identifying the variables/column headers of the input data. Out of 15 columns, ChatGPT-4o correctly identifies 14 header names. The only misidentified column is “sprtrn”, which should represent the S&P 500 Composite Index Return but is labeled as “Spread Return” in the responses.
Table 2.
Zero-shot prompting. This table reports the financial analysis with a zero-shot prompt. Panel A is the results generated using ChatGPT-4o. The “Prompt” refers to the request made by the authors, while the “Response” refers to the feedback provided by ChatGPT-4o). Panel B is an “Assessment” of the quality and accuracy of the ChatGPT-4o response provided by the authors.
ChatGPT-4o suggests that the statistical analysis should include five parts: descriptive statistics, distribution analysis, time series analysis, risk and return analysis, and correlation analysis. Next, ChatGPT-4o provides further details on what those analyses should entail. In implementing these analyses, ChatGPT-4o starts with descriptive statistics and distribution analysis. ChatGPT-4o mentions that the descriptive statistics should include the count, mean, standard deviation, minimum, 25th, 50th, 75th percentiles, maximum, skewness, and kurtosis without providing the actual numbers for these statistics. ChatGPT-4o suggests including a histogram, KDE plot, and QQ plot for the distribution statistics. In our example, it plots the distribution of the daily returns, QQ plots of the daily returns, and the time series of the daily returns for MSFT. In the subsequent time series analysis, ChatGPT-4o uses the time series plot for Microsoft’s daily returns as an example to illustrate return fluctuations over time. It suggests examining four implications: risk and return, distribution properties, volatility clustering, and modeling returns (such as using the GARCH model).
ChatGPT-4o proceeds with the analysis, including calculating annualized metrics (return, volatility, and Sharpe ratio), examining correlations, and performing additional time series analysis. In the correlation analysis, ChatGPT-4o provides a correlation matrix of daily returns for the Dow 30 components. It identifies stock pairs with high and low correlations and suggests implications for portfolio diversification, risk management, and investment strategy.
ChatGPT-4o demonstrates a robust understanding of financial data analysis and suggests comprehensive steps to analyze the Dow 30 daily returns. It accurately identifies most data columns and outlines essential descriptive and time series analyses. However, it falls short in providing actual descriptive statistics and mislabels one column, indicating room for improvement in accuracy and completeness. The suggested analyses include calculating annualized metrics, examining correlations, and performing additional time series analysis, which are crucial for understanding the underlying risk and return dynamics. ChatGPT-4o’s ability to identify high- and low-correlated stocks and suggest implications for portfolio diversification, risk management, and investment strategy highlights its potential utility in financial research despite some shortcomings in precision.
5. Summary Statistics
Summary statistics are typically reported in the first table in many finance papers. Descriptive statistics is also the first analysis we should conduct according to the zero-shot prompting analysis by ChatGPT-4o, as illustrated in the previous section. Therefore, we proceed with our analysis to compare the summary statistics generated from statistical software with those obtained from ChatGPT-4o prompts.
Table 3 reports the summary statistics for the “Dow 30 Daily Returns” file. We first present the results generated using Stata 18 Standard Edition in Panel A of Table 3. As shown in Panel B, we upload the “Dow 30 Daily Returns” file and ask ChatGPT-4o to “Act as a finance professor and statistician. Analyze this data for the daily returns data of Dow 30 components from 2019 to 2023, pulled from WRDS. Provide a summary statistic table that includes the mean, median, standard deviation, minimum, maximum, and number of observations for the following variables: return (ret), volume (vol), bid price (bid), ask price (ask), shares outstanding (shrout), and value-weighted return (vwretd)”.
Table 3.
Summary statistics. This table reports the summary statistics for the daily returns data of Dow 30 components from 2019 to 2023. Panel A presents the results generated using Stata. Panel B shows the results generated using ChatGPT-4o. The “Prompt” refers to the request made by the authors, while the “Response” refers to the feedback provided by ChatGPT-4o. Panel C offers an “Assessment” of the quality and accuracy of the ChatGPT-4o response provided by the authors.
ChatGPT-4o responds with the statistics it calculates. Interestingly, the statistics for return (ret) and volume (vol) do not match those generated by Stata, while those for bid price (bid), ask price (ask), shares outstanding (shrout), and value-weighted return (vwretd) do match. This discrepancy suggests that while ChatGPT-4o can accurately process and analyze most of the data, certain variables might be prone to errors, warranting further investigation and validation.
Overall, ChatGPT-4o provides a reasonably accurate set of summary statistics, demonstrating its utility in performing preliminary data analysis. However, the mismatches in key variables like return and volume highlight the need for cross-validation with traditional statistical software to ensure the accuracy and reliability of the results. These findings underscore the potential of integrating AI tools into financial research while also emphasizing the importance of careful verification.
6. Plotting and Insights
In Table 2, the response of ChatGPT-4o suggests that time series analysis should follow the descriptive and distribution statistics, and it plots the time series of daily returns for MSFT (Microsoft Corp). Therefore, we assess the ability of ChatGPT-4o to generate time series figures and conduct time series analysis (the latter is discussed in the regression analysis section, i.e., fitting an ARMA-GARCH model).
Table 4 evaluates a time series plot for Microsoft’s return (ret) and the market’s return (vwretd) from 2019 to 2023. Panel A shows the results generated using Stata, illustrating that Microsoft’s returns are more volatile than the market’s returns, with increased volatility during late 2019 and early 2020.
Table 4.
Time series analysis of returns of Microsoft and the market. This table evaluates a time series plot for Microsoft’s (ticker = MSFT) return (ret) and the market’s (vwretd) return from 2019 to 2023. Panel A presents the results generated using Stata. Panel B shows the results generated using ChatGPT-4o. The “Prompt” refers to the request made by the authors, while the “Response” refers to the feedback provided by ChatGPT-4o. Panel C offers an “Assessment” of the quality and accuracy of the ChatGPT-4o response provided by the authors.
As shown in Panel B, we upload the “Dow 30 Daily Returns” file and send the following prompts to ChatGPT-4o: “Act as a finance professor and statistician. Analyze this data for the daily returns data of Dow 30 components from 2019 to 2023, pulled from WRDS. Provide a time series plot for the return (ret) of Microsoft (ticker = MSFT) and the market return (vwretd)”.
Interestingly, ChatGPT-4o has an issue with producing a time series plot, although it does have the ability to generate a time series plot, as evidenced by Table 2. However, here, instead of generating figures, ChatGPT-4o provides a textual description of the data trends for the daily returns of Microsoft and the market return from 2019 to 2023. For Microsoft, it highlights the company’s typical fluctuations associated with a high-profile technology company. ChatGPT-4o notes that the daily return volatility reflects underlying market conditions, including periods of heightened uncertainty like the COVID-19 pandemic. ChatGPT-4o also mentions a positive correlation between Microsoft’s returns and the market returns, which was not requested but adds value to the analysis.
Overall, ChatGPT-4o shows potential in analyzing and describing time series data but faces challenges in generating visual representations directly. This limitation underscores the importance of using traditional statistical software or tools for precise visualizations. The descriptive insights provided by ChatGPT-4o demonstrate its analytical capabilities, yet the inconsistencies in generating figures highlight the need for further improvements in its application.
7. Risk and Return Analysis
In the previous two sections, both the results from Stata and the response from ChatGPT-4o indicate increased return volatility during the onset of the COVID-19 period. Following this observation, we examine stock returns around the COVID-19 lockdown. In Table 5, we aim to extract the daily returns of Dow 30 components before and after the first COVID-19 lockdown day (Sunday, 15 March 2020). Specifically, we need to obtain the daily returns on 13 March and 16 March, the sum of these two daily returns for each stock, and the average returns of all stocks on each day.
Table 5.
Returns around COVID-19 lockdown. This table assesses the stock returns on the trading days before and after the first COVID-19 lockdown day (Sunday, 15 March 2020). Panel A presents the results generated using Stata. Panel B shows the results generated using ChatGPT-4o. The “Prompt” refers to the request made by the authors, while the “Response” refers to the feedback provided by ChatGPT-4o. Panel C offers an “Assessment” of the quality and accuracy of the ChatGPT-4o response provided by the authors.
After uploading the “Dow 30 Daily Returns” file, we command ChatGPT-4o to “Act as a finance professor and statistician. Analyze this data for the daily returns data of Dow 30 components from 2019 to 2023, pulled from WRDS. Provide the return (ret) of each stock on the trading days before and after the first COVID lockdown day (Sunday, 15 March 2020). Compute the total return of each stock’s two trading days, and the average return of all stocks’ two trading days”.
ChatGPT-4o does not provide the returns directly but instead supplies step-by-step Python codes as a response. We need to copy and run the suggested code in Python. However, we only obtain partial results using the given Python code, specifically, the sum of the two daily returns for each stock, as shown in the first table in Panel C.
Upon reviewing the code, we find that it successfully extracts the daily returns for each stock from the original dataset but fails to display them. This oversight prevents the full visibility of the individual daily returns on 13 March and 16 March. To resolve this issue, we add a few lines of codes to ensure that both individual daily returns and their sums are displayed. The final results, incorporating these changes, are shown in the second table in Panel C.
To ensure the accuracy of the results provided by ChatGPT-4o, we perform a parallel analysis using Stata. The results generated by Stata are presented in Panel A. By comparing these results with those obtained from the modified Python code provided by ChatGPT-4o, we find that they are identical, confirming the reliability of the Python code after our adjustments.
Thus, overall, ChatGPT-4o demonstrates a strong capability in generating effective Python code for financial data analysis. However, users must be vigilant as the initial code may contain minor omissions or errors that require careful verification and potential adjustments.
In Table 2, the response from ChatGPT-4o suggests that we should analyze the risk and return of each stock and compute their annualized metrics. Specifically, it mentions three key metrics: annualized return, annualized volatility, and Sharpe ratio.
We focus on the Sharpe Ratio as it involves a multi-step calculation that tests the computation abilities of ChatGPT-4o. To obtain the Sharpe Ratio, we first need to compute the daily excess return (return minus risk-free rate), aggregate these to obtain the annual excess return, compute the annualized stock return volatility, and then use the annualized excess return to divide the annual return volatility. If ChatGPT-4o calculates the Sharpe Ratio correctly, it indicates it can accurately compute the preceding annualized return and volatility. Sharpe Ratio analysis is presented in Table 6.
Table 6.
Sharpe ratio analysis. This table assesses the Sharpe ratio, (stock return—risk-free rate)/(standard deviation of the stock return), for each stock each year. Panel A presents the results generated using Stata. Panel B shows the results generated using ChatGPT-4o. The “Prompt” refers to the request made by the authors, while the “Response” refers to the feedback provided by ChatGPT-4o. Panel C offers an “Assessment” of the quality and accuracy of the ChatGPT-4o response provided by the authors.
After uploading the “Dow 30 Daily Returns” file, we command ChatGPT-4o to “Act as a finance professor and statistician. Analyze this data for the daily returns data of Dow 30 components from 2019 to 2023, pulled from WRDS, and the daily risk-free rate, pulled from Kenneth R. French—Data Library. Compute the annualized excess return (raw return—risk-free rate) of each stock. Compute the standard deviation of excess returns for each stock each year. Then, calculate the Sharpe Ratio, the risk-adjusted return, for each stock each year. Report the Sharpe Ratio for each stock each year and the average Sharpe Ratio of all stocks each year”.
Again, ChatGPT-4o does not provide the results directly but instead supplies step-by-step solutions as a response. ChatGPT-4o suggests this step-by-step approach: (1) Load and merge the data. (2) Compute the annualized excess returns. (3) Calculate the standard deviation of the excess returns. (4) Compute the Sharpe ratios. (5) Report the Sharpe ratios and the average Sharpe ratio for each year. Each step comes with a Python code that we can copy and run on our machines.
However, we encountered some issues when executing the code provided by ChatGPT-4o. For instance, the excess returns were annualized, but the volatilities were not. Also, Python could not find the “date” column in the Fama–French dataset because the column name is actually “Date”. Finally, the initial commands suggested by ChatGPT-4o only displayed the first five rows of the results. We resolved these issues by annualizing the stock return volatilities, correcting the column name, and adjusting the code to display all results. After these adjustments, we ran the corrected code in Python and obtained results that were very similar to those from Stata, as shown in Panel B.5
Overall, ChatGPT-4o demonstrates a solid understanding of the process involved in calculating the Sharpe ratio and provides useful code for this purpose. However, users must be diligent in verifying and adjusting the provided code to ensure accuracy, particularly in handling data specifics such as column names and understanding annualization methods.
Next, we consider another commonly used risk measure for stocks, market beta, which is not suggested in ChatGPT-4o’s response in the zero-shot prompting, as shown in Table 2. Computing market beta involves multiple steps that can effectively test the computational abilities of ChatGPT-4o. To obtain the market beta, we first need to compute the daily excess return (return minus risk-free rate). Then, we run a regression by firm and year to store the slope coefficients. In Stata, this requires using local macros and loops to get the results. Table 7 reports the market beta analysis.
Table 7.
Market Beta Analysis. This table assesses the market (CAPM) beta for each stock each year. Panel A presents the results generated using Stata. Panel B shows the results generated using ChatGPT-4o. The “Prompt” refers to the request made by the authors, while the “Response” refers to the feedback provided by ChatGPT-4o. Panel C offers an “Assessment” of the quality and accuracy of the ChatGPT-4o response provided by the authors.
After uploading the “Dow 30 Daily Returns” file, illustrated in Panel B of Table 1, and the “FamaFrench 3Factors” file, illustrated in Panel C of Table 1, we command ChatGPT-4o to “Act as a finance professor and statistician. Analyze this data for the daily returns data of Dow 30 components from 2019 to 2023, pulled from WRDS, and the daily risk-free rate, pulled from Kenneth R. French—Data Library. Compute the CAPM’s Beta for each stock each year”.
ChatGPT-4o has an issue with inspecting the data columns, so it does not provide the results directly. Instead, ChatGPT-4o provides a detailed guide with Python code on how to analyze the data to compute the CAPM’s Beta for each stock each year on the local machine.
Unfortunately, in the initial trial, we cannot obtain any results using the provided Python code because it has a couple of errors. First of all, the name of the variable “excess market return” should be ‘MktRF’, not ‘Mkt-RF’. Second, the beta value for DOW in 2019 is missing because the value of the variable ‘ret’ is null on 2 April 2019. After correcting these errors in the Python code, we obtain the same results as those generated from Stata.
Overall, ChatGPT-4o demonstrates a solid understanding of the process involved in calculating market beta and provides useful code for this purpose. However, it has issues in reading the variable name precisely. Users must be diligent in verifying and adjusting the provided code to ensure accuracy, particularly in handling data specifics such as column names and dealing with missing values.
8. ARMA-GARCH Estimation
Lastly, to capture the volatility clustering and leverage effects in the market returns, we perform an autoregressive moving average (ARMA) model with generalized autoregressive conditional heteroskedasticity (GARCH) processes, namely ARMA–GARCH model, on the market return (vwretd). In Table 2, the response from the zero-shot prompting suggests that we should also model the return dynamics, explicitly mentioning GARCH. Estimating the ARMA-GARCH model is a complex task that involves multiple steps.
First, our data are daily, consisting of a panel of 30 firms over five years, meaning the vwretd values are repeated 30 times in the data. We need to keep only one set of the vwretd values over the 5 years (1258 trading days) and delete duplicated ones. Second, the estimation of the ARMA-GARCH model parameters is sensitive to option specifications. We obtain the results in Panel A of Table 8 in Stata using the following syntax: arch vwretd, arch(1) garch(1) ar(1) ma(1).
Table 8.
ARMA-GARCH Estimation. This table assesses the ARMA-GARCH estimation for the market return (vwretd). Panel A presents the results generated using Stata. Panel B shows the results generated using ChatGPT-4o. The “Prompt” refers to the request made by the authors, while the “Response” refers to the feedback provided by ChatGPT-4o. Panel C offers an “Assessment” of the quality and accuracy of the ChatGPT-4o response provided by the authors.
We upload the “Dow 30 Daily Returns” file and command ChatGPT-4o: “Act as a finance professor and statistician. Analyze this data for the daily returns data of Dow 30 components from 2019 to 2023, pulled from WRDS. Estimate an ARMA-GARCH model on the value-weighted return (vwretd)”.
ChatGPT-4o again has an issue with executing the model fitting directly. Instead, it provides the steps (load the data, estimate ARMA model, and estimate GARCH model) with Python codes. We copy the Python codes and run them on a local machine, but they come with error messages. We then copy and paste the error message into ChatGPT-4o, which provides another set of codes. However, it still has issues to work immediately.
Specifically, in the first trial, we did not obtain any results using the Python code provided by ChatGPT-4o because the ‘ARMA’ function in the ‘statsmodels.tsa.api’ module has been replaced by the ‘ARIMA’ function. After correcting this error, we obtain some estimation results, but these results are incorrect since they are produced using duplicated data. As mentioned above, the ‘vwretd’ values are repeated 30 times in the data. Thus, we need to delete duplicate data before fitting an ARMA-GARCH model. The final results are shown in Panel C of Table 8. However, they are still different from those produced by Stata. This is because ARMA and GARCH models are estimated simultaneously in Stata, while they are estimated sequentially in Python. Unfortunately, there is currently no package in Python that can estimate these two models jointly.
Overall, ChatGPT-4o demonstrates a fundamental understanding of the process involved in estimating ARMA-GARCH models and provides somewhat helpful code. However, users must be diligent in verifying and adjusting the provided code to ensure accuracy, particularly in handling deprecated functions and avoiding data duplication.
9. Conclusions
OpenAI’s ChatGPT-4o makes an ambitious leap forward in generative artificial intelligence. In this paper, we evaluate ChatGPT-4o’s performance in financial data analysis, using daily stock data from the CRSP database for 30 companies in the Dow Jones Industrial Average, along with the market return and risk-free data from Ken French’s library. We conduct tests including zero-shot prompting, time series analysis, risk and return analysis, and ARMA-GARCH estimation. Our findings reveal that ChatGPT-4o’s performance is comparable to traditional statistical software, such as Stata, with minor discrepancies due to differences in implementation methods. ChatGPT-4o demonstrates a robust understanding of financial data analysis, accurately interpreting complex datasets and delivering thorough evaluations. These capabilities may significantly enhance investment strategies and risk management practices.
Through detailed comparative analyses, we illustrate how ChatGPT-4o can transform financial research methodologies and enhance the efficiency of data-driven decision-making processes. Despite some initial challenges with data handling and coding errors, our adjustments show that ChatGPT-4o can be a powerful tool with appropriate verification and validation procedures. These results underscore ChatGPT-4o’s robust understanding of financial data, as evidenced by its ability to interpret complex datasets and offer insightful analyses. Nonetheless, human expertise remains indispensable in validating outputs, ensuring the contextual relevance of interpretations, and adapting to real-time market fluctuations. As highlighted throughout our discussion, ChatGPT-4o should be viewed as a complementary tool, rather than a replacement for seasoned analysts.
Our research underscores the practical utility of ChatGPT-4o in real-world financial analysis and sets the stage for its broader adoption in the finance industry. The model’s advanced natural language processing capabilities and analytical strength offer significant benefits for financial analysts, researchers, and practitioners. These results suggest that integrating ChatGPT-4o into financial research and practice can lead to more efficient data processing, improved analytical capabilities, and better-informed investment decisions.
Looking ahead, we aim to further explore ChatGPT-4o’s adaptability to different data types and contexts, including its application to other financial instruments and markets. Future research will also consider integrating ChatGPT-4o with emerging AI models, including DeepSeek, Google’s Gemini, and newer iterations of OpenAI’s o-series. Ultimately, continued innovation in GenAI-driven financial research promises to refine predictive models, optimize investment strategies, and promote a deeper understanding of market dynamics, all while underscoring the indispensable role of expert human oversight.
Author Contributions
Conceptualization, Z.F.; methodology, Z.F. and F.L.; software, Z.F. and F.L.; validation, W.-H.C. and B.L.; formal analysis, Z.F. and F.L.; writing—original draft preparation, F.L. and B.L.; writing—review and editing, W.-H.C. and Z.F.; visualization, Z.F. and F.L.; supervision, Z.F.; project administration, Z.F. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
Data was obtained through academic subscriptions obtained by our respective institutions, and the authors cannot, under the terms of the agreements, have the data sets available for deposit.
Conflicts of Interest
The authors declare no conflict of interest.
Notes
| 1 | OpenAI’s new multimodal “GPT-4 omni” combines text, vision, and audio in a single model (https://the-decoder.com/, accessed on 13 May 2024). |
| 2 | See, for example, Khan and Umer (2024); Shue et al. (2023); Yue et al. (2023). |
| 3 | https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html (accessed on 13 May 2024). |
| 4 | https://www.cnbc.com/dow-30/ (accessed on 13 May 2024). |
| 5 | The small difference is due to the method of annualization. The code in Stata uses the actual number of trading days to compute the annual excess return and its annual standard deviation (e.g., 252 days in 2019 and 253 days in 2020), whereas the code suggested by ChatGPT-4o uses 252 days for all years. If the actual trading days are used for annualization in both cases, the results from Stata and Python would be identical. |
References
- Aldridge, I. (2023). The AI revolution: From linear regression to ChatGPT and beyond and how it connects to finance. The Journal of Portfolio Management, 49(9), 64–77. [Google Scholar] [CrossRef] [Scilit]
- Ali, H., & Aysan, A. F. (2023). What will ChatGPT revolutionize in the financial industry? Modern Finance, 1(1), 116–129. [Google Scholar] [CrossRef] [Scilit]
- Alkaissi, H., & McFarlane, S. I. (2023). Artificial hallucinations in ChatGPT: Implications in scientific writing. Cureus, 15(2), e35179. [Google Scholar] [CrossRef] [Scilit]
- Alshater, M. M. (2022). Exploring the role of artificial intelligence in enhancing academic performance: A case study of ChatGPT. SSRN Electronic Journal, 4312358. [Google Scholar] [CrossRef] [Scilit]
- Anders, B. A. (2023). Is using ChatGPT cheating, plagiarism, or both, neither, or forward thinking? Patterns, 4(3), 100694. [Google Scholar] [CrossRef] [Scilit]
- Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Wilie, B., Lovenia, H., Ji, Z., Yu, T., Chung, W., & Fung, P. (2023). A multitasking, multilingual, multimodal evaluation of chatbot on reasoning, hallucination, and interactivity. arXiv, arXiv:2302.04023. [Google Scholar]
- Bhatia, G., Nagoudi, E. M., Cavusoglu, H., & Abdul-Mageed, M. (2024). FinTral: A family of GPT-4 level multimodal financial large language models. arXiv, arXiv:2402.10986. [Google Scholar]
- Borji, A. (2023). A categorical archive of ChatGPT failures. arXiv, arXiv:2302.03494. [Google Scholar]
- Cao, Y., & Zhai, J. (2023). Bridging the gap—The impact of ChatGPT on financial research. Journal of Chinese Economic and Business Studies, 21(2), 177–191. [Google Scholar] [CrossRef] [Scilit]
- Chen, B., Wu, Z., & Zhao, R. (2023). From fiction to fact: The growing role of generative AI in business and finance. Journal of Chinese Economic and Business Studies, 21(4), 471–496. [Google Scholar] [CrossRef] [Scilit]
- Dai, H., Liu, Z., Liao, W., Huang, X., Wu, Z., Zhao, L., Liu, W., Liu, N., Li, S., Zhu, D., & Li, X. (2023). ChatAug: Leveraging ChatGPT for text data augmentation. arXiv, arXiv:2302.13007. [Google Scholar]
- Dale, R. (2021). GPT-3: What’s it good for? Natural Language Engineering, 27(1), 113–118. [Google Scholar] [CrossRef] [Scilit]
- da Silva, J. A. T. (2023). Is ChatGPT a valid author? Nurse Education in Practice, 68, 103600. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dowling, M., & Lucey, B. (2023). ChatGPT for (finance) research: The Bananarama conjecture. Finance Research Letters, 53, 103662. [Google Scholar] [CrossRef] [Scilit]
- Fahad, S. A., Salloum, S. A., & Shaalan, K. (2024). The role of ChatGPT in knowledge sharing and collaboration within digital workplaces: A systematic review. In A. Al-Marzouqi, S. A. Salloum, M. Al-Saidat, A. Aburayya, & B. Gupta (Eds.), Artificial intelligence in education: The power and dangers of ChatGPT in the classroom (Vol. 144). Studies in Big Data. Springer. [Google Scholar] [CrossRef] [Scilit]
- Fatouros, G., Soldatos, J., Kouroumali, K., Makridis, G., & Kyriazis, D. (2023). Transforming sentiment analysis in the financial domain with ChatGPT. arXiv, arXiv:2308.07935v1. [Google Scholar] [CrossRef] [Scilit]
- Feng, Z., Hu, G., & Li, B. (2024). Unleashing the power of ChatGPT in finance research: Opportunities and challenges. SSRN Electronic Journal, 4424979. [Google Scholar] [CrossRef] [Scilit]
- Frieder, S., Pinchetti, L., Griffiths, R., Salvatori, T., Lukasiewicz, T., Petersen, P., & Berner, J. (2023). Mathematical capabilities of ChatGPT. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, & S. Levine (Eds.), Advances in neural information processing systems (Vol. 36, pp. 27699–27744). MIT Press. [Google Scholar]
- Frye, B. L. (2022). Should using an AI text generator to produce academic writing be plagiarism? Fordham intellectual property. Media & Entertainment Law Journal, 33, 946. [Google Scholar]
- Gao, C. A., Howard, F. M., Markov, N. S., Dyer, E. C., Ramesh, S., Luo, Y., & Pearson, A. T. (2022). Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers. bioRxiv. [Google Scholar] [CrossRef] [Scilit]
- González-Padilla, D. A. (2022). Concerns about the potential risks of artificial intelligence in manuscript writing. Journal of Urology, 209(4), 682–683. [Google Scholar] [CrossRef] [Scilit]
- Hosseini, M., Rasmussen, L. M., & Resnik, D. B. (2024). Using AI to write scholarly publications. Accountability in Research, 31, 715–723. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jha, M., Qian, J., Weber, M., & Yang, B. (2023). ChatGPT and corporate policies. Working paper. SSRN, 4521096. [Google Scholar] [CrossRef] [Scilit]
- Khalil, M., & Er, E. (2023). Will ChatGPT get you caught? Rethinking of plagiarism detection. arXiv, arXiv:2302.04335. [Google Scholar]
- Khan, M. S., & Umer, H. (2024). ChatGPT in finance: Applications, challenges, and solutions. Heliyon, 10(2), e24890. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kim, A. G., Muhn, M., & Nikoleav, V. V. (2024). Financial statement analysis with large language models. Chicago booth research paper forthcoming, fama-miller working paper. Available online: https://ssrn.com/abstract=4835311 (accessed on 15 May 2024).
- Korinek, A. (2023). Language models and cognitive automation for economic research. NBER working paper series. Available online: https://www.nber.org/papers/w30957 (accessed on 15 May 2024).
- Li, X., Feng, H., Yang, H., & Huang, J. (2024). Can ChatGPT reduce human financial analysts’ optimistic biases? Economic and Political Studies, 12(1), 20–33. [Google Scholar] [CrossRef] [Scilit]
- Liebrenz, M., Schleifer, R., Buadze, A., Bhugra, D., & Smith, A. (2023). Generating scholarly content with ChatGPT: Ethical challenges for medical publishing. The Lancet Digital Health, 5(3), e105–e106. [Google Scholar] [CrossRef] [Scilit]
- Lin, Z. (2023). Why and how to embrace AI such as ChatGPT in your academic life. Royal Society Open Science, 10(8), 230658. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, L. X., Sun, Z., Xu, K., & Chen, C. (2024). AI-driven financial analysis: Exploring ChatGPT’s capabilities and challenges. International Journal of Financial Studies, 12(3), 60. [Google Scholar] [CrossRef] [Scilit]
- Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., & Tang, J. (2023). GPT understands, too. AI Open, 5, 208–215. [Google Scholar] [CrossRef] [Scilit]
- Lopez-Lira, A., & Tang, Y. (2024). Can ChatGPT forecast stock price movements? Return predictability and large language models. Working paper. SSRN, 4412788. [Google Scholar] [CrossRef] [Scilit]
- Lund, B. D., & Wang, T. (2023). Chatting about ChatGPT: How may AI and GPT impact academia and libraries? Library Hi Tech News, 40(3), 26–29. [Google Scholar] [CrossRef] [Scilit]
- Pelster, M., & Val, J. (2024). Can ChatGPT assist in picking stocks? Finance Research Letters, 59, 104786. [Google Scholar] [CrossRef] [Scilit]
- Rahman, M. M., & Watanobe, Y. (2023). ChatGPT for education and research: Opportunities, threats, and strategies. Applied Sciences, 13, 5783. [Google Scholar] [CrossRef] [Scilit]
- Rice, S., Crouse, S. R., Winter, S. R., & Rice, C. (2024). The advantages and limitations of using ChatGPT to enhance technological research. Technology in Society, 76, 102426. [Google Scholar] [CrossRef] [Scilit]
- Schlosky, M. T. T., Karadas, S., & Raskie, S. (2024). ChatGPT, help! I am in financial trouble. Journal of Risk and Financial Management, 17(6), 241. [Google Scholar] [CrossRef] [Scilit]
- Shen, Y., Heacock, L., Elias, J., Hentel, K. D., Reig, B., Shih, G., & Moy, L. (2023). ChatGPT and other large language models are double-edged swords. Radiology, 307(2), 230163. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shue, E., Liu, L., Li, B., Feng, Z., Li, X., & Hu, G. (2023). Empowering beginners in bioinformatics with ChatGPT. Quantitative Biology, 11(2), 105–108. [Google Scholar] [CrossRef] [Scilit]
- Smales, L. A. (2023). Classification of RBA monetary policy announcements using ChatGPT. Finance Research Letters, 58(C), 104514. [Google Scholar] [CrossRef] [Scilit]
- Stokel-Walker, C. (2023). ChatGPT listed as author on research papers: Many scientists disapprove. Nature, 613, 620–621. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Thorp, H. H. (2023). ChatGPT is fun, but not an author. Science, 379(6630), 313. [Google Scholar] [CrossRef] [Scilit]
- van Dis, E. A., Bollen, J., Zuidema, W., van Rooij, R., & Bockting, C. L. (2023). ChatGPT: Five priorities for research. Nature, 614(7947), 224–226. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, J., Ye, Q., Liu, L., Guo, N. L., & Hu, G. (2024). Scientific figures interpreted by ChatGPT: Strengths in plot recognition and limits in color perception. NPJ Precision Oncology, 8(1), 84. [Google Scholar] [CrossRef] [Scilit]
- Yang, Z., Li, L., Lin, K., Wang, J., Lin, C., Liu, Z., & Wang, L. (2023). The dawn of LMMs: Preliminary explorations with GPT-4V(ision). arXiv, arXiv:2309.17421. [Google Scholar]
- Yeo-Teh, N. S. L., & Tang, B. L. (2023). Letter to editor: NLP systems such as ChatGPT cannot be listed as an author because these cannot fulfill widely adopted authorship criteria. Accountability in Research, 31(7), 968–970. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yue, T., Au, D., Au, C. C., & Iu, K. Y. (2023). Democratizing financial knowledge with ChatGPT by OpenAI: Unleashing the power of technology. SSRN Electronic Journal, 4346152. [Google Scholar] [CrossRef] [Scilit]
- Zaremba, A., & Demir, E. (2023). ChatGPT: Unlocking the future of NLP in finance. Modern Finance, 1(1), 93–98. [Google Scholar] [CrossRef] [Scilit]
- Zhu, G., Fan, X., Hou, C., Zhong, T., Seow, P., Shen-Hsing, A. C., Rajalingam, P., Yew, L. K., & Poh, T. L. (2023). Embrace opportunities and face challenges: Using ChatGPT in undergraduate students’ collaborative interdisciplinary learning. arXiv, arXiv:2305.18616. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).




