Next Article in Journal
Sustainability-Oriented Policy–Terrain-Coupled Mixed-Fleet Routing for Scenario-Based Green Urban Freight Logistics
Previous Article in Journal
Multi-Objective Optimization of Relief-Well Dewatering for Canals Under High Groundwater Levels Using NSGA-II and Entropy-Weighted TOPSIS
Previous Article in Special Issue
Agentic AI Deployment Readiness and Responsible Value Realization in Sustainable Banking
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Textual Sentiment and Financial Market Dynamics: Econometric Evidence from Eastern Europe’s Green and Digital Transition

by
Cristian-Valentin Hapenciuc
,
Daniela Mihaela Neamțu
*,
Teodora Cajvan
*,
Camelia Băeșu
and
Otilia-Maria Bordeianu
Department of Management, Business Administration and Tourism, Faculty of Economics, Administration and Business, “Ștefan cel Mare” University of Suceava, 720229 Suceava, Romania
*
Authors to whom correspondence should be addressed.
Sustainability 2026, 18(17), 9177; https://doi.org/10.3390/su18179177
Submission received: 7 July 2026 / Revised: 29 August 2026 / Accepted: 3 September 2026 / Published: 7 September 2026

Abstract

Conventional macroeconomic indicators are released with a time lag and may therefore provide limited information about rapidly changing market expectations, creating a need for complementary high-frequency indicators capable of capturing the informational content of financial narratives. This study examines whether sentiment extracted from unstructured financial text contains incremental predictive information for short-horizon market dynamics in the context of Eastern Europe’s green and digital transition. Using financial news and discourse collected between May 2025 and May 2026, a FinBERT-based Daily Sentiment Index (DSI) is constructed for selected technology, energy, and sustainability-related narratives and linked to market indicators and representative assets, including the DAX, UiPath, and OMV Petrom. The empirical strategy combines Granger predictability tests, vector autoregression (VAR), GARCH(1,1) models with sentiment effects, lead–lag analysis, Bai–Perron structural-break tests, and Markov-switching specifications to assess predictive temporal relationships, volatility dynamics, and time variation in sentiment–market interactions. Forecasting utility is further evaluated through an expanding-window out-of-sample exercise in which a sentiment-augmented VAR is compared with a benchmark autoregressive specification using one-day-ahead DAX returns, with relative predictive accuracy evaluated via the Diebold–Mariano test. The results indicate that lagged sentiment contains statistically significant predictive information for subsequent market returns and that incorporating the DSI significantly improves out-of-sample forecast accuracy relative to the benchmark specification. The evidence also reveals heterogeneous and time-varying responses across the selected cases: technology-related assets exhibit comparatively differentiated sensitivity to macro-sentiment conditions, whereas the energy case displays stronger exposure to transition-related narratives. Structural-break and regime-dependent estimates further suggest instability in the relationship between ESG sentiment and market performance over the sample period, but do not establish a permanent structural transformation in investor preferences. The study contributes to the narrative-economics and computational-finance literature by providing applied evidence that domain-specific textual sentiment contains incremental short-horizon information beyond conventional market dynamics and can complement lagged macroeconomic indicators in financial forecasting and risk-monitoring applications.

1. Introduction

Financial markets process information from multiple sources, including macroeconomic releases, corporate disclosures, regulatory announcements, financial news, and broader public narratives. Traditional asset-pricing approaches emphasize the incorporation of fundamental information into market prices, whereas behavioral finance and narrative economics additionally recognize that the interpretation, diffusion, and framing of information can shape investor expectations and risk perceptions. These perspectives are complementary rather than mutually exclusive: market prices may reflect both changes in underlying fundamentals and changes in expectations generated by the arrival and interpretation of new information. Within this setting, computational analysis of financial text provides a means of transforming continuously generated qualitative information into quantitative sentiment measures that can be evaluated jointly with observed market variables.
This informational mechanism is particularly relevant to environmental, social, and governance (ESG) investing. ESG-related information can affect expectations concerning future cash flows, regulatory compliance costs, financing conditions, transition risks, reputational exposure, and long-term competitive positioning. Importantly, the market interpretation of such information need not be homogeneous across sectors or stable over time. A decarbonization policy, for example, may imply additional transition costs for carbon-intensive firms while simultaneously improving expected opportunities for firms providing digital or low-carbon technologies. Changes in the tone of ESG and transition-related narratives may therefore alter investor expectations and risk assessments, potentially generating portfolio reallocation, return adjustments, and changes in conditional volatility. Textual sentiment and financial volatility are consequently linked through an information-processing channel: when new narratives modify the distribution of market expectations or increase disagreement regarding future economic conditions, their informational content may subsequently be reflected in market returns and volatility. Whether textual sentiment contains statistically significant information about these subsequent market dynamics, however, remains an empirical question. This question is particularly relevant in the context of Europe’s simultaneous green and digital transformation. The so-called twin transition exposes firms and investors to overlapping sources of uncertainty arising from decarbonization policies, energy-security concerns, technological change, regulatory intervention, financing conditions, and geopolitical developments. Romania provides a useful setting for examining these interactions because the selected market entities exhibit materially different exposures to energy transition, banking transformation, and digitalization. The empirical design therefore considers representative entities including OMV Petrom and Hidroelectrica in energy, Banca Transilvania in banking, and UiPath as a technology-related case, while the DAX is employed as a broader European industrial-market proxy. The purpose of this multi-entity design is not to infer sector-wide behavioral laws from individual firms, but to examine whether sentiment–market relationships differ across cases exposed to distinct dimensions of the green and digital transition.
A central empirical problem concerns the temporal mismatch between rapidly evolving market information and the availability of conventional macroeconomic indicators. Measures such as gross domestic product, industrial production, and inflation are released according to predetermined statistical calendars and therefore become observable only after part of the economic activity they describe has already occurred. Some indicators are also subsequently revised. In this study, this delay is referred to as reporting latency: the interval between the occurrence of economically relevant developments and their representation in official macroeconomic statistics. Reporting latency does not diminish the fundamental importance of official statistics; rather, it limits their ability to reflect changes in market expectations at very short horizons. Financial news is generated continuously and can incorporate new economic, regulatory, geopolitical, and firm-specific information as events unfold. This temporal asymmetry motivates the investigation of textual sentiment as a complementary high-frequency information source that may contain predictive information before corresponding developments are fully reflected in lower-frequency macroeconomic indicators.
Transforming this textual information into a reliable quantitative measure introduces a second methodological problem. Earlier sentiment approaches frequently relied on general-purpose dictionaries, word counts, or bag-of-words representations. Although useful as benchmarks, these approaches can misclassify financial language when polarity depends on context, domain-specific terminology, syntactic structure, or negation. Such semantic measurement error is consequential for subsequent econometric analysis. If the constructed sentiment variable systematically misrepresents the informational content of financial text, estimated sentiment–market relationships may be attenuated or distorted, and apparent lead–lag patterns may reflect measurement noise rather than economically meaningful information. Domain-specific Transformer architectures such as FinBERT address part of this limitation by generating context-sensitive sentiment classifications from financial language.
Nevertheless, superior text classification alone does not demonstrate that sentiment has economic or predictive value. A sentiment measure may classify financial language accurately while providing no incremental information for subsequent market outcomes. The methodological problem is therefore two-stage. First, unstructured financial narratives must be converted into a sufficiently reliable quantitative sentiment measure. Second, the resulting measure must be subjected to formal statistical and econometric evaluation to determine whether it contains information beyond contemporaneous descriptive association. The present study addresses the first stage by constructing a threshold-calibrated Daily Sentiment Index (DSI) from FinBERT sentiment probabilities and validating the underlying classifier against manually annotated financial headlines and a Loughran–McDonald lexicon benchmark. The second stage evaluates the temporal and predictive properties of the resulting DSI through complementary time-series and forecasting procedures.
Against this background, the primary objective of this study is to construct and evaluate a domain-specific textual sentiment measure and determine whether it contains incremental predictive information for short-horizon financial market dynamics in the context of Eastern Europe’s green and digital transition. More specifically, the study evaluates whether the DSI precedes subsequent market movements, whether sentiment is associated with conditional volatility, whether these relationships differ across the selected cases and evolve over time, and whether incorporating sentiment improves short-horizon forecasting performance relative to a benchmark specification that excludes the DSI. The analysis is explicitly predictive rather than structurally causal: temporal precedence or Granger predictability is interpreted as evidence of incremental forecasting information and not as identification of an exogenous causal effect of sentiment on asset prices. Accordingly, the study addresses the following research questions:
RQ1. 
Does the FinBERT-derived Daily Sentiment Index contain statistically significant predictive information for subsequent market movements beyond information contained in past market dynamics?
RQ2. 
Does textual sentiment provide statistically significant information about conditional market volatility?
RQ3. 
Are sentiment–market relationships heterogeneous across the selected technology, energy, and financial-market cases and time-varying over the sample period?
RQ4. 
Does incorporating the Daily Sentiment Index improve short-horizon out-of-sample forecasting accuracy relative to a benchmark specification that excludes textual sentiment?
The contribution of the study is primarily methodological and empirical. First, it develops a reproducible pipeline that converts financial-news headlines into a threshold-calibrated Daily Sentiment Index using a finance-specific Transformer architecture. The reliability of the sentiment measurement stage is assessed using manually annotated headlines, intercoder agreement, classification metrics, and comparison with the Loughran–McDonald financial lexicon. Second, rather than inferring predictive value from descriptive correlations, the study integrates the DSI into a complementary econometric framework comprising lead–lag analysis, Granger predictability tests interpreted in their predictive sense, vector autoregression, GARCH(1,1) specifications, and time-varying analyses designed to examine the stability and heterogeneity of sentiment–market relationships. Predictive usefulness is further evaluated through out-of-sample forecast comparison between a sentiment-augmented specification and a benchmark model excluding the DSI. Third, the study provides context-specific evidence from selected Eastern European transition-related cases, allowing the informational content of financial narratives to be examined across assets with different exposures to digitalization, energy transition, and broader European industrial conditions. Rather than proposing a new behavioral or asset-pricing theory, the study contributes empirical evidence on when and to what extent domain-specific textual sentiment contains incremental information relevant to short-horizon financial monitoring and forecasting.
The remainder of the paper is organized as follows. Section 2 develops the theoretical background and hypotheses. Section 3 describes the data, sentiment-index construction, validation procedures, and econometric methodology. Section 4 presents and discusses the empirical findings. Section 5 concludes with the implications, limitations, and directions for future research.

2. Literature Review

2.1. The Evolution of Financial Theory and Systemic Resilience

The theoretical foundations of modern financial economics are closely associated with the Efficient Market Hypothesis (EMH), according to which financial prices incorporate available information and persistent abnormal risk-adjusted returns should therefore be difficult to obtain systematically [1]. This proposition is closely related to the random-walk argument: if new information arrives unpredictably and is rapidly incorporated into market prices, successive price changes should also be difficult to predict systematically. The EMH distinguishes three forms of informational efficiency according to the information set reflected in prices [1]. Weak-form efficiency implies that current prices incorporate information contained in historical prices and trading data; semi-strong-form efficiency extends this information set to all publicly available information; and strong-form efficiency assumes that prices reflect all relevant information, including private information. Within this framework, competition among market participants and the rapid incorporation of new information constitute central mechanisms supporting informational efficiency.
A central debate surrounding the EMH concerns the rationality assumptions underlying investor decision-making. Traditional financial models commonly approximate investors as rational agents who process available information consistently and form expectations accordingly. Behavioral Finance (BF) challenges this representation by demonstrating that decision-making under risk and uncertainty can systematically depart from the predictions of conventional expected-utility models [2]. Prospect Theory demonstrates that individuals evaluate gains and losses asymmetrically and that loss aversion can materially influence financial decisions [2]. More broadly, bounded rationality, herding, overreaction, underreaction, and other behavioral mechanisms provide potential channels through which investor sentiment and narratives can temporarily influence market prices. Behavioral Finance therefore challenges the assumption that information is processed uniformly, rationally, and instantaneously by all market participants, while not necessarily implying that financial markets are permanently inefficient.
The apparent tension between informational efficiency and behavioral deviations is reconciled more explicitly by the Adaptive Markets Hypothesis (AMH) [3]. Rather than treating market efficiency and behavioral biases as mutually exclusive states, the AMH conceptualizes financial markets as evolutionary systems in which the degree of efficiency can vary with competition, environmental conditions, institutional structures, technological change, and investor adaptation [3]. Under this framework, behavior that appears suboptimal in one environment may represent an adaptive response to another, while profitable patterns may emerge and subsequently disappear as market participants learn and competitive conditions change. Consequently, periods characterized by relatively efficient information processing can alternate with periods in which behavioral biases, uncertainty, changing investor attention, or market narratives exert a stronger influence on observed market outcomes [3].
This adaptive interpretation provides the principal theoretical lens for the present study. If informational efficiency varies across time and market conditions, the relationship between textual sentiment and financial outcomes should likewise not be assumed to remain constant. Financial narratives may contain limited incremental information when relevant information is rapidly incorporated into prices, while their predictive content may become more pronounced during periods of heightened uncertainty, transition risk, geopolitical stress, or changing investor attention. The AMH therefore provides a coherent theoretical basis for examining time-varying sentiment–market relationships, conditional volatility, structural instability, and heterogeneous responses across selected assets. Behavioral Finance complements this perspective by providing the cognitive mechanisms through which narratives and sentiment can influence expectations and perceptions of risk.
This theoretical perspective is particularly relevant to systemic resilience and sustainability-related financial behavior. Resilience can be understood as the capacity of a system to absorb shocks, reorganize, and maintain essential functions under changing environmental conditions [4,5]. At the investor level, behavioral responses to uncertainty can affect this adaptive process. Loss aversion, for example, implies asymmetric responses to perceived gains and losses [2], while long-horizon environmental and transition risks may be evaluated differently from immediate compliance costs or short-term financial losses [6]. At the market level, the aggregation of heterogeneous investor responses can generate changing patterns of risk perception, asset repricing, and volatility rather than a single, time-invariant reaction to sustainability-related information.
The empirical ESG literature further supports treating these relationships as heterogeneous rather than predetermined. A broad synthesis covering more than 2000 empirical studies reports an overall tendency toward a non-negative relationship between ESG characteristics and corporate financial performance, while simultaneously documenting substantial heterogeneity across studies and empirical settings [7]. Other asset-pricing evidence indicates that realized green returns may reflect changes in environmental preferences and climate concerns rather than a stable and universally observable ESG premium [8]. These findings suggest that the financial relevance of sustainability information can vary according to market conditions, investor preferences, asset characteristics, and time. Such evidence is consistent with the adaptive perspective and motivates an empirical framework capable of evaluating whether ESG-related textual sentiment contains incremental information for financial returns and volatility and whether these relationships remain stable throughout the observation period.

2.2. ESG Narratives and the Role of “Soft Data” in Asset Value Formation

Financial asset valuation reflects not only firm fundamentals and macroeconomic conditions but also the way new information is interpreted and incorporated into investor expectations. Narrative economics provides a complementary perspective by emphasizing that informational flows and collective narratives can influence market dynamics alongside fundamental information [9]. Narratives concerning technological transformation, energy transition, geopolitical uncertainty, or corporate sustainability can diffuse rapidly across financial media and shape expectations before related developments are fully observable in official economic statistics. This mechanism is particularly relevant to ESG information, which has evolved from a predominantly ethical consideration into a financially relevant source of information concerning regulatory exposure, transition risk, financing conditions, reputation, and long-term corporate value [10].
The empirical literature nevertheless provides heterogeneous evidence regarding the financial implications of sustainability information. Evidence synthesized across a large body of empirical research indicates that ESG characteristics are frequently associated with non-negative corporate financial performance, although the magnitude and direction of the relationship differ considerably across empirical settings [7]. In contrast to approaches centered predominantly on firm-level ESG characteristics, asset-pricing evidence indicates that green-asset returns may also reflect changes in environmental preferences and climate concerns, suggesting that observed ESG-related performance cannot necessarily be interpreted as a constant premium [8]. Research using climate-related news further demonstrates that textual information concerning climate change contains financially relevant information and may contribute to dynamic portfolio-hedging strategies [11]. Collectively, these findings indicate that sustainability-related information is relevant to financial markets, while simultaneously demonstrating substantial differences in the information measured, methodological design, time horizon, and financial outcomes considered [7,8,11]. This heterogeneity leaves open the question of whether high-frequency, domain-specific textual sentiment provides incremental information for short-horizon market dynamics and whether its informational content varies over time and across transition-exposed assets.
The quantification of such narratives has increasingly relied on “soft data”, understood in this study as unstructured textual information extracted from financial news and related informational sources. Unlike conventional “hard data”, including GDP, inflation, and industrial-production indicators that are released according to predetermined statistical calendars, textual information is generated continuously as economic and financial events unfold. Consequently, soft data can provide higher-frequency information about changing expectations between official statistical releases. This characteristic creates the potential for constructing high-frequency indicators capable of monitoring changes in market sentiment associated with climate risks, regulatory developments, geopolitical shocks, and technological transformation.
Early empirical evidence demonstrates that financial-media content contains market-relevant information, with pessimistic narratives being associated with subsequent price pressure and changes in market activity [12]. This finding established an important empirical connection between textual information and financial outcomes. However, association between news content and market variables does not by itself demonstrate that textual sentiment provides incremental forecasting information. The relevant empirical question is whether sentiment extracted from continuously generated financial narratives contains information about subsequent market dynamics beyond that already embedded in historical market variables.
The informational relevance of textual data is particularly important in sectors undergoing structural transformation, such as energy and technology. Climate-related news can alter investor assessments of regulatory exposure, transition costs, future cash flows, and hedging requirements [11]. Similarly, technological narratives can affect expectations regarding productivity, competitive positioning, and long-term growth. The same narrative shock may therefore be interpreted differently across assets depending on their sectoral exposure and underlying economic characteristics. This heterogeneity provides a direct link between textual sentiment, ESG transition narratives, and the adaptive-market perspective developed in Section 2.1.
From this perspective, textual sentiment should not be interpreted as a substitute for fundamental financial or macroeconomic information. Rather, it represents a complementary high-frequency information source whose empirical relevance must be formally evaluated. The principal analytical challenge is therefore to determine whether unstructured narratives contain measurable information associated with subsequent market returns and volatility and whether such relationships are stable across time and the selected transition-related cases [9,11].

2.3. Computational Monitoring: From Lexical Counting to FinBERT Architectures

The methodological approach to quantifying qualitative financial information has evolved from basic word-frequency methods toward increasingly context-sensitive language models. Early automated approaches frequently relied on bag-of-words representations and sentiment dictionaries, in which document polarity was inferred from the frequency of predefined positive and negative terms. Foundational empirical research using media content demonstrated that pessimistic financial narratives are associated with downward price pressure and subsequent market activity, establishing the economic relevance of textual information [12]. However, dictionary-based approaches are constrained by their limited ability to account for financial context, syntactic structure, negation, and domain-specific terminology.
The importance of this limitation is demonstrated by evidence showing that general-purpose sentiment dictionaries can systematically misclassify terminology commonly used in financial documents [13]. Words that carry negative connotations in ordinary language may have neutral or context-dependent meanings in financial communication. Consequently, lexical classification errors can introduce measurement error into the sentiment variable used in subsequent financial analysis. The literature therefore establishes both the economic relevance of textual information and a fundamental measurement problem: market-relevant sentiment cannot necessarily be captured reliably through context-insensitive word counts [12,13].
This methodological limitation motivated the transition toward contextual language representations. Transformer architectures employ attention mechanisms that evaluate relationships among words within their surrounding linguistic context, allowing sentiment classification to account for syntactic dependencies and contextual meaning. BERT-based models extend this capability by deriving representations from bidirectional linguistic context, while finance-specific architectures such as FinBERT adapt contextual language modeling to the terminology and semantic structures characteristic of financial communication. This development is particularly relevant where polarity depends on negation, conditional statements, or domain-specific meanings that cannot be adequately represented through static lexical scores [13].
The methodological progression from lexical analysis toward contextual NLP therefore addresses an important measurement limitation, but improved sentiment classification does not itself establish financial or forecasting relevance. A model may classify financial language with greater accuracy while the resulting sentiment measure contains little or no incremental information about subsequent market behavior. Consequently, classification performance and economic predictive performance constitute distinct empirical questions. This distinction is central to the present study because the construction of a context-sensitive sentiment indicator represents only the first stage of the empirical analysis; the resulting measure must subsequently be evaluated within an econometric framework.
Recent developments in computational monitoring provide a framework through which NLP-derived indicators can support high-frequency financial risk assessment [14]. By transforming unstructured text into quantitative variables, such approaches permit textual information to be integrated with conventional time-series data. However, the use of increasingly complex machine-learning architectures also raises issues concerning transparency, interpretability, computational filtering, and reproducibility [15]. These considerations are particularly important when model-generated sentiment measures are subsequently used as explanatory or predictive variables in financial applications. Methodological transparency therefore requires explicit reporting of corpus construction, preprocessing, classification procedures, model validation, sentiment aggregation, temporal alignment, and subsequent econometric testing.
Within sustainability-related applications, contextual sentiment models can provide high-frequency measures of how ESG and transition narratives evolve over time [14,16]. Their principal analytical advantage is the ability to distinguish contextual meaning that may be lost under dictionary-based approaches. Nevertheless, such models should not be interpreted automatically as early-warning systems or as evidence that narratives causally determine financial outcomes. Their economic relevance depends on whether the extracted sentiment measure demonstrates statistically meaningful temporal and predictive relationships with observed market variables.
Taken together, the reviewed literature reveals three interrelated research gaps. First, empirical evidence concerning the financial implications of ESG information remains heterogeneous, indicating that sentiment–market relationships may depend on market conditions, sectoral exposure, investor preferences, and time [7,8,11]. Second, although financial-text research establishes that news sentiment contains market-relevant information, conventional dictionary-based measures remain vulnerable to domain-specific semantic misclassification [12,13]. Third, improvements in NLP-based sentiment classification do not by themselves establish incremental economic or forecasting value; the resulting sentiment measure must be evaluated within a formal time-series framework and against an appropriate predictive benchmark.
The present study addresses these gaps by constructing a context-sensitive FinBERT-derived Daily Sentiment Index (DSI) and subsequently evaluating its relationship with selected financial-market outcomes through complementary econometric procedures. The empirical framework distinguishes contemporaneous association from predictive temporal dependence and further examines conditional volatility and time variation in sentiment–market relationships. In addition, the predictive usefulness of the DSI is evaluated against a benchmark specification that excludes textual sentiment. Accordingly, the contribution is primarily methodological and empirical: the study integrates contextual NLP-based sentiment measurement with formal econometric and out-of-sample predictive evaluation in selected transition-exposed cases, without claiming structural causal identification or proposing a new theory of financial behavior (Table 1).

3. Research Methodology

3.1. Research Hypotheses and the Systemic Resilience Framework

To evaluate the efficacy of algorithm-assisted risk analysis and its impact on financial stability, this research formalizes three core hypotheses designed to test the responsiveness of “soft data” (narrative sentiment) relative to “hard data” (market prices and volatility indices). These hypotheses focus specifically on the intersection of the “Twin Transition” and systemic resilience. The table below provides a developed view of the core hypotheses, the supporting academic literature, and the specific methodology applied in the study (Table 2):
To rigorously evaluate H1 and determine if the DSI functions as an early-warning indicator rather than merely a coincidental metric, the methodology must extend beyond visual mirroring effects. Specifically, we test the ‘early-warning’ capacity by employing Granger predictability tests to evaluate incremental predictive temporal precedence and lead-lag analysis to determine if the DSI statistically precedes VIX and asset price movements. Additionally, a Vector Autoregression (VAR) model is utilized to examine dynamic predictive interactions, while forecast comparisons (evaluating a DSI-augmented model against a baseline model without DSI) via out-of-sample prediction tests ensure the robustness of our findings. It is a fundamental methodological premise of this study that the argument of an early-warning mechanism can be empirically sustained only if the DSI exhibits statistically significant predictive power across these specific inferential tests.
By establishing these hypotheses within a longitudinal 12-month framework (May 2025–May 2026), the methodology shifts from a static analysis to an exploratory evaluation of time-varying market behavior. The study validates these propositions to determine the extent to which algorithmic sentiment analysis can serve as a sustainable tool for modern risk management.
While computational sentiment extraction and descriptive visualizations provide essential context, they serve strictly as a preliminary exploratory analysis. To move beyond correlational associations and formally evaluate complex predictive and time-varying statistical relationships, this study integrates the NLP outputs into a robust, hybrid methodological framework: exploratory analysis statistical testing econometric modeling machine learning validation. This sequential approach ensures that descriptive trends are subjected to rigorous empirical validation before predictive interpretations are formulated.

3.2. Research Design: The Data Science Paradigm and Textual Epistemology

The methodological framework of this research employs a quantitative design grounded in the Data Science paradigm, representing a strategic departure from the rigid constraints of traditional econometrics. By shifting focus from retrospective financial ratios to unstructured “soft data”, specifically financial news feeds, this study treats informational narratives as a high-frequency information source associated with market volatility and systemic resilience. This approach recognizes that in a digitalized environment, the narratives surrounding assets may contain information about price trajectories before official economic indicators are made public, effectively transforming language into a quantifiable economic force.
The epistemological basis of this study is the “Text as Data” framework, which posits that unstructured textual information contains high-dimensional, latent data capable of predicting market behavior with greater immediacy than lagged official reports. This methodology aligns with the computational monitoring mechanisms of statistical modeling, which prioritizes the extraction of complex, non-linear patterns from real-world data over simple theory validation [20]. Furthermore, these Big Data methodologies provide a sophisticated toolkit for identifying complex and potentially nonlinear statistical relationships that standard linear regression models may overlook [21].
To operationalize this framework, the study utilizes Computational Content Analysis powered by Natural Language Processing (NLP), evolving beyond the limitations of traditional “Bag-of-Words” methods. Because dictionary-based approaches often ignore linguistic context, failing to distinguish, for instance, between a “liability” and its reduction, this methodology employs the FinBERT model. This deep learning architecture uses an attention-based mechanism to decode the syntax of sustainability, transforming raw text into a Daily Sentiment Index. Ultimately, this system functions as a real-time monitor of investor psychology, serving as a candidate high-frequency early-warning indicator whose predictive usefulness is evaluated against traditional “hard data” metrics.
To ensure complete empirical replicability, we explicitly recognize that methodological transparency requires the comprehensive reporting of the corpus construction, the total number of observations, the technical implementation of the FinBERT architecture, the sentiment classification procedures, the data purification workflow, and the precise mathematical parameters governing the model.
To ensure strict methodological clarity and bridge the theoretical framework with our empirical design, we explicitly distinguish between measurable constructs and interpretive heuristics. The Twin Transition is operationalized mathematically through our purposive sample selection (UiPath representing digitalization; OMV Petrom representing the green transition). Systemic resilience and sectoral decoupling are quantitatively evaluated via GARCH(1,1) conditional volatility associations and cross-asset correlation matrices. The Adaptive Market Hypothesis (AMH) is empirically tested through time-varying interactions, specifically utilizing Markov-switching regime models and dynamic rolling correlations to capture shifting investor behaviors. Conversely, concepts such as organizational adaptive capacity, cognitive desensitization, and computational monitoring are not independent mathematical variables; rather, they are explicitly employed throughout the discussion as heuristic interpretive frameworks to theoretically contextualize the observed econometric results.

3.3. Sample Justification Through the Sustainability and Resilience Lens

The research universe is defined by a purposive sampling strategy designed to analyze context-specific evidence of market behavior during the “Twin Transition”, the concurrent shift toward green and digital economies. This study focuses on strategic proxies serving as an illustrative case selection that captures specific sectoral dynamics: energy security, digital inclusion, and operational efficiency. The methodology employs a 12-month longitudinal window spanning from May 2025 to May 2026, creating a high-resolution environment to test how narrative-associated volatility interacts with sustainability-oriented assets during a period of industrial friction and geopolitical instability.
It is crucial to clarify that this selection functions as a multi-sectoral case study rather than a representative sample of the entire European market. These specific entities were chosen due to their distinct sectoral relevance (technology, energy, finance), high data availability, pronounced exposure to ESG and digitalization narratives, and their significant roles in local and regional markets. However, we explicitly acknowledge that these entities cannot be treated as a universal proxy for broader European dynamics. The sample’s small size, geographic concentration, and sectoral heterogeneity inherently limit external validity. Consequently, the results offer sector-bounded, exploratory insights rather than broadly generalizable claims across the continent.
The selection of the Romanian market and its regional proxies is justified by its empirical characteristics, which provide a particularly informative setting for narrative economics. Compared with mature, highly liquid Western markets, this specific Eastern European ecosystem exhibits structural rigidities, shallower liquidity pools, and susceptibility to governmental regulatory interventions. Furthermore, its direct geographical proximity to the ongoing geopolitical conflict in Ukraine, combined with a distinct national energy mix ranging from state-backed renewables to major fossil fuel players, creates an environment in which narrative analysis is particularly relevant. Information concerning energy security or digital transitions may be associated with market responses of different timing and magnitude than in dominant global economies [22]. By isolating this specific market, the study examines high-frequency sentiment responses and whether localized market dynamics are associated with relative divergence from broader regional patterns (Table 3).
It is imperative to state that a 12-month timeframe cannot capture a full economic cycle, given that macroeconomic cycles typically unfold over 3 to 10 years and encompass structurally distinct phases, including expansion, peak, contraction, and recovery. Consequently, any attempt to extrapolate these results to complete economic cycles is strictly avoided, and our findings must be interpreted as context-specific patterns representing short-term dynamics. Nevertheless, this 12-month longitudinal window is highly adequate for the specific objectives of this study, which does not aim to model long-term macroeconomic evolution. Instead, this specific timeframe is explicitly tailored to capture high-frequency sentiment responses, real-time market reactions to volatile ESG, geopolitical, and technological narratives, short-term predictive relationships involving conditional volatility, and short-horizon market responses to newly available information.
The specific 12-month longitudinal window, concluding in May 2026, was deliberately selected because it contains a relatively high concentration of relevant informational shocks. This period captures narrative variation concerning the European energy transition, the deployment of critical digital infrastructure, and shifting geopolitical realities. These informational flows provide sufficient variation to examine the proposed short-horizon relationships and systemic-resilience hypotheses, while the limited temporal window remains an acknowledged constraint. Extending the observation period therefore constitutes a valuable avenue for future research and would allow the stability of the present findings to be assessed across a broader range of market conditions.
The energy sector’s transition is examined through the contrast between Hidroelectrica and OMV Petrom. As a pure-play renewable provider, Hidroelectrica serves as the regional benchmark for decarbonization, while OMV Petrom reflects the fossil-fuel sector’s exposure to climate regulations and geopolitical supply risks. This comparative approach aligns with established transition frameworks suggesting that the speed of energy shifts is determined by the friction between existing regimes and green innovations [23,24]. This selection allows the research to determine whether market sentiment prioritizes long-term renewable stability or remains tied to the volatility of fossil-fuel supply chains.
Social sustainability and financial democratization are analyzed through the digital transformation of the Romanian domestic economy, represented by Transilvania Bank (BT). Its strategic implementation of the EU ID wallet is treated as a core mechanism for digital inclusion and social resilience. Such digital transitions are considered essential for sustainable recovery, as they facilitate financial access for underbanked populations while reducing the carbon footprint of physical infrastructure [25]. Consequently, BT serves as a key control variable representing the intersection of digital efficiency and social equity in a volatile climate.
Finally, UiPath is selected to represent the “growth” factor within the digital resilience paradigm, testing how digital-first assets perform against macroeconomic headwinds. Economically, UiPath embodies “operational sustainability,” utilizing automation and AI to optimize resource allocation [18]. By decoupling corporate value from physical resource constraints and traditional industrial cycles, this selection measures whether innovation-driven narratives can temporarily insulate specific assets from the systemic negativity surrounding the regional ‘old economy.’ To effectively measure this divergence, the German DAX index was explicitly selected as the baseline macroeconomic proxy. Because the DAX is heavily weighted toward traditional manufacturing, automotive, and heavy industries, it provides a highly accurate and structural benchmark for the European ‘old economy’ industrial cycle, serving as a necessary counterweight to evaluate the relative stability of digital and renewable case studies. However, as an inherent limitation, it must be explicitly noted that the DAX strictly represents a specific heavy-industry structure and should not be interpreted as a universal proxy for the entirety of the European macroeconomy.

3.4. Data Collection and Context-Aware Purification (Data Mining)

The empirical integrity of computational analysis depends entirely on the transparency, quality, and rigorous purification of its input data. To ensure strict methodological transparency and replicability, the data retrieval procedure was formalized through an automated programmatic pipeline. Specifically, the primary textual dataset was extracted via The Guardian Open Platform API and Yahoo Finance API, targeting a strictly bounded 12-month longitudinal period from May 2025 to May 2026. The selection of The Guardian as the primary source for unstructured narrative discourse is fundamentally justified by its highly stable API, exceptionally clean metadata architecture, and consistent, high-density editorial coverage of European ESG transitions and macroeconomic dynamics. This source provides a coherent and well-structured corpus highly suitable for preliminary exploratory analysis. However, we explicitly caution that its specific editorial stance and narrative style should not be interpreted as fully representative of all diverse European media ecosystems, which is why rigorous automated deduplication and contextual filtering procedures were strictly applied prior to algorithmic ingestion to minimize structural noise.
To ensure absolute methodological integrity and prevent any form of look-ahead bias, the temporal boundaries of the dataset were strictly defined and enforced prior to the commencement of the analysis. The automated data extraction protocol officially concluded on May 12, 2026. We explicitly confirm that all textual and financial data utilized in this study were publicly available and fully accessible at the exact time of analysis; no subsequent data revisions or late-published indicators were retroactively included. To systematically eliminate look-ahead bias, our computational pipeline incorporated rigorous timestamp verification. This programmatic constraint ensured that no market metrics, revised macroeconomic reports, or media articles published after the established cutoff date were ingested into the model. Consequently, the algorithmic evaluations were executed relying strictly on the chronological flow of information sequentially available to real-world market participants during the designated timeframe, thereby preserving the authenticity of the associative and predictive frameworks.
Following extraction, the raw dataset underwent a rigorous, context-aware purification process. We applied automated filtering algorithms to eliminate syndicated media duplicates, non-English articles, and items with insufficient informational density (specifically, headlines containing fewer than 5 words, ensuring the text provided sufficient syntactic structure for the Transformer’s attention mechanism), reducing the corpus to a final, highly representative analytical sample of 14,210 unique headline observations. The distributional breakdown of this final dataset was strictly categorized to ensure balanced sectoral representation: 4150 articles pertained to the technology and AI sector (UiPath), 5320 articles focused on the energy transition and geopolitical dynamics (OMV Petrom and Hidroelectrica), 2840 articles covered banking and digital inclusion (Transilvania Bank), while the remaining 1900 articles captured broad European macroeconomic sentiment associated with the DAX index. Data retrieval was executed using strict entity-matching queries (e.g., \”UiPath\” AND (\”AI\” OR \”earnings\” OR \”tech\”); \”OMV Petrom\” AND (\”energy transition\” OR \”geopolitics\” OR \”supply\”)). To ensure precise temporal alignment with the financial markets, article publication timestamps (UTC) were strictly synchronized with local market closing times (16:00 GMT+2). Articles published post-market close, or during weekends and non-trading holidays, were algorithmically rolled over to the subsequent trading day ( t + 1 ) to accurately reflect when the information could realistically be priced into the assets.
Regarding textual preprocessing, it is critical to distinguish between the primary Transformer model and the baseline dictionary. For the FinBERT architecture, the text (specifically the article headlines, to capture high-density sentiment without the noise of full-text body paragraphs) was preserved in its raw, unlemmatized state, retaining stop-words and natural punctuation. This is imperative because Transformer attention mechanisms rely on complete syntactic structures to resolve contextual nuances like negations. Conversely, traditional NLP preprocessing (including lowercasing, stop-word removal, and lemmatization) was applied exclusively to the baseline comparison dataset used for the Loughran-McDonald dictionary test [26].
The most critical phase involved context-aware data cleaning to address the “semantic noise” inherent in unstructured web data. Exploratory analysis identified significant lexical ambiguities; for instance, the acronym “BVB” refers to the Bucharest Stock Exchange in a local context but frequently denotes a football club in global news. Without a context-aware filtering protocol, sports-related headlines would introduce “sentiment contamination,” erroneously signaling negative volatility for the Romanian capital market.
To rectify this, a protocol was implemented to systematically remove entries containing non-financial keywords. Furthermore, the purification process addressed the removal of duplicates resulting from news syndication. Eliminating these redundancies was essential to prevent the over-weighting of specific media events during the calculation of the Daily Sentiment Index [27].

3.5. Algorithmic Architecture: FinBERT and the Construction of the Daily Sentiment Index (DSI)

The operationalization of the “Text as Data” framework is achieved through a specialized computational pipeline that moves beyond basic word counting toward a contextual understanding of the “syntax of finance.” Central to this process is FinBERT, a Large Language Model (LLM) based on the Bidirectional Encoder Representations from Transformers (BERT) architecture. FinBERT is utilized instead of generic models due to the unique nature of financial terminology; in a market context, words that typically carry negative connotations in general prose are often neutral or even positive. Standard dictionaries frequently misclassify terms like “liability” or “tax,” whereas FinBERT is trained to accurately interpret their roles within corporate filings and earnings transcripts [13].
To ensure rigorous methodological transparency and exact replication, we explicitly define our computational framework as a sequential pipeline: unstructured text input preprocessing   tokenization model inference aggregation DSI generation. Prior to model ingestion, the unstructured textual data underwent a minimal preprocessing pipeline strictly limited to UTF-8 encoding normalization and the removal of residual HTML tags, explicitly preserving natural punctuation, stop-words, and raw syntactic structures. [28]. This unlemmatized text was subsequently tokenized utilizing the native BERT WordPiece tokenizer, strictly configured with a max_length parameter constrained to 128 tokens, dynamic padding, and explicit truncation. We deployed the pre-trained ProsusAI/finbert model accessed via the HuggingFace repository. Given that the selected model (ProsusAI/finbert) is already rigorously pre-trained on large-scale financial corpora (e.g., Financial PhraseBank), it was deployed directly for zero-shot inference without further task-specific weight updates. The computational pipeline was accelerated within a dedicated GPU environment utilizing an NVIDIA Tesla T4 (16 GB VRAM), operating on PyTorch 2.0 and CUDA 11.8. To guarantee complete reproducibility across identical experimental setups, a strict computational protocol was enforced by establishing a global random seed of 42 alongside a deterministic backend. Ultimately, the predictive efficacy of this pre-trained architecture on our specific corpus was rigorously validated through comprehensive evaluation metrics. During inference, the attention mechanism evaluates the contextual syntax (e.g., negations), outputting discrete softmax probabilities (Ppositive, Pnegative, Pneutral) which are subsequently mapped to a continuous polarity scale [−1.0, +1.0].
The deployment of the FinBERT architecture is strictly necessitated by the complex informational environment of the analyzed market. In an economic sector characterized by high volatility and heavy regulatory intervention, standard lexicon-based approaches are fundamentally inadequate for capturing the nuanced syntax of finance. While FinBERT itself is an established tool, the specific applied contribution of this study lies in the context-specific preprocessing and aggregation pipeline designed to handle localized market discourse. The custom algorithm integrates class balance calibration and strict linguistic normalization to explicitly isolate and filter out the ‘sentiment contamination’ prevalent in regional news syndication. This contextual aggregation mechanism reduces the influence of peripheral mentions and ambiguous local acronyms on the Daily Sentiment Index, thereby improving the contextual relevance of the sentiment measure used in the subsequent econometric analysis.
Since financial markets are driven by aggregate consensus, raw output scores are synthesized into a unified time-series metric to correlate “soft data” with daily asset prices. The Daily Sentiment Index (DSI) is defined as the arithmetic mean of the polarity scores for all relevant news items published within a 24 h trading window. This method smooths out intraday noise and identifies the prevailing market narrative.
The mathematical formulation for the Index for a specific entity on day t is defined as follows:
D S I = 1 N t i = 1 N t S i , t
where N t represents the total volume of news articles associated with a specific entity or topic on day t, and S i , t is the individual polarity score of article i. It is critical to emphasize that the DSI is not a ‘raw’ metric, but a derived econometric construct that requires robust theoretical and technical justification. While the unweighted arithmetic mean serves as a baseline approximation, our architecture incorporates necessary model calibration and normalization. To differentiate event intensity and filter ambient semantic noise, we implemented a threshold tuning mechanism (class balance calibration) wherein low-probability softmax outputs (p < 0.65) are linguistically normalized to strict neutrality (Si,t = 0). Furthermore, strict entity-relevance filtering ensures that peripheral or brief mentions do not disproportionately skew the daily index. The final DSI remains a transparent, threshold-calibrated arithmetic mean, prioritizing methodological reproducibility over opaque weighting schemes. The validity of this calibrated DSI was confirmed through benchmark validation, demonstrating superior contextual resolution when compared against traditional lexicon-based alternatives such as the Loughran-McDonald financial dictionary.
Moving beyond the structural architecture, ensuring the empirical reliability of the FinBERT output required rigorous validation protocols and robustness checks. To validate the inference performance of the pre-trained FinBERT model on our localized financial corpus, a representative pool of 500 headlines was manually annotated. To ensure high domain validity, the annotation was conducted independently by two financial researchers with expertise in European market dynamics. The annotators followed strict guidelines to classify narrative sentiment regarding fundamental asset valuation, achieving a high intercoder agreement (Cohen’s Kappa = 0.82). The resulting pool of 500 manually annotated headlines provided the broader human-labeled validation set from which the 100-headline evaluation subset reported in Table 4 was selected. The selection and composition of this evaluation subset, together with the corresponding classification metrics, are described below.
To further assure methodological robustness and verify the reliability of the generated sentiment scores, continuous stability checks were performed. The FinBERT outputs were formally benchmarked against traditional lexicon-based alternative models, specifically the Loughran-McDonald dictionary. FinBERT consistently outperformed the lexicon baseline by effectively resolving contextual ambiguities, such as negations and domain-specific jargon, which traditional models frequently misclassified.
The superiority of the contextual embedding approach is quantitatively demonstrated in Table 4, which contrasts the predictive performance of the pre-trained ProsusAI/finbert model against the traditional Loughran-McDonald (LM) financial lexicon on the out-of-sample test set.
Methodologically, this 100-headline evaluation subset was selected from the broader pool of 500 manually annotated headlines to ensure representation of all three sentiment classes (40 negative, 35 positive and 25 neutral). The subset was used exclusively for evaluation and was not used for model training, fine-tuning, or parameter adjustment. Accordingly, it constitutes an evaluation holdout relative to the model-estimation process, rather than a separately collected independent sample. This composition ensured that all three sentiment classes were represented in the evaluation and facilitated comparison with the Loughran–McDonald baseline. Because the pre-trained FinBERT model was deployed without task-specific fine-tuning, these 100 observations were used exclusively for evaluation and were not employed for any model-weight updates. The complete confusion matrix is reported in Table 4. Based on this 100-headline evaluation subset, the negative class achieved a precision of 92.3% and a recall of 90.0% (F1 = 0.91); the neutral class achieved a precision of 78.6% and a recall of 88.0% (F1 = 0.83); and the positive class achieved a precision of 93.9% and a recall of 88.6% (F1 = 0.91). Overall classification accuracy was 89.0%, with a macro-F1 score of 0.88.
Furthermore, temporal stability checks were conducted by analyzing sentiment score variance across known non-volatile market periods. This confirmed that the algorithmic distribution accurately captures genuine shifts in investor psychology rather than reacting to ambient semantic noise or inherent data biases. Consequently, through out-of-sample ground-truth validation (Table 4), baseline lexicon benchmarking, and formal stationarity testing (ADF), the statistical properties and the academic robustness of the constructed Daily Sentiment Index (DSI) are rigorously verified prior to its integration into the econometric models.
We explicitly recognize that simple descriptive correlations cannot empirically demonstrate predictive power, structural sectoral decoupling, or complex market learning effects. Consequently, to establish robust inferential depth and evaluate temporal dynamics, the descriptive Daily Sentiment Index (DSI) vectors are subjected to advanced econometric modeling. This comprehensive framework incorporates Granger predictability tests to establish predictive temporal precedence (explicitly noting that this evaluates forecasting ability, not definitive economic causality). This comprehensive framework incorporates Granger predictability tests to assess incremental predictive temporal precedence, Vector Autoregression (VAR) models to capture joint dynamic interactions, GARCH(1,1) specifications to evaluate conditional volatility associations, and dynamic lead-lag analyses to uncover lead-lag narrative relationships, all supported by rigorous robustness and statistical significance testing.
To avoid the over-interpretation of strictly descriptive associations, it is imperative that all correlational and econometric results are validated through formal statistical significance testing. Consequently, our analytical procedures have been expanded to include normality and stationarity testing, the calculation of precise p-values for all correlation matrices, formal t-tests for regression coefficients, and the systematic reporting of 95% confidence intervals.
We explicitly recognize that relying solely on static modeling substantially limits the explanatory power of the study. Financial sentiment and market volatility rarely exhibit strictly linear, time-invariant relationships. To overcome the limitations of descriptive, static frameworks, this study introduces advanced analytical modeling to rigorously examine interaction effects. Specifically, we model conditional volatility using GARCH(1,1) specifications to evaluate the predictive association between sentiment and market variance. Furthermore, regime-dependent interactions are analyzed using Markov-switching techniques to evaluate structural shifts between sentiment shocks and market behavior.

4. Results

It should be explicitly noted that the decision to focus the primary econometric reporting on UiPath, OMV Petrom, and the macro-regional DAX index was made post hoc, following preliminary data exploration. While Hidroelectrica and Transilvania Bank remain central to the conceptual framework of the ‘Twin Transition’, their early correlational profiles largely mirrored broader systemic macro-trends without displaying the acute structural divergences observed in the other assets. To maintain analytical conciseness in the main text without compromising methodological transparency, the supplementary Granger predictability and GARCH(1,1) results for the prespecified entities Hidroelectrica and Transilvania Bank are reported in Appendix B.
Prior to examining the dynamic interdependencies and conditional volatility associations, it is essential to establish the distributional characteristics of the dataset. Table 5 presents the descriptive summary statistics for all primary time-series variables over the 252-trading-day sample period. The financial assets are expressed in continuous compounding returns ( Δ ln P t ), while the Daily Sentiment Index (DSI) variables are expressed in their stationary level forms.
The static sectoral interdependencies and baseline correlations between the Daily Sentiment Index and the selected market assets are visually summarized in Figure 1.
A key finding is the negative correlation identified between sentiment regarding the energy transition and traditional assets: −0.47 for OMV Petrom and −0.52 for the DAX index. Throughout this section, results are reported alongside their respective p-values and 95% confidence intervals to rigorously evaluate the robustness of the observed relationships. For the aforementioned energy transition narratives, the relationship is statistically significant at conventional levels (p < 0.01). This inverse relationship indicates the presence of a ‘carbon transition risk’; as global narratives on greening intensify, the valuation of fossil fuel-based assets and the traditional industrial economy tends to face downward pressure. In contrast, the technology sector (UiPath) exhibits a moderate positive correlation with the macroeconomic baseline proxy (DAX, 0.35), indicating that it is not entirely decoupled from broad European industrial cycles. However, its near-zero correlation (−0.07) with digital resilience sentiment and its negative correlation (−0.24) with general ESG narratives suggests a differentiated risk profile. For interactions such as this, where p-values exceed the standard 0.05 threshold, we explicitly acknowledge that the association is not statistically significant and the evidence remains inconclusive without stronger significance. Nevertheless, this provides exploratory evidence that the market may recalibrate its optimism regarding innovation based on solid financial fundamentals, not just media narratives.
However, further econometric testing is required to evaluate whether these descriptive associations persist in predictive and time-varying specifications. It is imperative to note that any claims regarding structural decoupling, market learning, or definitive narrative effects within this study are considered valid only when directly supported by the subsequent advanced econometric results, rather than relying strictly on these initial visual or static correlations.
Beyond these static sectoral interdependencies, the potential predictive relevance of investor sentiment is further evaluated under Hypothesis 1, which examines whether sentiment contains incremental information for subsequent market dynamics through the lens of algorithmic governance. This is clearly demonstrated by the complex dynamics observed in the relationship between ESG sentiment and the DAX index, where time-series analysis reveals periods of acute divergence followed by phases of convergence as the market progressively internalizes sustainability criteria. These dynamic trajectories of market performance plotted against macroeconomic sustainability narratives are graphically illustrated in Figure 2.
The analysis of the 60-day dynamic correlation (Figure 3) provides exploratory evidence regarding Hypothesis H1. We observe a descriptive pattern that we heuristically term a “regime shift” in the data associations: while in the second half of 2025 the correlation was strongly negative (reaching −0.8), the first part of 2026 exhibits a transition toward a positive correlation (+0.8). It is crucial to emphasize that this “regime shift” is utilized here strictly as an analytical metaphor to describe shifting correlational trends within the sample, rather than a proven structural transformation. Descriptive rolling correlations provide a visual proxy for evolving data relationships, but they inherently cannot demonstrate causality. These variations may be heavily influenced by unobserved macroeconomic factors, and any inferences regarding definitive changes in investor preferences remain speculative without a causal identification design.
This time-varying pattern in the association between sustainability sentiment and market performance provides a useful comparison with the differentiated behavior of the innovation-led case, providing a logical transition to the evaluation of Hypothesis 2 concerning the resilience of the technology sector and the observed limits of decoupling. Consequently, the analysis of UiPath’s performance evaluates whether the technology case exhibits relative statistical divergence from traditional macroeconomic cycles, illustrating how innovation narratives can insulate specific assets from the stagnation of the broader industrial landscape.
It is imperative to emphasize that rolling correlations alone cannot establish causal changes in investor preferences or structural market evolution; they strictly indicate dynamic modifications in the co-movement of the time series. Because the study does not employ a causal identification strategy, these initial visual associations must be interpreted with caution. Consequently, to move beyond descriptive co-movements, the analysis proceeds to formal predictive and time-series tests.
While the descriptive rolling correlation in Figure 3 visually suggests a ‘regime shift’ in investor behavior, relying solely on exploratory visualizations is insufficient to confirm structural market transformations.
Prior to estimating the Markov-switching regime model, it is mandatory to satisfy specific preliminary econometric conditions: stationarity in the presence of breaks, and non-linearity. To validate these assumptions, we conducted a suite of diagnostic tests. First, alongside the standard ADF tests, Kwiatkowski-Phillips-Schmidt-Shin (KPSS) tests and ADF Breakpoint tests were applied, confirming that the variables remain stationary in their transformed states even when structural breaks are accounted for ( p < 0.05 ). Second, to formally justify the use of a non-linear regime-switching framework, we applied the Brock-Dechert-Scheinkman (BDS) test on the residuals of a baseline linear specification. The BDS test strongly rejected the null hypothesis of independent and identically distributed (i.i.d.) linear series across all embedding dimensions ( p < 0.01 ), confirming robust non-linear dependence. Finally, the presence of these distinct variance states was formally identified using the Bai-Perron structural break analysis.
Furthermore, to rigorously validate the regime specification of the Markov model, we estimated both two-regime and three-regime specifications. The optimal number of unobserved states was determined utilizing standard information criteria. The two-regime model was definitively selected for the final estimation, as it strictly minimized both the Akaike Information Criterion (AIC = −854.3) and the Bayesian Information Criterion (BIC = −832.1) compared to the three-regime alternative (AIC = −841.5, BIC = −805.2), effectively capturing the structural shift without overparameterization.
To objectively investigate this transition without relying solely on visual biases or simple changes in correlation signs, we applied a Bai-Perron structural break analysis alongside a Markov-switching regime model (Table 6). The Markov-switching dynamic regression was estimated using maximum likelihood, with the unobserved regimes identified strictly based on distinct conditional variance states ( σ ). The qualitative regime labels presented in the results, ‘High Volatility/Risk Penalty’ and ‘Low Volatility/Value Alignment’, were assigned post-estimation. They simply map the mathematically derived high-variance state to the period of negative sentiment correlation, and the low-variance state to the period of positive sentiment alignment, providing an economic interpretation for the statistically identified structural break. While these econometric tests successfully identify a mathematical breakpoint, we explicitly acknowledge that a statistical shift does not automatically confirm a fundamental economic transition from a ‘cost’ to a ‘value driver.’ The observed changes in correlations could be driven by several alternative explanations, including transient macroeconomic interventions, structural variations in media reporting, base effects following highly volatile periods, or broader shifts in overall investor risk appetite that are independent of ESG metrics. Therefore, rather than definitively confirming a structural transformation, we state that the observed patterns are merely consistent with a potential transition toward a value-driven ESG perception. Robustness tests (such as VAR stability diagnostics and rolling window coefficients) are required to confirm a genuine structural regime change.
The relative pricing resilience of the selected technology asset during periods of fluctuating digital innovation sentiment is depicted in Figure 4.
Although UiPath’s price shows visible resilience during periods of stress in innovation sentiment, the regression analysis (Figure 5) reveals an almost flat trend slope (−0.07 correlation). Given the moderate 0.35 correlation with the DAX, claims of absolute structural decoupling are unsupported; rather, the technology asset demonstrates a differentiated sensitivity. It remains relatively insulated from ESG-related systemic risk narratives (as evidenced by the −0.24 correlation) while exhibiting distinct vulnerability to its own technological momentum cycles. To visually assess these differing sensitivities across all three structural pillars, Figure 5 plots the cross-sectional regression trends for the macroeconomic, technological, and energy cases. In these scatter plots, each individual dot represents a daily paired observation of the FinBERT-derived narrative sentiment score and the corresponding market asset performance. The solid colored lines delineate the linear regression trend (the line of best fit) for each sector, indicating the directional association between textual sentiment and market valuation. Furthermore, the surrounding shaded areas represent the 95% confidence intervals, illustrating the degree of statistical uncertainty around the estimated trend lines.
The exploratory regression models presented in Figure 5 provide preliminary associations regarding sectoral decoupling and transition risks. However, to assess predictive temporal precedence and evaluate asymmetric conditional-volatility associations, it is imperative to move beyond static correlations. Consequently, we subjected the time-series variables to formal Granger predictability testing (Table 7) and GARCH(1,1) conditional volatility modeling (Table 8). These inferential frameworks evaluate whether lagged sentiment contains information about subsequent market variance. Crucially, they provide robust statistical support for the differentiated volatility profile of the technology sector (which exhibits an insignificant variance response to macro-sentiment) alongside the structural vulnerabilities of traditional energy assets.
For all subsequent time-series models, the sample size consists of n = 252 trading day observations. To strictly satisfy the stationarity requirements of VAR and Granger predictability frameworks, all financial asset prices (DAX, OMV Petrom, UiPath) were transformed into continuous compounding returns using first-logarithmic differencing (ΔlnPt), while the Daily Sentiment Index (DSI) was utilized in its stationary level form. Augmented Dickey–Fuller (ADF) tests confirmed the absence of unit roots (p < 0.01) across all transformed series. Optimal lag lengths for both the Granger predictability tests and the VAR system were selected dynamically by minimizing the Bayesian Information Criterion (BIC), which consistently identified a lag order of p = 2. All models include a constant as the sole deterministic term.
Prior to estimating the conditional volatility models, preliminary diagnostic testing was conducted to justify the GARCH framework. Engle’s ARCH-LM test was applied to the ordinary least squares (OLS) residuals of the baseline mean equations, strongly rejecting the null hypothesis of homoscedasticity ( p < 0.01 ) and confirming the presence of significant ARCH effects. Furthermore, to ensure optimal model specification, alternative asymmetric volatility variants (specifically EGARCH and GJR-GARCH) were estimated. The standard GARCH(1,1) specification was ultimately selected for the final estimation as it strictly minimized both the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC), providing the most parsimonious fit for the given sample size without overparameterization.
To fully capture the volatility dynamics, the GARCH(1,1) model with an exogenous sentiment regressor is specified with a standard mean equation, R t = μ + ϵ t , and a conditional variance equation defined as: σ t 2 = ω + α ϵ t 1 2 + β σ t 1 2 + γ D S I t 1 . A Student-t distributional assumption was utilized to account for the heavy tails typically observed in financial returns. The complete estimation parameters, including persistence ( α + β ) and residual diagnostics, are presented in Table 8.
It must be explicitly noted that because the specific econometric objective of this stage is to measure the predictive association between lagged narrative sentiment and individual-asset variance from an exogenous narrative index to individual asset variance, rather than estimating the dynamic conditional covariance matrix between multiple asset returns, a univariate GARCH specification with an exogenous regressor (GARCH-X) is the methodologically appropriate choice, distinct from multivariate frameworks such as ADCC-GARCH.
The interpretation of the exogenous sentiment coefficient ( γ ) in the variance equation warrants specific clarification. Because the Daily Sentiment Index (DSI) is scaled continuously between −1.0 (extreme pessimism) and +1.0 (extreme optimism), the negative γ coefficients reported in Table 8 are highly economically intuitive. Within this specification, a negative coefficient indicates that negative DSI values are associated with higher estimated conditional variance, whereas positive DSI values are associated with lower estimated conditional variance. This pattern is consistent with asymmetric sentiment–volatility associations, without implying a causal effect of sentiment on volatility. We formally verified that the estimated conditional variance ( σ t 2 ) remained strictly positive across all observations in our sample, as the baseline persistence parameters ( ω , α , β ) consistently dominated the exogenous narrative term. This functional form was selected to evaluate whether the sign of narrative sentiment is statistically associated with differences in estimated conditional variance.
We explicitly note that any categorical claim regarding early-warning capabilities requires strict statistical confirmation; therefore, the evidence presented here merely suggests potential predictive value. Preliminary results indicate that DSI may precede volatility movements, but forecasting capabilities are established strictly through the out-of-sample accuracy metrics presented subsequently.
The requirement for such sophisticated sentiment filtering in the digital space highlights the multifaceted nature of narrative associations across different industries, leading directly into the re-evaluation of Hypothesis 3, which examines the distinct dynamics of the energy transition and the emergence of diminishing marginal impact of economic narratives. Consequently, Hypothesis 3 is re-examined through the lens of the green transition (Figure 6), where the empirical results partially refute the expectation of a simple positive correlation between market sentiment and price discovery in the energy sector.
The negative correlation illustrated in Figure 1 and the downward trajectory in Figure 6 are consistent with an interpretation in which energy-transition narratives are associated with perceived long-term profitability risks for the oil and gas sector. Geopolitical developments formed part of the broader informational environment during the sample period; however, geopolitical sentiment was not modeled as a separate explanatory series and was not formally compared with transition-related sentiment. Accordingly, no inference is made regarding the relative importance of transition-related versus geopolitical narratives for OMV Petrom. The analysis of Hypothesis 3 is therefore restricted to the empirically tested association between energy-transition sentiment and the selected energy asset. The distributional density and media polarization across the three distinct narrative themes are compared in Figure 7.
We observe that narratives about AI/Digital exhibit very high density and low volatility (a sharp curve), indicating a relative consensus in public discourse. In contrast, narratives about Energy and ESG exhibit much broader and more irregular distributions, reflecting intense polarization and high informational uncertainty. This dispersion in sentiment is consistent with the more unstable correlations observed for energy assets and why FinBERT-based NLP-based market surveillance is essential for filtering out noise in a fragmented media landscape.
Finally, to ensure that the FinBERT-derived Daily Sentiment Index (DSI) possesses genuine forecasting utility rather than mere historical data fit, we conducted a rigorous out-of-sample predictive validation. The specific target variable for this analysis was the daily logarithmic return of the DAX index ( Δ DAX Returns). To rigorously prevent any same-day information leakage, a strict t 1 lag structure was enforced; the model utilizes only sentiment information explicitly published and aggregated prior to the prediction timestamp (day t ) to compute 1-day-ahead forecasts.
The evaluation utilized a rigorous expanding-window approach. We designated the first 200 trading days as the initial training set (the first forecast origin). The window then expanded by one day at a time, re-estimating the model parameters at each step, yielding 52 out-of-sample one-day-ahead forecasts. The benchmark model is explicitly specified as a univariate Autoregressive AR(2) model of DAX returns, matching the BIC-selected lag length of the competing system to ensure a strictly fair comparison. The complete system estimates for this sentiment-augmented VAR model are presented in Table 9. Subsequently, Table 10 compares the forecasting errors of this AR(2) benchmark against the enhanced VAR-based model incorporating the DSI sentiment metrics. Because the Mean Absolute Percentage Error (MAPE) is mathematically unstable when evaluating daily returns that cross zero, predictive performance was evaluated strictly using Root Mean Square Error (RMSE) and Mean Absolute Error (MAE). The statistical significance of the predictive improvement was formally evaluated utilizing the Diebold-Mariano (DM) test, implemented with a squared-error loss function and Newey-West robust standard errors to account for potential autocorrelation in the forecast errors.
Pre-estimation Diagnostics: Augmented Dickey–Fuller (ADF) confirms stationarity for all series (p < 0.01). Lag order p = 2 selected via minimum Bayesian Information Criterion (BIC).

5. Discussion

The empirical findings are consistent with the narrative-economics perspective and the AMH framework, indicating that textual sentiment can contain incremental information relevant to short-horizon market dynamics. By complementing lagged hard data with higher-frequency soft data, the analysis supports the use of algorithmic sentiment measures as supplementary tools for financial monitoring and forecasting.
By transitioning from static correlations to dynamic modeling, the interpretation of our results significantly deepens. The conditional volatility associations identified via the GARCH(1,1) specifications and the predictive temporal precedence indicated by Granger predictability testing suggest that lagged sentiment indicators contain statistically significant predictive information for conditional market variance. Furthermore, the regime-dependent interactions modeled through Markov-switching techniques indicate that the interaction strength between narratives and market behavior is not static, evolving significantly across different phases of macroeconomic stress.
In this analytical context, it must be clarified that this study does not intend to model long-term structural macroeconomic trends, but rather to examine short-horizon behavioral and informational patterns observable in the data.
These empirical findings are consistent with the Adaptive Market Hypothesis [3] insofar as they indicate that the relationship between ESG sentiment and market outcomes may vary over time. While past literature, such as Pástor et al. (2022) [8], discusses changes in green-asset performance and environmental preferences in mature markets, the present results show a comparatively rapid adjustment in correlation structures within this specific empirical setting. This descriptive “regime shift” should not be interpreted as direct evidence of a structural change in investor preferences or capital allocation. Rather, it is consistent with the possibility that market participants reassessed transition-related information, regulatory liabilities, and expected cost structures during the analyzed period. The evidence therefore supports a time-varying interpretation of sentiment–market relationships without establishing the underlying causal mechanism.
The evaluation of Hypothesis 1 provides preliminary insights into market responsiveness. The observed fluctuation in the correlation between ESG sentiment and the DAX index (shifting from a stark −0.80 in late 2025 to a positive +0.80 in 2026) suggests a potential trend in how sustainability narratives align with market performance over the analyzed period. Rather than claiming a definitive structural maturation of the market, we interpret this descriptive “regime shift” strictly as an exploratory hypothesis. It suggests that sustainability metrics may have aligned with value-driving factors during this specific timeframe. Beyond the theoretical frameworks, these findings suggest critical implications for sustainable finance and responsible investment strategies. The dynamic nature of ESG sentiment is associated with changes in systemic risk perception, indicating that sustainability narratives may provide informative signals about market stability during periods of transition. For ESG-oriented investors, the evidence indicates that narrative-driven sentiment metrics provide a crucial operational advantage; they can act as early signals for shifts in market perception, assist in dynamically managing reputational risks, and support resilient capital allocation strategies in highly volatile macroeconomic contexts. Therefore, Hypothesis 1 (H1) is partially supported: the empirical evidence confirms predictive temporal precedence and forecasting utility, but stops short of proving a definitive, causal early-warning mechanism.
Regarding Hypothesis 2, the results indicate a differentiated statistical sensitivity for the technology sector (exemplified by UiPath) relative to broader market downturns. This finding extends the theoretical assertions of Brynjolfsson & McAfee (2014) [18] regarding digital operational sustainability into a high-frequency financial context. While the near-zero correlation observed with German industrial stagnation does not provide sufficient evidence to support absolute structural decoupling, it indicates a relative statistical divergence. The observed pattern suggests that digital-first assets may exhibit differentiated sensitivity to traditional industrial-market conditions. During periods of weakness in traditional industrial indicators, the observed relative performance of innovation-driven assets may be consistent with investors assigning greater value to operational efficiency and automation; however, the present analysis does not directly measure portfolio flows or establish a structural hedging mechanism. Therefore, Hypothesis 2 (H2) is partially supported: the results indicate differentiated statistical sensitivity in the selected technology case, but do not provide sufficient evidence to establish absolute structural decoupling.
The findings for Hypothesis 3 indicate that transition-related sentiment is statistically associated with the market dynamics of the selected traditional energy asset, OMV Petrom, over the analyzed period. This interpretation is consistent with the transition-risk perspective discussed by Engle et al. (2020) [11]. Importantly, geopolitical sentiment was not modeled as a separate explanatory series and was not formally tested against transition-related sentiment. The study therefore does not draw conclusions about whether OMV Petrom responds more strongly to transition narratives than to geopolitical conflict. Hypothesis 3 (H3) is supported only with respect to the observed association and predictive information linked to energy-transition sentiment within the specified empirical framework.
The broader implications of these results suggest that financial markets function as complex adaptive systems. The high density and low volatility found in AI/Digital narratives, contrasted with the intense polarization of ESG and Energy discourse, highlight the necessity of algorithmic governance. Tools like FinBERT are essential for filtering “sentiment contamination”, such as misinterpreting sports news for market signals, ensuring that liquidity and risk perception are based on accurate context.
Furthermore, the results may profoundly inform corporate sustainability communication and the formulation of sustainable finance policies. Quantitative narrative indicators serve as vital tools for policymakers to monitor emerging transition risks and gauge the public reception of ESG regulatory interventions. In the context of non-financial disclosure, real-time sentiment indicators can effectively complement the traditional, static compliance metrics mandated by frameworks such as the Corporate Sustainability Reporting Directive (CSRD) [29] or the European Sustainability Reporting Standards (ESRS). By applying computational narrative analysis, stakeholders can detect critical structural gaps between official corporate reporting and actual public perception, thereby ensuring higher transparency and market integrity.

6. Conclusions

The study provides evidence that high-frequency textual sentiment contains quantifiable incremental information relevant to short-horizon market dynamics and conditional variance. Synthesizing the empirical evidence, we find heterogeneous sentiment–market relationships across the selected cases during the twin green and digital transition. The results indicate that transition-related and industrial-demand narratives contain context-specific information associated with the observed market dynamics over the analyzed period. Specifically, the selected traditional energy case exhibits statistically significant associations with energy-transition sentiment, whereas the selected technology case displays differentiated sensitivity to broader industrial-market conditions. Ultimately, these findings indicate that integrating context-sensitive NLP metrics into econometric frameworks can complement lagged ‘hard data’ with higher-frequency textual information.
Implications for Practice and Policy The synthesized findings present immediate, actionable implications for key market participants. For investors and fund managers, the FinBERT-derived sentiment index serves as a high-frequency risk management tool, potentially supporting dynamic portfolio monitoring and risk assessment alongside conventional economic indicators. For policymakers and regulatory bodies, computational narrative monitoring provides a real-time feedback mechanism to gauge market reception to ESG regulatory interventions (such as the CSRD). This capability allows for more adaptive sustainable finance policies and facilitates the rapid identification of structural gaps between official corporate disclosures and actual public perception.
Limitations The empirical scope of this study is inherently bounded by its 12-month longitudinal window and concentrated regional focus, precluding the immediate generalization of these short-term dynamics to full macroeconomic cycles or broader continental scales. Additionally, as noted in the sample design, utilizing a limited number of specific companies as proxies introduces the risk of capturing firm-specific idiosyncratic shocks (e.g., leadership changes or specific contracts) rather than broad sectoral trends. Methodologically, while the deployment of univariate GARCH-X and Markov-switching models effectively captures directional variance and structural shifts, these exploratory inferences remain sensitive to model specification and parameter tuning, and cannot establish structural or causal decoupling without expanded multivariate frameworks. Furthermore, any NLP-derived index remains inherently constrained by the underlying corpus quality, necessitating rigorous threshold calibration to mitigate semantic noise and media bias.
Future Research Directions Building upon these constraints, several avenues remain for technological and behavioral refinement. Future research should extend this framework across a full economic cycle to determine if sectoral divergence persists during global recessions. Investigating the granularity between institutional sentiment (e.g., professional terminals) and retail discourse (e.g., social media) could reveal unique short-term behavioral arbitrage opportunities. Finally, exploring the temporal “latency period” in the era of algorithmic trading could provide a precise map of how rapidly distinct ESG narratives are internalized into asset price discovery. Ultimately, the future of sustainable finance relies on the synergy between massive computational data processing and strategic human intuition.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/su18179177/s1, Archive S1: Complete Replication Package (including purified datasets, manually annotated validation subset, and Python/econometric codebases).

Author Contributions

Conceptualization, C.B.; methodology, D.M.N. and T.C.; software, T.C.; validation, C.B., C.-V.H., D.M.N., T.C. and O.-M.B.; formal analysis, D.M.N., T.C. and O.-M.B.; investigation, C.-V.H. and T.C.; resources, C.B.; data curation, T.C.; writing—original draft preparation, T.C. and D.M.N.; writing—review and editing, C.B., C.-V.H. and O.-M.B.; visualization, T.C.; supervision, C.B. and C.-V.H.; project administration, C.B. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data supporting the findings of this study were derived from publicly available sources (The Guardian Open Platform, Yahoo Finance). The complete Replication Package, including the purified datasets, the manually annotated validation subset, and the comprehensive Python V3.10/econometric codebases required to reproduce all textual inferences, time-series models (VAR, GARCH), forecasting metrics, and figures, is provided as a compressed archive in the Supplementary Materials accompanying this manuscript, ensuring absolute methodological transparency.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AbbreviationFull Definition/Meaning
AIArtificial Intelligence, Used in the context of innovation-driven narratives and operational optimization.
AMHAdaptive Market Hypothesis, The evolutionary framework conceptualizing financial ecosystems where bounded rationality adapts to systemic shocks.
BERTBidirectional Encoder Representations from Transformers, A deep learning architecture used for language understanding.
BVBBursa de Valori București, The Bucharest Stock Exchange, noted for context-aware purification to avoid confusion with sports clubs.
CEECentral and Eastern Europe, Regional designation used in the context of OMV Petrom and transitional markets.
DAXDeutscher Aktienindex, The primary stock market index for the German economy, used as a macro regional control.
DSIDaily Sentiment Index, A quantitative time-series metric derived from aggregated FinBERT polarity scores.
EGARCHExponential General Autoregressive Conditional Heteroskedasticity, An asymmetric volatility model used to capture leverage effects.
EMHEfficient Market Hypothesis, The traditional theory that asset prices reflect all available fundamental information.
ESGEnvironmental, Social, and Governance, Core systematic risk factors and narratives associated with corporate resilience.
EUEuropean Union, Referring to regional projects such as the digital ID wallet initiative.
GDPGross Domestic Product, A “Hard Data” macroeconomic indicator cited for its reporting latency.
LLMLarge Language Model, Advanced computational models capable of distinguishing subtle semantic nuances in text.

Appendix A. Methodological and Computational Specifications

To ensure complete replicability, this appendix details the exact computational workflow utilized in the study. The primary textual corpus was constructed by aggregating data from The Guardian Open Platform API and Yahoo Finance API over a strict 12-month longitudinal period ending on 12 May 2026. The initial extraction yielded a raw corpus of 18,452 financial news articles. The data cleaning workflow applied automated filtering to remove syndicated media duplicates, non-English articles, and items with insufficient informational density (specifically, headlines containing fewer than 5 words), resulting in a final analytical sample of 14,210 unique observations. The distributional breakdown of these observations included 4150 articles for UiPath, 5320 for the energy sector (OMV Petrom and Hidroelectrica), 2840 for Transilvania Bank, and 1900 for the German DAX index.
To eliminate the “black box” nature often associated with machine learning, the computational framework is explicitly defined as a transparent sequential pipeline. Prior to model ingestion, the text underwent a minimal preprocessing phase restricted strictly to UTF-8 normalization and HTML tag stripping; natural punctuation and stop-words were explicitly preserved to ensure the Transformer’s attention mechanism could accurately resolve contextual syntax and negations. Text was then tokenized using the native BERT WordPiece tokenizer with a maximum sequence length of 128 tokens, dynamic padding, and explicit truncation.
The model implementation utilized the pre-trained ProsusAI/finbert architecture accessed via the HuggingFace repository. Given its robust baseline on financial corpora, it was deployed directly for zero-shot inference without further task-specific weight updates. The computational pipeline was accelerated on an NVIDIA Tesla T4 GPU running PyTorch 2.0 and CUDA 11.8, with strict reproducibility enforced through a global random seed set to 42. Sentiment classification was achieved by evaluating contextual syntax through the model’s attention mechanism, which outputted discrete softmax probabilities mapped to a continuous polarity scale. Model inference performance was rigorously validated against a manually annotated 100-headline holdout subset, achieving an exact overall accuracy of 89.0% and a macro-F1 score of 0.88, confirmed by a robust confusion matrix. Prior to econometric modeling, the generated DSI and associated financial time series were subjected to stationarity testing and logarithmic differencing where necessary to ensure valid inferential testing.
The complete, fully commented Python codebase utilized to execute this sequential pipeline, alongside the purified datasets and evaluation benchmarks, is provided as a compressed Replication Package in the Supplementary Materials accompanying this manuscript.

Appendix B. Supplementary Econometric Results (Hidroelectrica and Transilvania Bank)

To ensure complete methodological transparency and alignment with the prespecified sample design (Section 3.2), this appendix provides the core econometric outputs for the domestic control entities: Hidroelectrica (representing the pure-play green energy transition) and Transilvania Bank (representing digital banking inclusion).
Table A1. Granger Predictability Testing Results (Supplementary Entities, Lag = 2).
Table A1. Granger Predictability Testing Results (Supplementary Entities, Lag = 2).
Null Hypothesis (H0)F-Statisticp-ValueDecision
Energy Sentiment does not Granger-predict Hidroelectrica2.1450.118Fail to Reject
Digital/Banking Sentiment does not Granger-predict Transilvania Bank3.0120.049 *Reject H0
Notes: Unlike OMV Petrom, which exhibited high sensitivity to transition narratives, the pure-play renewable asset (Hidroelectrica) showed no statistically significant Granger-predictive relationship with energy-transition sentiment in this specification, providing no evidence of predictive sensitivity over the analyzed sample. Conversely, Transilvania Bank exhibited a marginal but statistically significant temporal sensitivity to digital inclusion narratives (p < 0.05). * denotes statistical significance at the 5% level (p < 0.05).
Table A2. GARCH(1,1) Conditional Volatility Association Parameters (Supplementary Entities). Mean Equation: R t = μ + ϵ t | Variance Equation: σ t 2 = ω + α ϵ t 1 2 + β σ t 1 2 + γ D S I t 1 .
Table A2. GARCH(1,1) Conditional Volatility Association Parameters (Supplementary Entities). Mean Equation: R t = μ + ϵ t | Variance Equation: σ t 2 = ω + α ϵ t 1 2 + β σ t 1 2 + γ D S I t 1 .
ParameterHidroelectrica (Exog: Energy Trans.)Transilvania Bank (Exog: Digital Sentiment)
Constant ( ω )0.0002 (0.01)0.0001 (0.01)
ARCH lag 1 ( α )0.085 (0.02) **0.092 (0.03) **
GARCH lag 1 ( β )0.860 (0.04) **0.845 (0.05) **
Sentiment Exog ( γ )−0.031 (0.07) [Not Sig.]−0.105 (0.04) **
Notes: Standard errors in parentheses. ** denotes p < 0.01 . The statistically insignificant exogenous sentiment parameter ( γ ) for Hidroelectrica provides no evidence of a significant conditional-volatility association with transition-related sentiment over the sample period, whereas the Transilvania Bank estimate indicates a slight negative association between positive digitalization sentiment and conditional variance.

References

  1. Fama, E.F. Efficient Capital Markets: A Review of Theory and Empirical Work. J. Financ. 1970, 25, 383–417. [Google Scholar] [CrossRef] [Scilit]
  2. Kahneman, D.; Tversky, A. Prospect Theory: An Analysis of Decision under Risk. Econometrica 1979, 47, 263–291. [Google Scholar] [CrossRef] [Scilit]
  3. Lo, A.W. The Adaptive Markets Hypothesis: Market Efficiency from an Evolutionary Perspective. J. Portf. Manag. 2004, 30, 15–29. [Google Scholar] [CrossRef] [Scilit]
  4. Folke, C. Resilience (Republished). Ecol. Soc. 2016, 21, 44. [Google Scholar] [CrossRef] [Scilit]
  5. Walker, B.; Salt, D. Resilience Thinking: Sustaining Ecosystems and People in a Changing World; Island Press: Washington, DC, USA, 2020. [Google Scholar]
  6. Giglio, S.; Maggiori, M.; Rao, K.; Stroebel, J.; Weber, A. Climate Change and Long-Run Discount Rates: Evidence from Real Estate. Rev. Financ. Stud. 2021, 34, 3527–3571. [Google Scholar] [CrossRef] [Scilit]
  7. Friede, G.; Busch, T.; Bassen, A. ESG and Financial Performance: Aggregated Evidence from More than 2000 Empirical Studies. J. Sustain. Financ. Invest. 2015, 5, 210–233. [Google Scholar] [CrossRef] [Scilit]
  8. Pástor, Ľ.; Stambaugh, R.F.; Taylor, L.A. Dissecting Green Returns. J. Financ. Econ. 2022, 146, 361–402. [Google Scholar] [CrossRef] [Scilit]
  9. Shiller, R.J. Narrative Economics: How Stories Go Viral and Drive Major Economic Events; Princeton University Press: Princeton, NJ, USA, 2019. [Google Scholar] [CrossRef] [Scilit]
  10. Albuquerque, R.; Koskinen, Y.; Zhang, C. Corporate Social Responsibility and Firm Risk: Theory and Empirical Evidence. Manag. Sci. 2019, 65, 4451–4469. [Google Scholar] [CrossRef] [Scilit]
  11. Engle, R.F.; Giglio, S.; Kelly, B.; Lee, H.; Johannes, S. Hedging Climate Change News. Rev. Financ. Stud. 2020, 33, 1184–1216. [Google Scholar] [CrossRef] [Scilit]
  12. Tetlock, P.C. Giving Content to Investor Sentiment: The Role of Media in the Stock Market. J. Financ. 2007, 62, 1139–1168. [Google Scholar] [CrossRef] [Scilit]
  13. Loughran, T.; McDonald, B. When Is a Liability Not a Liability? Textual Analysis, Dictionaries, and 10-Ks. J. Financ. 2011, 66, 35–65. [Google Scholar] [CrossRef] [Scilit]
  14. Hassan, T.A.; Hollander, S.; Kalyani, A.; van Lent, L.; Schwedeler, M.; Tahoun, A. Text as Data in Economic Analysis. J. Econ. Perspect. 2025, 39, 193–220. [Google Scholar] [CrossRef] [Scilit]
  15. Morosan-Danila, L.; Grigoras-Ichim, C.-E.; Bordeianu, O.-M.; Neamtu, D.-M.; Agheorghiesei, D.-T.; Filipeanu, D.; Tugui, A. Explainable AI for Predicting and Justifying Firm-Level Financial Resilience in Healthcare Services. Electronics 2026, 15, 1022. [Google Scholar] [CrossRef] [Scilit]
  16. Henfridsson, O.; Bygstad, B. The Generative Mechanisms of Digital Infrastructure Evolution. MIS Q. 2013, 37, 907–931. [Google Scholar] [CrossRef] [Scilit]
  17. Baker, S.R.; Bloom, N.; Davis, S.J. Measuring Economic Policy Uncertainty. Q. J. Econ. 2016, 131, 1593–1636. [Google Scholar] [CrossRef] [Scilit]
  18. Brynjolfsson, E.; McAfee, A. The Second Machine Age: Work, Progress, and Prosperity in a Time of Brilliant Technologies; W. W. Norton & Company: New York, NY, USA, 2014. [Google Scholar] [CrossRef] [Scilit]
  19. Caldara, D.; Iacoviello, M. Measuring Geopolitical Risk. Am. Econ. Rev. 2022, 112, 1194–1225. [Google Scholar] [CrossRef] [Scilit]
  20. Breiman, L. Statistical Modeling: The Two Cultures. Stat. Sci. 2001, 16, 199–231. [Google Scholar] [CrossRef] [Scilit]
  21. Varian, H.R. Big Data: New Tricks for Econometrics. J. Econ. Perspect. 2014, 28, 3–28. [Google Scholar] [CrossRef] [Scilit]
  22. OECD. Policy Framework for Resilience in the Energy Sector; OECD Publishing: Paris, France, 2023. [Google Scholar]
  23. Geels, F.W. Socio-technical transitions to sustainability: A review of criticisms and elaborations of the Multi-Level Perspective. Curr. Opin. Environ. Sustain. 2019, 39, 187–201. [Google Scholar] [CrossRef] [Scilit]
  24. Sovacool, B.K.; Hess, D.J.; Amir, S.; Geels, F.W.; Hirsh, R.; Medina, L.R.; Miller, C.; Palavicino, C.A.; Phadke, R.; Ryghaug, M.; et al. Sociotechnical Agendas: Reviewing Future Directions for Energy and Climate Research. Energy Res. Soc. Sci. 2020, 70, 101617. [Google Scholar] [CrossRef] [Scilit]
  25. World Bank. World Development Report 2022: Finance for an Equitable Recovery; World Bank: Washington, DC, USA, 2022. [Google Scholar] [CrossRef] [Scilit]
  26. Manning, C.D.; Schütze, H. Foundations of Statistical Natural Language Processing; MIT Press: Cambridge, MA, USA, 1999. [Google Scholar] [CrossRef] [Scilit]
  27. Han, J.; Kamber, M.; Pei, J. Data Mining: Concepts and Techniques, 3rd ed.; Morgan Kaufmann: Waltham, MA, USA, 2011. [Google Scholar] [CrossRef] [Scilit]
  28. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. arXiv 2017, arXiv:1706.03762. [Google Scholar] [CrossRef] [Scilit]
  29. Butnaru, G.I.; Neamţu, D.-M.; Dragolea, L.-L. The Impact of the CSRD on Managerial Strategies and Sustainable Competitive Advantages in the Tourism Industry. Sustainability 2026, 18, 2174. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Year Correlation Matrix. Macro Sentiment vs. Markets.
Figure 1. Year Correlation Matrix. Macro Sentiment vs. Markets.
Sustainability 18 09177 g001
Figure 2. ESG Sentiment vs. German Market (DAX).
Figure 2. ESG Sentiment vs. German Market (DAX).
Sustainability 18 09177 g002
Figure 3. Exploratory 60-Day Rolling Correlation ESG Sentiment vs. DAX Index. The solid purple line traces the dynamic correlation coefficient over the analyzed period. The shaded pink area highlights the phase where this correlation exceeds the +0.5 threshold (marked by the dashed red line), visually illustrating the period of strong positive alignment between sustainability narratives and market performance.
Figure 3. Exploratory 60-Day Rolling Correlation ESG Sentiment vs. DAX Index. The solid purple line traces the dynamic correlation coefficient over the analyzed period. The shaded pink area highlights the phase where this correlation exceeds the +0.5 threshold (marked by the dashed red line), visually illustrating the period of strong positive alignment between sustainability narratives and market performance.
Sustainability 18 09177 g003
Figure 4. Tech Sector Resilience (UiPath).
Figure 4. Tech Sector Resilience (UiPath).
Sustainability 18 09177 g004
Figure 5. Scatter plots (ESG -> DAX, AI -> UiPath, Energy -> OMV).
Figure 5. Scatter plots (ESG -> DAX, AI -> UiPath, Energy -> OMV).
Sustainability 18 09177 g005
Figure 6. The Energy Transition Effect (OMV Petrom).
Figure 6. The Energy Transition Effect (OMV Petrom).
Sustainability 18 09177 g006
Figure 7. Media Polarization: Sentiment Volatility.
Figure 7. Media Polarization: Sentiment Volatility.
Sustainability 18 09177 g007
Table 1. Critical Synthesis of Existing Literature and Identified Research Gaps.
Table 1. Critical Synthesis of Existing Literature and Identified Research Gaps.
Domain/Theoretical ApproachPrevailing MethodologyGeneral Results and FindingsMethodological LimitationsIdentified Research Gaps
Behavioral Finance and Narrative Economics (Sentiment vs. Risk-based pricing)Lexicon-based sentiment analysis, Survey data, Linear regressionsMedia pessimism correlates with downward price pressure.Relies on static dictionaries; cannot resolve financial semantic ambiguity.Lack of advanced NLP integration; inability to process high-dimensional narrative velocity.
ESG and Systemic Resilience (Adaptive Market Hypothesis)Panel data regressions, Historical financial ratios (Hard Data)Empirical evidence is mixed (positive, negative, and neutral impacts on value).Retrospective bias (lagged reporting); fails to capture real-time market regime shifts.Need for real-time “nowcasting” of ESG perception using unstructured Soft Data.
NLP-Finance and Algorithmic Governance (Machine Learning/Deep Learning)Transformer models (BERT, FinBERT), Simple correlationsHigh accuracy in sentiment classification; successful short-term forecasting.“Black box” implementations; limited formal evaluation of temporal and incremental predictive relationships; focus solely on forecasting.Limited integration of context-sensitive NLP with transparent econometric testing, time-varying analysis, and out-of-sample predictive evaluation.
Table 2. Research Hypotheses, Supporting Authors, and Methodology.
Table 2. Research Hypotheses, Supporting Authors, and Methodology.
HypothesisAuthorsMethodology
H1. The Predictive Precedence of Sentiment. Aggregated sentiment scores derived from real-time financial news provide incremental predictive information for short-horizon market movements and volatility (proxied by the DAX index) beyond the information already contained in historical price dynamics.Shiller (2019) [9], Tetlock (2007) [12], Baker et al. (2016) [17].The study utilizes the FinBERT architecture to construct a Daily Sentiment Index (DSI). This analysis tests whether the DSI can function as a ‘nowcast’ of investor sentiment, evaluating its predictive capacity regarding broad market movements and volatility (proxied by the DAX index) ahead of lagged fundamental corrections.
H2. Technological Divergence and Sensitivity. Technology-focused assets (e.g., UiPath) exhibit a differentiated sensitivity to broad macroeconomic and ESG narratives, demonstrating a relative statistical divergence from traditional industrial market cycles rather than absolute structural decoupling.Brynjolfsson & McAfee (2014) [18], Hassan, T.A.
(2025) [14]
The methodological approach involves Pearson Correlation Matrices and Scatter Plot analysis. It specifically measures the near-zero correlation between digital-first assets (e.g., UiPath) and industrial stagnation narratives of dominant economies like Germany.
H3. Energy Transition Sensitivity. Market valuation and conditional volatility in the traditional energy case (e.g., OMV Petrom) exhibit statistically significant associations with energy-transition narratives, with lagged sentiment potentially providing incremental predictive information for subsequent market dynamics.Engle et al. (2020) [11], Caldara & Iacoviello (2022) [19], Folke (2016) [4].This evaluates the time-varying association between energy-transition sentiment and the market dynamics of energy entities such as OMV Petrom.
Table 3. Sample Design, Justification, and External Validity Constraints.
Table 3. Sample Design, Justification, and External Validity Constraints.
EntitySectorCountry/RegionJustification for InclusionAssociated Limitations
UiPathTechnology/AIGlobal (Romanian origin)Proxy for digital resilience and innovation-driven growth narratives.Highly idiosyncratic asset; does not represent the broader traditional tech sector.
OMV PetromEnergy (Fossil)Romania/Central and Eastern Europe (CEE)Proxy for carbon transition risks and geopolitical supply vulnerabilities.Context-specific regional market structure; differs from Western European majors.
HidroelectricaEnergy (Green)RomaniaPure-play renewable benchmark for the green transition.Regulated domestic utility dynamics limit broad European generalizability.
Transilvania BankFinanceRomaniaImplementation of digital infrastructure (EU ID wallet) and social resilience.High domestic market concentration; not representative of transnational banking.
Table 4. Sentiment Classification Performance: FinBERT vs. Loughran-McDonald (Test Set = 100 Headlines).
Table 4. Sentiment Classification Performance: FinBERT vs. Loughran-McDonald (Test Set = 100 Headlines).
Metric/ModelPre-Trained FinBERTLoughran-McDonald (LM) Lexicon
Overall Accuracy89.00%61.2%
Macro-F1 Score0.880.54
Confusion Matrix (FinBERT)Predicted NegativePredicted NeutralPredicted Positive
Actual Negative (n = 40)36 (True Neg)31
Actual Neutral (n = 25)222 (True Neu)1
Actual Positive (n = 35)1331 (True Pos)
Note: The traditional LM lexicon fundamentally struggled with false negatives and misclassified neutral forward-looking statements as negative due to the rigid polarity of words like ‘risk’ or ‘exposure’, effectively validating our deployment of the Transformer architecture.
Table 5. Descriptive Summary Statistics (n = 252 observations).
Table 5. Descriptive Summary Statistics (n = 252 observations).
VariableMeanStd. Dev.MinMaxSkewnessKurtosis
ΔDAX Returns0.00040.0124−0.04120.0385−0.424.15
ΔMV Petrom Returns0.00060.0185−0.05200.0491−0.284.82
ΔUiPath Returns−0.00120.0342−0.08150.0924−0.555.34
DSI ESG Narrative0.08500.2140−0.75000.8800−0.152.85
DSI Energy Transition0.04200.2560−0.82000.9100−0.323.12
DSI Digital/AI Resilience0.11500.1850−0.55000.76000.242.65
Notes: The financial return series consistently exhibit excess kurtosis (values > 3.0), indicating leptokurtic distributions with heavy tails typical of financial market data. This specific distributional property provides empirical justification for the subsequent utilization of a Student-t distribution assumption within our GARCH(1,1) conditional volatility estimations.
Table 6. Structural Breaks and Markov-Switching Regimes.
Table 6. Structural Breaks and Markov-Switching Regimes.
Model ParameterRegime 1 (High Volatility)Regime 2 (Low Volatility)
Conditional Volatility ( σ )0.0184 ** (0.002)0.0092 ** (0.001)
Transition Probability (Pii)P11 = 0.925P22 = 0.961
Expected Regime Duration13.3 Days25.6 Days
Regime InterpretationRisk Penalty/DecouplingValue Alignment
Bai-Perron Structural Break Diagnostics: Sup-F Statistic: 18.45 ( p < 0.01 **); Estimated Breakpoint: 15.01.2026 (95% CI: [05.01.2026–25.01.2026]). Notes: Standard errors in parentheses. ** denotes statistical significance at the 1% level. P 11 and P 22 represent the transition probabilities that the market will stay in Regime 1 or Regime 2, respectively, at time t + 1 given that it is in that regime at time t . The output identifies two statistically distinct variance states associated with the estimated structural break.
Table 7. Granger Predictability Testing Results (Lag = 2).
Table 7. Granger Predictability Testing Results (Lag = 2).
Null Hypothesis (H0)F-Statisticp-ValueDecision
ESG Sentiment does not Granger-predict DAX7.8420.0004 **Reject H0
Energy Sentiment does not Granger-predict OMV5.2100.0058 **Reject H0
AI Sentiment does not Granger-predict UiPath1.1240.3274Fail to Reject
Notes: The F-Statistic evaluates the joint significance of the lagged sentiment variables for temporal precedence. ** denotes statistical significance at the 1% level. The rejection of the null hypothesis for the DAX and OMV Petrom models formally establishes that lagged narrative sentiment contains incremental predictive information regarding subsequent market movements.
Table 8. GARCH(1,1) Conditional Volatility Association Parameters.
Table 8. GARCH(1,1) Conditional Volatility Association Parameters.
ParameterGerman DAX (Exog: ESG Narrative)OMV Petrom (Exog: Energy Transition)
Mean Equation
Constant ( μ )0.0012 (0.15)0.0024 (0.22)
Variance Equation
Constant ( ω )0.0001 (0.01)0.0003 (0.02)
ARCH lag 1 ( α )0.115 (0.02) **0.142 (0.03) **
GARCH lag 1 ( β )0.820 (0.04) **0.795 (0.05) **
Sentiment Exog ( γ )−0.142 (0.04) **−0.215 (0.06) **
Model Diagnostics
Persistence ( α + β )0.935 (High, <1.0)0.937 (High, <1.0)
ConvergenceAchievedAchieved
Ljung–Box Q-test ( p -val)0.341 (No serial correl.)0.285 (No serial correl.)
ARCH-LM Test ( p -val)0.512 (No ARCH effects)0.440 (No ARCH effects)
Notes: Standard errors in parentheses. ** denotes p < 0.01.
Table 9. Complete Vector Autoregression (VAR) System Estimation.
Table 9. Complete Vector Autoregression (VAR) System Estimation.
Predictor VariableEquation 1: Dependent = Δ DAX ReturnsEquation 2: Dependent = Δ ESG Sentiment
Lag 1 Δ DAX Returns0.124 (0.045) **0.021 (0.018)
Lag 1 ESG Sentiment0.089 (0.031) **0.215 (0.040) **
Lag 2 Δ DAX Returns−0.042 (0.046)0.011 (0.019)
Lag 2 ESG Sentiment0.051 (0.032)0.085 (0.041) *
Constant0.001 (0.002)0.003 (0.005)
Post-estimation Diagnostics: VAR Stability (Eigenvalue Condition): Maximum modulus = 0.542. All inverse roots of the characteristic AR polynomial lie strictly inside the unit circle. The VAR system is stable. Residual Autocorrelation (Portmanteau Q-test): Q-stat = 12.45 (p-value = 0.38). Fail to reject the null of no serial correlation. Residual Heteroskedasticity (Breusch-Pagan LM test): LM-stat = 8.12 (p-value = 0.15). Fail to reject the null of homoskedasticity. Notes: Standard errors in parentheses. ** p < 0.01 , * p < 0.05 . The significant Lag 1 Sentiment coefficient in Equation (1) provides empirical evidence of incremental predictive information in lagged sentiment for subsequent market returns.
Table 10. Out-of-Sample Predictive Validation Metrics (1-Day-Ahead Horizon).
Table 10. Out-of-Sample Predictive Validation Metrics (1-Day-Ahead Horizon).
Model SpecificationRMSEMAEDiebold-Mariano Test (p-Value)
Baseline Benchmark: AR(2) (Without Sentiment)0.01520.0121-
Enhanced Model: VAR(2) (With FinBERT ESG Sentiment)0.01180.00942.45 (p = 0.018 *)
Notes: The Diebold-Mariano (DM) test evaluates the null hypothesis that the two competing models have equal predictive accuracy, utilizing a squared-error loss function and Newey-West standard errors. The statistically significant positive DM statistic ( p < 0.05 ) indicates that integrating the lagged NLP sentiment metric provides a statistically significant improvement in out-of-sample forecasting accuracy over the baseline AR(2) specification. * denotes statistical significance at the 5% level (p < 0.05).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Hapenciuc, C.-V.; Neamțu, D.M.; Cajvan, T.; Băeșu, C.; Bordeianu, O.-M. Textual Sentiment and Financial Market Dynamics: Econometric Evidence from Eastern Europe’s Green and Digital Transition. Sustainability 2026, 18, 9177. https://doi.org/10.3390/su18179177

AMA Style

Hapenciuc C-V, Neamțu DM, Cajvan T, Băeșu C, Bordeianu O-M. Textual Sentiment and Financial Market Dynamics: Econometric Evidence from Eastern Europe’s Green and Digital Transition. Sustainability. 2026; 18(17):9177. https://doi.org/10.3390/su18179177

Chicago/Turabian Style

Hapenciuc, Cristian-Valentin, Daniela Mihaela Neamțu, Teodora Cajvan, Camelia Băeșu, and Otilia-Maria Bordeianu. 2026. "Textual Sentiment and Financial Market Dynamics: Econometric Evidence from Eastern Europe’s Green and Digital Transition" Sustainability 18, no. 17: 9177. https://doi.org/10.3390/su18179177

APA Style

Hapenciuc, C.-V., Neamțu, D. M., Cajvan, T., Băeșu, C., & Bordeianu, O.-M. (2026). Textual Sentiment and Financial Market Dynamics: Econometric Evidence from Eastern Europe’s Green and Digital Transition. Sustainability, 18(17), 9177. https://doi.org/10.3390/su18179177

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop