1. Introduction
The quick growth of the Web3 environment has also resulted in the emergence of digital collectibles known as non-fungible tokens (NFTs) [
1]. An NFT combines various sources of information: image data of the artwork itself, descriptive information about the artwork, artist information, social information, transaction information, and network information regarding the blockchain relationships [
2]. Analyzing the interaction of the various information streams in modulating market trends has been the challenge of information analysis of Web3 platforms. The price of NFTs has been driven by nonlinear and non-stationary trends that cannot be accounted for solely through image data. Recent empirical analyses have identified several stylized facts in NFT markets that are directly relevant to our modeling choices. First, NFT price/return fluctuations are heavy-tailed and deviate from Gaussian assumptions, implying that large moves and volatility bursts are non-negligible rather than rare. Second, NFT price dynamics can exhibit long-range temporal correlations and, in certain regimes, multifractal (multi-scale) organization, indicating that both short-term shocks and persistent memory effects may coexist. Third, cross-collection and cross-instrument dependencies are time-varying and may contain non-random global modes beyond noise, suggesting that a relational structure is essential for capturing systemic co-movements. These findings motivate our design of a temporal encoder capable of modeling long-range dependencies and a heterogeneous interaction graph that explicitly represents co-movement and cross-entity relations [
3,
4].
The existing research work concerning price prediction of NFTs mostly adopts unimodal attributes from the price histories of the concerned items, image embeddings, and/or metadata attributes [
5]. The above-mentioned studies address only partial aspects of the problem because they do not consider the multimodal and relational characteristics of the involved market of heterogeneous information networks. Besides the above considerations, existing models mostly fail to consider the complex interrelated graphical structures emerging due to the flow of money, buyer–seller relationships, and ownership information in the form of contract relationships of the involved items while ignoring the co-movement of items within the same collection [
6]. In the context of information systems research, the above-described graphs represent the structural information about the information flows through the concerned market [
7].
In addition to pointwise prediction performance, information-centric risk analysis in NFT markets also lacks systematic methodologies [
8]. Web3 marketplaces are prone to abrupt price crashes, wash trading, and behavior-driven anomalies in which regional shocks can spread through address-level or collection-level graphs [
9]. In conventional financial risk analysis, variance-driven volatility and Value-at-Risk (VaR) represent the phenomena only partially because they do not model the transaction paths involving multiple steps, the whale-dominated graphs, and the intercollection linkages. It is a requirement to develop unified models that can encode NFTs and their contexts as multimodal and graph-data information entities, model the dynamics of the information entities over time, and extract risk indicators from the abstracted representations.
To overcome the above challenges, this work presents the multimodal temporal fusion and graph-structured modeling framework, a framework that combines four diverse information sources—visual embeddings, textual semantics, trading time series data, and graphs of on-chain relationships—that can be applied to the information-driven valuation and risk assessment of NFTs. The framework incorporates the fusion of temporal information at various scales through fusion modules that consider short-term market impulses and long-term behavioral patterns. The framework also utilizes the fusion of address-level and collection-level graphs through the application of graph attention layers to allow the model to reason about information flow and interasset correlations. Finally, the framework incorporates a contrastive multimodal learning scheme to improve the robustness of the framework against domain shifts from its learning environment to its testing environment.
In addition to the above representation learning tasks, this paper aims to develop a graph-based multi-source risk index (GMRI) measure of the instability of the NFT market pertaining to three sources: anomalous behavior patterns, disruptions of the relationship graphs, and overall multimodal sensitivity. The reason to consider this measure relevant to the paper’s contributions can be explained from the point of view of information system research and applications: in the context of information system research and applications, the presented work combines predictability with the structured understanding of the role of various information sources jointly impacting the valuation.
The major contributions of this work are summarized as follows:
A unified multimodal temporal graph modeling framework that integrates visual, textual, transactional, and relational information for NFT valuation, overcoming the limitations of unimodal or sequence-only information models.
A graph-structured relational learning module that captures liquidity flow, behavioral dependencies, and cross-collection co-movement through heterogeneous graph attention mechanisms, offering an information network view of NFT ecosystems.
A multimodal contrastive alignment mechanism that enhances representation consistency across modalities and improves robustness under highly non-stationary and cross-market Web3 environments.
A graph-based multi-source risk analysis index (GMRI) that provides interpretable insights into anomaly propagation, structural vulnerabilities, and multimodal drivers of market volatility, supporting risk-aware information system design.
Extensive experiments on real-world NFT datasets demonstrating significant improvements in prediction accuracy, cross-market generalization, and interpretability over strong baselines, together with detailed analyses of how different information channels contribute to valuation and risk.
In summary, this work sets up a paradigm of information modeling concerning the valuation of digital collectibles based on the combination of multimodal information, time dynamics, and graphical relationships. The model presented above advances the methods of understanding and analyzing risks relevant to the Web3 information system of digital assets.
Author Contributions
Conceptualization, Y.Y. and F.L.; methodology, Y.Y. and F.L.; software, Y.Y. and F.L.; validation, Y.Y. and F.L.; formal analysis, Y.Y. and F.L.; investigation, F.L.; resources, F.L.; data curation, F.L.; writing—original draft preparation, F.L.; writing—review and editing, J.H.; visualization, J.H.; supervision, J.H.; project administration, J.H. All authors have read and agreed to the published version of the manuscript.
Funding
This work was supported by the Beijing Language and Culture University under the University-Level Research Project “Digital-Intelligent Cultural Tourism Innovation Talent Training” (Project No. 2025HX02).
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Costa, D.; La Cava, L.; Tagarelli, A. Show me your NFT and I tell you how it will perform: Multimodal representation learning for NFT selling price prediction. In Proceedings of the ACM Web Conference 2023, Austin, TX, USA, 30 April–4 May 2023; pp. 1875–1885. [Google Scholar]
- Pala, M.; Sefer, E. NFT price and sales characteristics prediction by transfer learning of visual attributes. J. Financ. Data Sci. 2024, 10, 100148. [Google Scholar] [CrossRef] [Scilit]
- Szydło, P.; Wątorek, M.; Kwapień, J.; Drożdż, S. Characteristics of price related fluctuations in non-fungible token (NFT) market. Chaos Interdiscip. J. Nonlinear Sci. 2024, 34, 0185306. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wątorek, M.; Szydło, P.; Kwapień, J.; Drożdż, S. Correlations versus noise in the NFT market. Chaos Interdiscip. J. Nonlinear Sci. 2024, 34, 073112. [Google Scholar]
- Li, Z. Temporal Graph Neural Networks for NFT Valuation and Recommendation: A Multimodal Approach to Cold-Start and Market Dynamics. In Proceedings of the Machine Learning on Graphs in the Era of Generative Artificial Intelligence, Toronto, ON, Canada, 4 August 2025. [Google Scholar]
- Song, M.; Liu, Y.; Shah, A.; Chava, S. Abnormal trading detection in the nft market. arXiv 2023, arXiv:2306.04643. [Google Scholar] [CrossRef] [Scilit]
- Colavizza, G. Seller-buyer networks in NFT art are driven by preferential ties. Front. Blockchain 2023, 5, 1073499. [Google Scholar] [CrossRef] [Scilit]
- Upadhyay, N.; Upadhyay, S. The dark side of non-fungible tokens: Understanding risks in the NFT marketplace from a fraud triangle perspective. Financ. Innov. 2025, 11, 62. [Google Scholar] [CrossRef] [Scilit]
- Niu, Y.; Li, X.; Peng, H.; Li, W. Unveiling wash trading in popular NFT markets. In Proceedings of the Companion Proceedings of the ACM Web Conference 2024, Singapore, 13–17 May 2024; pp. 730–733. [Google Scholar]
- Kang, H.J.; Lee, S.G. Market Phases and Price Discovery in NFTs: A Deep Learning Approach to Digital Asset Valuation. J. Theor. Appl. Electron. Commer. Res. 2025, 20, 64. [Google Scholar] [CrossRef] [Scilit]
- Russell, F. NFTs and value. M/C J. 2022, 25, 2. [Google Scholar] [CrossRef] [Scilit]
- Seyhan, B.; Sefer, E. NFT primary sale price and secondary sale prediction via deep learning. In Proceedings of the Fourth ACM International Conference on AI in Finance, Brooklyn, NY, USA, 27–29 November 2023; pp. 116–123. [Google Scholar]
- Hajek, P.; Novotny, J.; Munk, M.; Munkova, D. Multimodal Financial Sentiment for Stock Return Prediction. Procedia Comput. Sci. 2025, 270, 582–591. [Google Scholar] [CrossRef] [Scilit]
- Fataliyev, K.; Liu, W. MCASP: Multi-modal cross attention network for stock market prediction. In Proceedings of the 21st Annual Workshop of the Australasian Language Technology Association, Melbourne, Australia, 29 November–1 December 2023; pp. 67–77. [Google Scholar]
- Jiang, Y.; Ning, K.; Pan, Z.; Shen, X.; Ni, J.; Yu, W.; Schneider, A.; Chen, H.; Nevmyvaka, Y.; Song, D. Multi-modal time series analysis: A tutorial and survey. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Toronto, ON, Canada, 3–7 August 2025; Volume 2, pp. 6043–6053. [Google Scholar]
- Wang, J.; Zhang, S.; Xiao, Y.; Song, R. A review on graph neural network methods in financial applications. arXiv 2021, arXiv:2111.15367. [Google Scholar] [CrossRef] [Scilit]
- Hu, L.; Wang, Q. A Study of Dynamic Stock Relationship Modeling and S&P500 Price Forecasting Based on Differential Graph Transformer. arXiv 2025, arXiv:2506.18717. [Google Scholar] [CrossRef] [Scilit]
- Song, J.; Zhang, S.; Zhang, P.; Park, J.; Gu, Y.; Yu, G. Illicit Social Accounts? Anti-Money Laundering for Transactional Blockchains. IEEE Trans. Inf. Forensics Secur. 2024, 20, 391–404. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Chan, S.; Chu, J.; Sulieman, H. On the market efficiency and liquidity of high-frequency cryptocurrencies in a bull and bear market. J. Risk Financ. Manag. 2020, 13, 8. [Google Scholar] [CrossRef] [Scilit]
- Tošić, A.; Vičič, J.; Hrovatin, N. Beyond the surface: Advanced wash-trading detection in decentralized NFT markets. Financ. Innov. 2025, 11, 1–21. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Zhu, Y.; Wang, G.J.; Xie, C.; Wang, Q. Risk contagion of NFT: A time-frequency risk spillover perspective in the Carbon-NFT-Stock system. Financ. Res. Lett. 2024, 59, 104765. [Google Scholar] [CrossRef] [Scilit]
- Su, X.; Yan, X.; Tsai, C.L. Linear regression. Wiley Interdiscip. Rev. Comput. Stat. 2012, 4, 275–294. [Google Scholar] [CrossRef] [Scilit]
- Rigatti, S.J. Random forest. J. Insur. Med. 2017, 47, 31–39. [Google Scholar] [CrossRef] [Scilit]
- Chen, T.; Guestrin, C. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar]
- Koonce, B. ResNet 50. In Convolutional Neural Networks with Swift for Tensorflow: Image Recognition and Dataset Categorization; Springer: Berlin/Heidelberg, Germany, 2021; pp. 63–72. [Google Scholar]
- Koroteev, M.V. BERT: A review of applications in natural language processing and understanding. arXiv 2021, arXiv:2103.11943. [Google Scholar] [CrossRef] [Scilit]
- Luo, D.; Wang, X. Moderntcn: A modern pure convolution structure for general time series analysis. In Proceedings of the Twelfth International Conference on Learning Representations, Vienna, Austria, 7–11 May 2024; pp. 1–43. [Google Scholar]
- Lim, B.; Arık, S.Ö.; Loeff, N.; Pfister, T. Temporal fusion transformers for interpretable multi-horizon time series forecasting. Int. J. Forecast. 2021, 37, 1748–1764. [Google Scholar] [CrossRef] [Scilit]
- Joseph, S.; Parthi, A.G.; Maruthavanan, D.; Veerapaneni, P.K.; Jayaram, V.; Pothineni, B. A Concatenation-Based Convolutional Network. In Proceedings of the 2024 4th International Conference on Robotics, Automation and Artificial Intelligence (RAAI), Singapore, 19–21 December 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 386–391. [Google Scholar]
- Tsai, Y.H.H.; Bai, S.; Liang, P.P.; Kolter, J.Z.; Morency, L.P.; Salakhutdinov, R. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; Volume 2019, p. 6558. [Google Scholar]
- Kipf, T. Semi-supervised classification with graph convolutional networks. arXiv 2016, arXiv:1609.02907. [Google Scholar]
- Velickovic, P.; Cucurull, G.; Casanova, A.; Romero, A.; Lio, P.; Bengio, Y. Graph attention networks. Stat 2017, 1050, 10–48550. [Google Scholar]
- Rossi, E.; Chamberlain, B.; Frasca, F.; Eynard, D.; Monti, F.; Bronstein, M. Temporal graph networks for deep learning on dynamic graphs. arXiv 2020, arXiv:2006.10637. [Google Scholar] [CrossRef] [Scilit]
- Jiang, B.; Zhang, Z.; Lin, D.; Tang, J.; Luo, B. Semi-supervised learning with graph learning-convolutional networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 11313–11320. [Google Scholar]
- Bi, W.; Du, L.; Fu, Q.; Wang, Y.; Han, S.; Zhang, D. Mm-gnn: Mix-moment graph neural network towards modeling neighborhood feature distribution. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining, Singapore, 27 February–3 March 2023; pp. 132–140. [Google Scholar]
- Yang, R.; Yang, B.; Ouyang, S.; She, T.; Feng, A.; Jiang, Y.; Lecue, F.; Lu, J.; Li, I. Graphusion: Leveraging large language models for scientific knowledge graph fusion and construction in nlp education. arXiv 2024, arXiv:2407.10794. [Google Scholar] [CrossRef] [Scilit]
Figure 1.
The overall architecture of the proposed framework.
Figure 2.
Time–frequency characterization of NFT price dynamics. Top: normalized historical log-price series for a representative NFT over time. Bottom: continuous wavelet scalogram (CWT) of the same series, where the x-axis is time, the y-axis corresponds to wavelet scales (mapped to frequency bands), and color intensity indicates spectral energy. High energy at low-frequency bands reflects long-horizon market cycles, while intermittent high-frequency bursts indicate short-term liquidity shocks and speculative fluctuations. (Top): normalized NFT price trajectory exhibiting a mixture of long-horizon cycles and short-term volatility bursts. (Bottom): continuous wavelet scalogram showing the time–frequency decomposition.
Figure 3.
Multimodal explainability and risk visualization. (a) Price trajectory overlaid with temporal saliency scores (normalized to [0, 1]), highlighting time segments that contribute most to valuation. (b) Relation-specific attention heatmap from the graph attention module, where each cell indicates the normalized attention weight assigned to a neighbor under a given relation type, revealing structurally influential counterparts (e.g., shared creator, co-transaction, wallet relation). (c) SHAP-based modality-level attribution showing the relative contribution of image, text, time series, and blockchain features aggregated over the test set. (d) Multi-source risk index over time, demonstrating early-warning behavior prior to major price drawdowns.
Figure 4.
(a) Prediction error across NFT collections. Boxplots with swarm overlays show the MAE distribution for five representative NFT collections (Art, Pixel, 3D, Photography, and Anime). (b) Prediction error stratified by risk tier.
Figure 5.
(a) Predicted vs. true NFT log-prices. Each point corresponds to an NFT, colored by its multi-source risk index. The strong alignment with the diagonal demonstrates high valuation accuracy, with high-risk assets exhibiting larger deviations. (b) Multimodal similarity vs. prediction divergence. Each point represents a quasi-symmetric NFT pair . Higher similarity leads to smaller prediction gaps, revealing that the symmetry-preserving constraint effectively enforces valuation consistency across structurally related assets.
Table 1.
Summary of representative NFT/Web3 valuation studies and the positioning of this work.
| Category | Representative Line of Work | Main Assumption(s) | Advantages | Limitations (What Is Missing) | Our Position |
|---|
| Statistical/classical ML on price series | AR/linear/XGBoost on historical prices | Past prices dominate; weak cross-asset coupling | Simple, efficient | Cannot capture multimodal signals; weak under drift | We add multimodal + graph + temporal reasoning |
| Visual-only NFT valuation | CNN/ViT aesthetics/rarity | Visual style correlates with value | Captures intrinsic appearance | Ignores text, market dynamics, and on-chain relations | Visual is one branch in our unified model |
| Text-only/metadata-driven | BERT-based description encoding | Text narrative drives value | Captures semantics and creator cues | Weak on market dynamics; no relational propagation | Text is aligned with other modalities (contrastive) |
| Shallow multimodal fusion | Concat/late-fusion of img+txt | Simple complementarity is enough | Easy to implement | Fusion is brittle; ignores temporal non-stationarity | We use gated fusion + alignment + temporal modeling |
| Graph-only blockchain analytics | GCN/GAT on address graphs | Structure dominates outcomes | Models neighborhood effects | Often static; weak temporal forecasting; ignores content | We use heterogeneous relations + temporal–structural fusion |
| Temporal forecasting models | Informer/ Autoformer/ PatchTST | Time series carries most signal | Strong long-range modeling | No multimodal content; no heterogeneous relations | We integrate temporal modeling with multimodal and graph |
| Risk/anomaly detection in Web3 | Heuristic/forensics anomaly detectors | Abnormality detectable from patterns | Useful for alerts | Not unified with valuation; weak explanation stability | We unify valuation + risk index + interpretability metrics |
Table 2.
Dataset statistics for MultiNFT-T and multimodal NFT benchmarks. We report the covered time span, collection distribution, price statistics, and modality missingness rates.
| Item | MultiNFT-T | NFT Marketplace Analytics |
|---|
| Time span | January 2021–December 2023 | June 2020–June 2024 |
| #NFTs/#Transactions | 312,450/1,870,000 | 420,000/3,950,000 |
| #Collections | 18,632 | 26,410 |
| Top-10 collection share | 11.8% | 14.5% |
| Price (log) min/median/p95/max | 0.120/5.010/8.420/12.260 | −0.050/5.180/8.710/12.900 |
| Missing rate (Image/Text/TS/On-chain) | 6.70%/28.40%/0.00%/3.20% | 4.00%/22.00%/0.00%/2.50% |
Table 3.
Architectural configuration of MM-Temporal-Graph for reproducibility.
| Module | Backbone/Block | Depth | Key Dimensions | Heads | Dropout | Notes |
|---|
| Image encoder | ResNet-50 or ViT-B/16 | 50/12 | proj to | – | 0.1 | pretrained; fine-tune last stage |
| Text encoder | BERT-base | 12 | 768 → 256 | 12 | 0.1 | max length 128 |
| Temporal encoder | Transformer encoder | | , FFN | 8 | 0.1 | window 64/128 |
| Relational GAT | relation-aware GAT | | | 8 | 0.1 | per-relation attention |
| Global fusion transformer | Transformer over nodes | | , FFN | 8 | 0.1 | cross-asset dependencies |
| Fusion gate | MLP + sigmoid | 2 | 512→256→1 | – | 0.1 | adaptive local/global mixing |
| Valuation head | MLP | 2 | 256→128→1 | – | 0.1 | regression |
Table 4.
Hyperparameter settings for the proposed framework.
| Hyperparameter | Value |
|---|
| Optimizer | AdamW |
| Learning rate | |
| Batch size | 64 |
| Training epochs | 120 (with early stopping) |
| Hidden dimension | 256 |
| Dropout rate | 0.3 |
| Number of graph attention heads | 8 |
| Number of transformer heads | 8 |
| Weight decay | |
| Graph consistency coefficient | 0.05 |
Table 5.
Overall performance on three NFT-related datasets (mean ± std over 5 seeds). Lower MAE/RMSE and higher indicate better performance. Bold numbers denote the best result for each metric in each dataset.
| Model | NFT Marketplace Analytics | MultiNFT-T |
|---|
| MAE↓ | RMSE↓ | ↑ | MAE↓ | RMSE↓ | ↑ |
|---|
| Linear Regression [22] | 0.228 ± 0.006 | 0.311 ± 0.008 | 0.702 ± 0.010 | 0.214 ± 0.006 | 0.296 ± 0.007 | 0.721 ± 0.009 |
| Random Forest [23] | 0.214 ± 0.005 | 0.297 ± 0.007 | 0.728 ± 0.009 | 0.201 ± 0.005 | 0.282 ± 0.006 | 0.748 ± 0.008 |
| XGBoost [24] | 0.208 ± 0.005 | 0.289 ± 0.007 | 0.739 ± 0.008 | 0.193 ± 0.004 | 0.271 ± 0.006 | 0.765 ± 0.008 |
| ResNet-50 (image) [25] | 0.203 ± 0.004 | 0.284 ± 0.006 | 0.747 ± 0.008 | 0.189 ± 0.004 | 0.268 ± 0.006 | 0.772 ± 0.007 |
| BERT-Regressor (text) [26] | 0.205 ± 0.004 | 0.286 ± 0.006 | 0.743 ± 0.008 | 0.192 ± 0.004 | 0.270 ± 0.006 | 0.768 ± 0.007 |
| ModernTCN (time series only) [27] | 0.199 ± 0.004 | 0.279 ± 0.006 | 0.756 ± 0.007 | 0.181 ± 0.003 | 0.261 ± 0.005 | 0.785 ± 0.006 |
| TFT (multimodal covariates) [28] | 0.196 ± 0.004 | 0.275 ± 0.006 | 0.762 ± 0.007 | 0.178 ± 0.003 | 0.257 ± 0.005 | 0.792 ± 0.006 |
| ConcatNet (img+txt) [29] | 0.194 ± 0.004 | 0.273 ± 0.006 | 0.765 ± 0.007 | 0.176 ± 0.003 | 0.254 ± 0.005 | 0.796 ± 0.006 |
| MM-Transformer [30] | 0.193 ± 0.004 | 0.272 ± 0.006 | 0.768 ± 0.007 | 0.175 ± 0.003 | 0.253 ± 0.005 | 0.799 ± 0.006 |
| GCN (graph only) [31] | 0.206 ± 0.004 | 0.287 ± 0.006 | 0.745 ± 0.007 | 0.188 ± 0.003 | 0.269 ± 0.005 | 0.770 ± 0.006 |
| GAT [32] | 0.198 ± 0.004 | 0.278 ± 0.006 | 0.758 ± 0.007 | 0.180 ± 0.003 | 0.260 ± 0.005 | 0.786 ± 0.006 |
| TGN (temporal graph) [33] | 0.184 ± 0.003 | 0.263 ± 0.005 | 0.784 ± 0.006 | 0.166 ± 0.002 | 0.245 ± 0.004 | 0.820 ± 0.005 |
| GLCN [34] | 0.192 ± 0.004 | 0.271 ± 0.006 | 0.769 ± 0.007 | 0.173 ± 0.003 | 0.252 ± 0.005 | 0.801 ± 0.006 |
| MM-GNN [35] | 0.187 ± 0.003 | 0.266 ± 0.005 | 0.778 ± 0.006 | 0.168 ± 0.002 | 0.247 ± 0.004 | 0.812 ± 0.005 |
| MTGNN (graph-temporal) [30] | 0.183 ± 0.003 | 0.262 ± 0.005 | 0.785 ± 0.006 | 0.165 ± 0.002 | 0.244 ± 0.004 | 0.821 ± 0.005 |
| GraphFusion-XL [36] | 0.181 ± 0.003 | 0.260 ± 0.005 | 0.787 ± 0.006 | 0.162 ± 0.002 | 0.241 ± 0.004 | 0.823 ± 0.005 |
| Ours (MM-Temporal-Graph) | 0.172 ± 0.002 | 0.251 ± 0.004 | 0.804 ± 0.004 | 0.153 ± 0.002 | 0.232 ± 0.003 | 0.841 ± 0.004 |
| p-value (Ours vs. GraphFusion-XL) | 0.006 | 0.009 | 0.008 | 0.0007 | 0.0009 | 0.0008 |
Table 6.
Validation of GMRI against standard risk metrics.
| Risk Score | Spearman ↑ | AUC ↑ | RDA (%) ↑ |
|---|
| Realized Volatility (RV) | 0.31 | 0.71 | 66.4 |
| Historical VaR ( = 0.05) | 0.28 | 0.69 | 64.9 |
| Max Drawdown (MDD) | 0.35 | 0.73 | 68.1 |
| GMRI (full) | 0.48 | 0.84 | 78.3 |
| GMRI w/o temporal instability | 0.39 | 0.77 | 72.4 |
| GMRI w/o structural exposure | 0.41 | 0.79 | 73.1 |
| GMRI w/o feature uncertainty | 0.37 | 0.75 | 71.0 |
Table 7.
The comparison results with standard risk metrics.
| Risk Score | Spearman ↑ | AUC ↑ | RDA (%) ↑ |
|---|
| Realized Volatility (RV) | 0.31 | 0.71 | 66.4 |
| Historical VaR () | 0.28 | 0.69 | 64.9 |
| Max Drawdown (MDD) | 0.35 | 0.73 | 68.1 |
| GMRI (ours) | 0.48 | 0.84 | 78.3 |
Table 8.
Ablation study of the proposed framework on the MultiNFT-T dataset. We report valuation accuracy and risk-related performance when removing different components.
| Model Variant | MAE ↓ | RMSE ↓ | ↑ | RDA (%) ↑ |
|---|
| Full model (ours) | 0.153 | 0.232 | 0.841 | 87.4 |
| w/o time series branch | 0.167 | 0.247 | 0.815 | 78.9 |
| w/o visual modality | 0.161 | 0.241 | 0.826 | 84.2 |
| w/o textual modality | 0.160 | 0.239 | 0.829 | 83.7 |
| w/o blockchain behavioral features | 0.159 | 0.238 | 0.831 | 81.5 |
| w/o heterogeneous graph (no GNN) | 0.171 | 0.252 | 0.808 | 75.6 |
| w/o temporal transformer (ModernTCN only) | 0.165 | 0.245 | 0.819 | 80.1 |
| w/o risk-aware regularizer () | 0.157 | 0.236 | 0.835 | 71.3 |
| Early fusion (no adaptive graph–temporal gate) | 0.160 | 0.240 | 0.828 | 82.0 |
Table 9.
Sensitivity analysis of key hyperparameters on the MultiNFT-T validation set.
| Window Length | Graph Heads | MAE ↓ | ↑ | RDA (%) ↑ |
|---|
| 0.00 | 64 | 8 | 0.157 | 0.835 | 71.3 |
| 0.02 | 64 | 8 | 0.155 | 0.838 | 79.6 |
| 0.05 | 64 | 8 | 0.153 | 0.841 | 87.4 |
| 0.10 | 64 | 8 | 0.156 | 0.837 | 86.1 |
| 0.05 | 32 | 8 | 0.159 | 0.829 | 82.7 |
| 0.05 | 64 | 8 | 0.153 | 0.841 | 87.4 |
| 0.05 | 128 | 8 | 0.154 | 0.839 | 86.9 |
| 0.05 | 64 | 4 | 0.156 | 0.836 | 84.5 |
| 0.05 | 64 | 8 | 0.153 | 0.841 | 87.4 |
| 0.05 | 64 | 12 | 0.154 | 0.840 | 87.0 |
Table 10.
Computational and communication overhead comparison. Params and FLOPs are measured per forward pass on a batch of 64 NFTs.
| Model | Params (M) | FLOPs (G) | GPU Mem (GB) | Infer. Time/1k NFTs (ms) |
|---|
| ModernTCN (time series only) | 8.3 | 11.5 | 2.1 | 21 |
| MM-GNN | 18.7 | 22.9 | 4.3 | 37 |
| GraphFusion-XL | 24.5 | 29.8 | 5.6 | 45 |
| Ours (MM-Temporal-Graph, small) | 21.2 | 26.4 | 4.9 | 42 |
| Ours (MM-Temporal-Graph, base) | 27.9 | 33.7 | 6.1 | 49 |
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |