HierFinRAG—Hierarchical Multimodal RAG for Financial Document Understanding
Abstract
1. Introduction
- Table-Text Graph Neural Network (TTGNN): Instead of splitting the document into independent chunks, we build a graph that connects related parts. We model sections, paragraphs, tables, and cells as nodes in this graph. We create links (edges) based on the document hierarchy (e.g., a table belongs to a section) and explicit cross-references (e.g., text saying “See Table 5”). This allows our system to “read” the document like an expert, jumping from a summary to the supporting data table to verify facts.
- Symbolic–Neural Fusion Reasoning: LLMs are great at reading but poor at math. We use a hybrid approach: if a question asks for a summary, the LLM answers it. If the question requires calculation (e.g., “What is the percentage growth?”), our system identifies the correct numbers and sends them to a calculator engine. This ensures the math is always correct while keeping the flexible language capabilities of the LLM.
- We propose a Structure-Aware graph approach that explicitly models accounting relationships, unlike generic document graphs.
- We introduce a Probabilistic Routing mechanism that strictly separates calculation from text generation, preventing calculation errors.
- We achieve a new state-of-the-art accuracy of 82.5% on the FinQA benchmark.
- We provide a comprehensive evaluation showing our method uses 40% fewer tokens than standard methods by retrieving only the relevant data.
2. Related Work
2.1. Retrieval-Augmented Generation Systems
2.2. Multimodal Document Understanding
2.3. Financial Document Analysis and NLP
2.4. Table Understanding and Reasoning
2.5. Graph Neural Networks for Document Processing
2.6. Symbolic–Neural Hybrid Reasoning
2.7. Attribution and Explainability in RAG
2.8. Evaluation and Benchmarking
2.9. Positioning of Our Work
- Unified Multimodal Architecture: First hierarchical RAG specifically designed for financial documents, jointly modeling tables, text, cross-references, and accounting constraints through TTGNN.
- Symbolic–Neural Fusion: Introduces reasoning router and constraint checking with accounting identities, going beyond pure neural or pure symbolic approaches.
- Hierarchical Retrieval: Three-level attention mechanism (document → section → cell) reduces search space by 1000× while maintaining accuracy.
- Verifiable Attribution: Cell-level and sentence-level evidence tracking with reasoning chain extraction, suitable for regulatory compliance.
- Comprehensive Evaluation: Introduces FinTable-X benchmark with 5000 questions and evaluates on 7 datasets with 15 metrics, including novel attribution metrics.
3. Methods
| Algorithm 1 Symbolic–Neural Fusion Routing. |
| Require: Query q, Retrieved Context |
| Ensure: Answer A |
| 1: ▹ e.g., {Table, Text} |
| 2: |
| 3: |
| 4: if then |
| 5: |
| 6: |
| 7: ▹ Symbolic Mode |
| 8: else if then |
| 9: ▹ Neural Mode |
| 10: else |
| 11: |
| 12: |
| 13: ▹ Hybrid Mode |
| 14: end ifreturn A |
3.1. Problem Formulation
3.2. Table-Text Graph Construction
3.2.1. Node Representation
3.2.2. Edge Formation
- Structural Edges (): Deterministic links reflecting the document hierarchy. We connect each cell to its column header and row header . Tables and paragraphs are linked to their parent Section .
- Semantic Edges (): Established between a paragraph and a table row to capture implicit alignment (e.g., a paragraph discussing “revenue growth” implicitly linking to the “2023 Revenue” row). We create an edge if , where determines the density of semantic connectivity.
- Cross-Reference Edges (): Explicit links detected via high-precision regex matching (e.g., patterns like “(see Table \d+)” or “As shown in Figure \d+”).
3.3. Table-Text Graph Neural Network (TTGNN)
3.4. Hierarchical Attention Retrieval
3.5. Symbolic–Neural Fusion
4. Experimental Results
4.1. Comparative Analysis on Financial Benchmarks
4.2. Retrieval Efficacy
4.3. Efficiency and Deployment Feasibility
4.4. Ablation Study
4.5. Analysis of Probabilistic Routing
4.6. Qualitative Case Study
4.7. Cost Analysis
4.8. Error Distribution and Limitations
5. Discussion
5.1. The Necessity of Hybrid Reasoning
5.2. Structure as a First-Class Citizen
5.3. Implications for Agentic Workflows
5.4. Limitations and Future Work
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Chen, Z.; Chen, W.; Smiley, C.; Shah, S.; Borova, I.; Langdon, D.; Moussa, R.; Beane, M.; Huang, T.H.; Routledge, B.R.; et al. Finqa: A dataset of numerical reasoning over financial data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Punta Cana, Dominican Republic, 2021; pp. 3697–3711. [Google Scholar]
- Islam, P.; Kannappan, A.; Kiela, D.; Qian, R.; Scherrer, N.; Vidgen, B. Financebench: A new benchmark for financial question answering. arXiv 2023, arXiv:2311.11944. [Google Scholar] [CrossRef] [Scilit]
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.T.; Rocktäschel, T.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 9459–9474. [Google Scholar]
- Karpukhin, V.; Oğuz, B.; Min, S.; Lewis, P.; Wu, L.; Edunov, S.; Chen, D.; Yih, W.T. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP); Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 6769–6781. [Google Scholar]
- Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; Zettlemoyer, L. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, 5–10 July 2020; Jurafsky, D., Chai, J., Schluter, N., Tetreault, J.R., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 7871–7880. [Google Scholar] [CrossRef] [Scilit]
- Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Wang, H.; Wang, H. Retrieval-augmented generation for large language models: A survey. arXiv 2023, arXiv:2312.10997. [Google Scholar]
- Gao, L.; Ma, X.; Lin, J.; Callan, J. Precise zero-shot dense retrieval without relevance labels. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 1762–1777. [Google Scholar]
- Sarthi, P.; Abdullah, S.; Tuli, A.; Khanna, S.; Goldie, A.; Manning, C.D. Raptor: Recursive abstractive processing for tree-organized retrieval. In Proceedings of the ICLR, Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Dang, Q.-V.; Nguyen, N.-S.-A. Predicting the Stock Price Using a Bayesian Graph Neural Networks-Based Architecture. In International Conference on Future Data and Security Engineering, Chi Minh City, Vietnam, 27–29 November 2025; Springer: Singapore, 2025; pp. 330–345. [Google Scholar]
- Asai, A.; Wu, Z.; Wang, Y.; Sil, A.; Hajishirzi, H. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. In Proceedings of the Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, 7–11 May 2024; OpenReview.net: Alameda, CA, USA, 2024. [Google Scholar]
- Singh, A.; Ehtesham, A.; Kumar, S.; Khoei, T.T. Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG. arXiv 2025, arXiv:2501.09136. [Google Scholar] [CrossRef] [Scilit]
- Yan, S.Q.; Gu, J.C.; Zhu, Y.; Ling, Z.H. Corrective Retrieval Augmented Generation. arXiv 2024, arXiv:2401.15884. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Zhang, W.; Yang, Y.; Huang, W.C.; Wu, Y.; Luo, J.; Bei, Y.; Zou, H.P.; Luo, X.; Zhao, Y.; et al. Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs. arXiv 2025, arXiv:2507.09477. [Google Scholar] [CrossRef] [Scilit]
- Guo, D.; Yang, D.; Zhang, H.; Song, J.; Wang, P.; Zhu, Q.; Xu, R.; Zhang, R.; Ma, S.; Bi, X.; et al. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. Nature 2025, 645, 633–638. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Dong, G.; Jin, J.; Zhang, Y.; Zhou, Y.; Zhu, Y.; Zhang, P.; Dou, Z. Search-o1: Agentic search-enhanced large reasoning models. arXiv 2025, arXiv:2501.05366. [Google Scholar]
- Jin, B.; Zeng, H.; Yue, Z.; Yoon, J.; Arik, S.; Wang, D.; Zamani, H.; Han, J. Search-r1: Training llms to reason and leverage search engines with reinforcement learning. arXiv 2025, arXiv:2503.09516. [Google Scholar] [CrossRef] [Scilit]
- Faysse, M.; Sibille, H.; Wu, T.; Omrani, B.; Viaud, G.; Hudelot, C.; Colombo, P. ColPali: Efficient Document Retrieval with Vision Language Models. In Proceedings of the Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, 24–28 April 2025; OpenReview.net: Alameda, CA, USA, 2025. [Google Scholar]
- Wang, P.; Bai, S.; Tan, S.; Wang, S.; Fan, Z.; Bai, J.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv 2024, arXiv:2409.12191. [Google Scholar]
- Mei, L.; Mo, S.; Yang, Z.; Chen, C. A survey of multimodal retrieval-augmented generation. arXiv 2025, arXiv:2504.08748. [Google Scholar] [PubMed]
- Xia, P.; Zhu, K.; Li, H.; Wang, T.; Shi, W.; Wang, S.; Zhang, L.; Zou, J.; Yao, H. MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models. In Proceedings of the Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, 24–28 April 2025; OpenReview.net: Alameda, CA, USA, 2025. [Google Scholar]
- Mao, M.; Perez-Cabarcas, M.M.; Kallakuri, U.; Waytowich, N.R.; Lin, X.; Mohsenin, T. Multi-RAG: A Multimodal Retrieval-Augmented Generation System for Adaptive Video Understanding. arXiv 2025, arXiv:2505.23990. [Google Scholar]
- Wan, X.; Yu, H. MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs. arXiv 2025, arXiv:2507.20804. [Google Scholar] [CrossRef] [Scilit]
- Yuan, X.; Ning, L.; Fan, W.; Li, Q. mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering. arXiv 2025, arXiv:2508.05318. [Google Scholar]
- Zhou, Y.; Su, Y.; Sun, Y.; Wang, S.; Wang, T.; He, R.; Zhang, Y.; Liang, S.; Liu, X.; Ma, Y.; et al. In-depth Analysis of Graph-based RAG in a Unified Framework. Proc. VLDB Endow. 2025, 18, 5623–5637. [Google Scholar] [CrossRef] [Scilit]
- Wu, S.; Irsoy, O.; Lu, S.; Dabravolski, V.; Dredze, M.; Gehrmann, S.; Kambadur, P.; Rosenberg, D.; Mann, G. Bloomberggpt: A large language model for finance. arXiv 2023, arXiv:2303.17564. [Google Scholar] [CrossRef] [Scilit]
- Liu, X.Y.; Wang, G.; Yang, H.; Zha, D. Fingpt: Democratizing internet-scale data for financial large language models. arXiv 2023, arXiv:2307.10485. [Google Scholar]
- Chen, Z.; Li, S.; Smiley, C.; Ma, Z.; Shah, S.; Wang, W.Y. Convfinqa: Exploring the chain of numerical reasoning in conversational finance question answering. arXiv 2022, arXiv:2210.03849. [Google Scholar]
- Alam, M.Z.; Zaman, K.A.U.; Miraz, M.H. AstuteRAG-FQA: Task-Aware Retrieval-Augmented Generation Framework for Proprietary Data Challenges in Financial Question Answering. arXiv 2025, arXiv:2510.27537. [Google Scholar] [CrossRef] [Scilit]
- Zha, Y. SMARTFinRAG: Interactive Modularized Financial RAG Benchmark. arXiv 2025, arXiv:2504.18024. [Google Scholar] [CrossRef] [Scilit]
- Dadopoulos, M.; Ladas, A.; Moschidis, S.; Negkakis, I. Metadata-Driven Retrieval-Augmented Generation for Financial Question Answering. arXiv 2025, arXiv:2510.24402. [Google Scholar]
- Elahi, A. Identifying Financial Risk Information Using RAG with a Contrastive Insight. arXiv 2025, arXiv:2510.03521. [Google Scholar] [CrossRef] [Scilit]
- Kothandapani, H.P. AI-Driven Regulatory Compliance: Transforming Financial Oversight through Large Language Models and Automation. Emerg. Sci. Res. 2025, 3, 12–24. [Google Scholar]
- Singh, G.; Singh, P.; Singh, M. Advanced Real-Time Fraud Detection Using RAG-Based LLMs. arXiv 2025, arXiv:2501.15290. [Google Scholar]
- Yin, P.; Neubig, G.; Yih, W.T.; Riedel, S. TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 8413–8426. [Google Scholar]
- Herzig, J.; Nowak, P.K.; Müller, T.; Piccinno, F.; Eisenschlos, J. TaPas: Weakly supervised table parsing via pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 4320–4333. [Google Scholar]
- Liu, Q.; Chen, B.; Guo, J.; Ziyadi, M.; Lin, Z.; Chen, W.; Lou, J.G. TAPEX: Table pre-training via learning a neural SQL executor. In Proceedings of the International Conference on Learning Representations, Virtual Event, 25–29 April 2022. [Google Scholar]
- Zhu, F.; Lei, W.; Huang, Y.; Wang, C.; Zhang, S.; Lv, J.; Feng, F.; Chua, T.S. TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 3277–3287. [Google Scholar]
- Zhao, Y.; Li, Y.; Li, C.; Zhang, R. MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual Data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 6588–6600. [Google Scholar]
- Zhao, Y.; Nan, L.; Qi, Z.; Zhang, R.; Radev, D. ReasTAP: Injecting table reasoning skills during pre-training via synthetic reasoning examples. arXiv 2022, arXiv:2210.12374. [Google Scholar]
- Edge, D.; Trinh, H.; Cheng, N.; Bradley, J.; Chao, A.; Mody, A.; Truitt, S.; Larson, J. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv 2024, arXiv:2404.16130. [Google Scholar] [CrossRef] [Scilit]
- Guo, Z.; Xia, L.; Yu, Y.; Ao, T.; Huang, C. LightRAG: Simple and Fast Retrieval-Augmented Generation. arXiv 2024, arXiv:2410.05779. [Google Scholar]
- Hu, Z.; Dong, Y.; Wang, K.; Sun, Y. Heterogeneous Graph Transformer. In Proceedings of the Web Conference 2020; Association for Computing Machinery: New York, NY, USA, 2020; pp. 2704–2710. [Google Scholar]
- Schlichtkrull, M.S.; Kipf, T.N.; Bloem, P.; van den Berg, R.; Titov, I.; Welling, M. Modeling Relational Data with Graph Convolutional Networks. In Proceedings of the Semantic Web—15th International Conference, ESWC 2018, Heraklion, Crete, Greece, 3–7 June 2018; Gangemi, A., Navigli, R., Vidal, M., Hitzler, P., Troncy, R., Hollink, L., Tordai, A., Alam, M., Eds.; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2018; Volume 10843, pp. 593–607. [Google Scholar] [CrossRef] [Scilit]
- Veličković, P.; Cucurull, G.; Casanova, A.; Romero, A.; Liò, P.; Bengio, Y. Graph Attention Networks. In Proceedings of the International Conference on Learning Representations; OpenReview.net: Alameda, CA, USA, 2018. [Google Scholar]
- Brody, S.; Alon, U.; Yahav, E. How Attentive are Graph Attention Networks? In Proceedings of the Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, 25–29 April 2022; OpenReview.net: Alameda, CA, USA, 2022. [Google Scholar]
- Xu, Y.; Li, M.; Cui, L.; Huang, S.; Wei, F.; Zhou, M. LayoutLM: Pre-training of Text and Layout for Document Image Understanding. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining; Association for Computing Machinery: New York, NY, USA, 2020; pp. 1192–1200. [Google Scholar]
- Huang, Y.; Lv, T.; Cui, L.; Lu, Y.; Wei, F. LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking. In Proceedings of the 30th ACM International Conference on Multimedia; Association for Computing Machinery: New York, NY, USA, 2022; pp. 4083–4091. [Google Scholar]
- Appalaraju, S.; Jasani, B.; Kota, B.U.; Xie, Y.; Manmatha, R. DocFormer: End-to-End Transformer for Document Understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2021; pp. 993–1003. [Google Scholar]
- d’Avila Garcez, A.; Lamb, L.C. Neurosymbolic AI: The 3rd wave. Artif. Intell. Rev. 2023, 56, 12387–12406. [Google Scholar] [CrossRef] [Scilit]
- Andreas, J.; Rohrbach, M.; Darrell, T.; Klein, D. Neural Module Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2016; pp. 39–48. [Google Scholar]
- Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Zettlemoyer, L.; Cancedda, N.; Scialom, T. Toolformer: Language Models Can Teach Themselves to Use Tools. arXiv 2024, arXiv:2302.04761. [Google Scholar]
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.R.; Cao, Y. ReAct: Synergizing Reasoning and Acting in Language Models. In Proceedings of the Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, 1–5 May 2023; OpenReview.net: Alameda, CA, USA, 2023. [Google Scholar]
- Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q.V.; Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Adv. Neural Inf. Process. Syst. 2022, 35, 24824–24837. [Google Scholar]
- Chen, W.; Ma, X.; Wang, X.; Cohen, W.W. Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks. Trans. Assoc. Comput. Linguist. 2023, 11, 155–176. [Google Scholar]
- Gao, L.; Madaan, A.; Zhou, S.; Alon, U.; Liu, P.; Yang, Y.; Callan, J.; Neubig, G. PAL: Program-aided Language Models. In Proceedings of the International Conference on Machine Learning, ICML 2023, Honolulu, HI, USA, 23–29 July 2023; Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J., Eds.; Proceedings of Machine Learning Research; PMLR: Cambridge, MA, USA, 2023; Volume 202, pp. 10764–10799. [Google Scholar]
- Imani, S.; Du, L.; Shrivastava, H. MathPrompter: Mathematical Reasoning using Large Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics: Industry Track, ACL 2023, Toronto, ON, Canada, 9–14 July 2023; Sitaram, S., Klebanov, B.B., Williams, J.D., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 37–42. [Google Scholar] [CrossRef] [Scilit]
- Wallat, J.; Heuss, M.; de Rijke, M.; Anand, A. Correctness is not Faithfulness in Retrieval Augmented Generation Attributions. In Proceedings of the 2025 International ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval, ICTIR 2025, Padua, Italy, 18 July 2025; Zamani, H., Dietz, L., Piwowarski, B., Bruch, S., Eds.; ACM: New York, NY, USA, 2025; pp. 22–32. [Google Scholar] [CrossRef] [Scilit]
- Niu, C.; Wu, Y.; Zhu, J.; Xu, S.; Shum, K.; Zhong, R.; Song, J.; Zhang, T. Ragtruth: A hallucination corpus for developing trustworthy retrieval-augmented language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 10862–10878. [Google Scholar]
- Hu, H.; He, C.; Xie, X.; Zhang, Q. Lrp4rag: Detecting hallucinations in retrieval-augmented generation via layer-wise relevance propagation. arXiv 2024, arXiv:2408.15533. [Google Scholar]
- Joren, H.; Zhang, J.; Ferng, C.; Juan, D.; Taly, A.; Rashtchian, C. Sufficient Context: A New Lens on Retrieval Augmented Generation Systems. In Proceedings of the Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, 24–28 April 2025; OpenReview.net: Alameda, CA, USA, 2025. [Google Scholar]
- ES, S.; James, J.; Anke, L.E.; Schockaert, S. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2024—System Demonstrations, St. Julians, Malta, 17–22 March 2024; Aletras, N., Clercq, O.D., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 150–158. [Google Scholar]
- Lin, J.; Zhang, C.; Liu, S.Y.; Li, H. RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems. arXiv 2025, arXiv:2510.13910. [Google Scholar] [CrossRef] [Scilit]
- Strich, J.; Isgorur, E.K.; Trescher, M.; Biemann, C.; Semmann, M. T2-RAGBench: Text-and-Table Benchmark for Evaluating Retrieval-Augmented Generation. arXiv 2025, arXiv:2506.12071. [Google Scholar] [CrossRef] [Scilit]
- Choi, C.; Kwon, J.; Ha, J.; Choi, H.; Kim, C.; Lee, Y.; Sohn, J.Y.; Lopez-Lira, A. Finder: Financial dataset for question answering and evaluating retrieval-augmented generation. In Proceedings of the 6th ACM International Conference on AI in Finance; ACM: New York, NY, USA, 2025; pp. 638–646. [Google Scholar]
- Anugraha, D.; Irawan, P.A.; Singh, A.; Lee, E.S.A.; Winata, G.I. M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG. arXiv 2025, arXiv:2512.05959. [Google Scholar]






| Symbol | Description |
|---|---|
| Heterogeneous Table-Text Graph | |
| Paragraph, Table, and Cell nodes | |
| Structural and Semantic edges | |
| Embedding vector for node | |
| Synthesized symbolic program |
| Configuration | FinQA EM | FinanceBench Acc | Retrieval R@5 |
|---|---|---|---|
| Full Model | 82.5% | 74.0% | 78.0% |
| No Hierarchy | 78.0% | 69.5% | 65.0% |
| No Graph (TTGNN) | 75.5% | 66.0% | 62.0% |
| No Symbolic | 70.0% | 62.5% | 78.0% |
| No Reranking | 79.2% | 71.0% | 72.0% |
| Step | Vanilla RAG (Baseline) | HierFinRAG (Ours) |
|---|---|---|
| Retrieval | Retrieves generic “Operating Expenses” paragraph. Misses footnote 3 pages away due to lack of structural link. | TTGNN traverses Structural edge to header, then Cross-Ref edge to “Note 12: Stock-Based Compensation”. |
| Reasoning | LLM attempts generation from incomplete context. | Router triggers Hybrid mode. Decomposes query into “Gross R&D” and “Stock Comp”. |
| Execution | Hallucinates a value or returns gross figure (incorrect). | Symbolically executes using extracted values. |
| Outcome | Failure (Incorrect/Hallucinated) | Success (Exact Derived Match) |
| Dataset | Avg Input Tokens | Avg Output Tokens | Est. Cost (1k Queries) |
|---|---|---|---|
| FinQA | 1250 | 150 | $6.20 |
| FinanceBench | 4500 | 300 | $21.50 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Dang, Q.-V.; Nguyen, N.-S.-A.; Vo, T.-B.-D. HierFinRAG—Hierarchical Multimodal RAG for Financial Document Understanding. Informatics 2026, 13, 30. https://doi.org/10.3390/informatics13020030
Dang Q-V, Nguyen N-S-A, Vo T-B-D. HierFinRAG—Hierarchical Multimodal RAG for Financial Document Understanding. Informatics. 2026; 13(2):30. https://doi.org/10.3390/informatics13020030
Chicago/Turabian StyleDang, Quang-Vinh, Ngoc-Son-An Nguyen, and Thi-Bich-Diem Vo. 2026. "HierFinRAG—Hierarchical Multimodal RAG for Financial Document Understanding" Informatics 13, no. 2: 30. https://doi.org/10.3390/informatics13020030
APA StyleDang, Q.-V., Nguyen, N.-S.-A., & Vo, T.-B.-D. (2026). HierFinRAG—Hierarchical Multimodal RAG for Financial Document Understanding. Informatics, 13(2), 30. https://doi.org/10.3390/informatics13020030

