Routed Prototype Adapters for Federated Financial Return Prediction with Frozen LLMs
Abstract
1. Introduction
- We formulate federated financial LLM adaptation as a compositional personalization problem. Instead of assuming a single global adapter or a fixed global–local split, we argue that financial clients should be able to share multiple adaptation directions and combine them according to local market signals.
- We propose a federated routed-adapter framework that integrates projected directional routing, simplex mirror-descent route updates, local residual learning, and least-norm residual-to-prototype decomposition. The individual mathematical tools are standard, but their integration provides a route-consistent adapter aggregation mechanism for frozen LLMs under heterogeneous financial clients.
- We provide extensive empirical evaluation across text-driven return prediction, structured stock-ranking benchmarks, and stock-movement classification tasks. Additional analyses examine routing stability, route entropy, partition robustness, communication and wall-clock cost, statistical significance, and the sensitivity to client participation and calibration batch size.
2. Related Work
2.1. Financial Forecasting with Structured and Textual Signals
2.2. Federated Learning and Personalization Under Heterogeneity
2.3. Federated Parameter-Efficient Tuning of Foundation Models
2.4. Relation to Personalized Federated PEFT and Routing-Based Adaptation
3. Method
3.1. Prototype Pool and Personalized Prediction
3.2. Projected Directional Routing
3.3. Local Residual Learning and Federated Prototype Update
3.4. Algorithm Description and Complexity
| Algorithm 1 Federated Projected Routing with Local Residual Adapters |
|
Optimization Interpretation
4. Experimental Results and Analysis
4.1. Experimental Setup
- Datasets. We evaluate our method on three groups of financial prediction benchmarks. For text-driven return prediction, we use FNSPID, which aligns financial news with stock prices and supports joint modeling of textual signals and market covariates [15]. For structured market prediction, we use the Qlib CSI300 and CSI800 benchmarks with Alpha158 and Alpha360 features, following the common stock-ranking protocol used by recent forecasting models such as MASTER [10,36]. For stock-movement classification, we further include the Open FinLLM forecasting suite [7,65,66]. All datasets are split chronologically to avoid look-ahead bias. In the federated setting, we construct clients by sector-level stock groups when sector metadata is available; otherwise, we use Dirichlet non-IID partitions over stock identities with concentration parameter .
- Evaluation metrics. For return regression, we report MSE and MAE. For cross-sectional stock ranking, we report IC, RankIC, ICIR, RankICIR, annualized return (AR), information ratio (IR), annualized volatility (AV), and maximum drawdown (MDD), following Qlib-style evaluation protocols [10,36]. For stock-movement classification, we report Accuracy and Matthews correlation coefficient (MCC), consistent with Open FinLLM forecasting tasks [65]. For federated learning, we additionally report macro-averaged client performance, worst-quartile client performance, trainable parameters, and communication cost per round.
- Compared methods. We compare with four groups of baselines. First, we include standard and personalized FL methods: Local-only, FedAvg [17], FedAvgM [67], FedProx [18], SCAFFOLD [19], FedPer [20], FedRep [21], and Ditto [22]. Second, we compare with federated PEFT and LLM fine-tuning methods, including FedIT [23], FFA-LoRA [46], FLoRA [24], FlexLoRA [25], FedBiOT [47], FwdLLM [48], FedDPA [49], FDLoRA [26], FedSA-LoRA [27], and FedEx-LoRA [28]. Third, we include financial forecasting baselines commonly used in stock-ranking benchmarks, including XGBoost [68], LightGBM [69], CatBoost [70], LSTM [71], GRU [72], TCN [73], Transformer [74], GAT [75], ALSTM [76], HIST [8], and MASTER [10]. Finally, we compare with financial LLM baselines, including FinBERT [11], FinGPT [12], PIXIU [13], PloutosGPT [77], and StockLLM [78].
- Implementation details. We use Llama-3.2-1B-Instruct as the frozen backbone and insert LoRA-style adapters into the attention projection layers. The adapter rank is , LoRA scaling is 16, and adapter dropout is . We use clients for FNSPID and Qlib, clients for smaller movement-prediction datasets, and sample of clients per communication round. The default number of prototypes is , the projection dimension is , the routing step size is , the server step size is , and each selected client performs local SGD steps per round. The calibration batch size and local training batch size are both 16. We optimize adapter residuals and prediction heads with AdamW using learning rates and , respectively. All experiments are repeated with three random seeds and run on four NVIDIA A100 80GB GPUs with bfloat16 training.
Fair Comparison Protocol
4.2. Main Results
- Overall comparison on stock-ranking benchmarks. Table 3 reports the main comparison on the Qlib CSI300 and CSI800 benchmarks. We include centralized stock-forecasting models as non-private reference baselines and compare our method with representative personalized FL and federated PEFT methods under the same client partition and communication protocol. Our method achieves the best performance across all reported ranking and portfolio metrics. Compared with the strongest centralized reference, MASTER, our method improves IC/RankIC from to on CSI300 and from to on CSI800. Among federated baselines, FedEx-LoRA is generally the strongest competitor, but it still underperforms our method, suggesting that exact adapter aggregation alone is insufficient for heterogeneous financial clients. The consistent gains of our method indicate that projected routing and local residual learning provide effective client-specific adaptation while preserving a shared prototype space.
- Results on text-driven return prediction and stock-movement classification. Table 4 further evaluates whether the proposed routed-adapter framework can exploit textual financial signals. On FNSPID, our method obtains the lowest MAE and MSE, outperforming both conventional sequence models and financial LLM baselines. On the Open FinLLM forecasting datasets, our method consistently achieves the highest ACC and MCC on BigData22, ACL18, and CIKM18. The improvement is particularly clear on MCC, which is more informative under class imbalance. These results suggest that the proposed method does not merely improve average prediction accuracy, but also produces more balanced directional decisions. The advantage over PloutosGPT and StockLLM indicates that federated prototype routing can provide additional robustness beyond direct financial instruction tuning, especially when client distributions are heterogeneous.
Statistical Significance
4.3. Ablation and Analysis
- Single-factor ablation. We conduct single-factor ablations to isolate the contribution of each key design component. All variants are evaluated under the same backbone, client partition, communication budget, and training protocol as the full model. Table 6 reports representative results on CSI300, FNSPID, and the averaged Open FinLLM movement-classification benchmarks. The small blue numbers beside each result denote the performance degradation relative to the full model; for error metrics, positive values indicate increased error.
- Parameter sensitivity analysis. We further study the sensitivity of our method to key hyperparameters, including the number of prototypes M, projection dimension q, routing step size , server update rate , and local residual steps E. For each group, we vary one hyperparameter while keeping the others at the default configuration. Figure 2 reports mean and standard deviation over three random seeds. Overall, the performance remains stable across a broad range of values, showing that our method is not sensitive to a narrow hyperparameter choice.
4.4. Routing Stability and Calibration Sensitivity
- Does routing capture client heterogeneity? To verify whether the learned routing distribution reflects client heterogeneity, we visualize the average final-round route of each client over the prototype pool. Specifically, for client i, we computewhere denotes the last 20 communication rounds. We also report each client’s normalized route entropy and its RankIC improvement over the uniform-routing variant. The clients are sorted by their dominant prototype and sector annotation.
4.4.1. Partition Robustness Beyond Sector Grouping
- Does the residual serve personalized adaptation? We further analyze whether the local residual mainly benefits clients with stronger distributional heterogeneity. For each client i, we compute a heterogeneity scorewhere denotes the local return-label distribution and denotes the mean structured-covariate vector of client i. We then compare the full model with the variant without local residual learning and measure the client-level gainWe also compute the normalized residual strength to quantify how much client-specific correction is learned around the routed adapter.
- Is our method more useful under stronger client heterogeneity? We further vary the degree of client heterogeneity using a Dirichlet partition over stock identities. Specifically, we sample client assignments with concentration parameter , where smaller indicates stronger non-IID heterogeneity. All methods are trained under the same backbone, communication rounds, and local-update budget. We report the macro-averaged client RankIC over three random seeds on the CSI300 benchmark.
- Is the method more stable across market regimes? We further evaluate whether the proposed routed prototype adaptation remains robust under changing market conditions. We partition the CSI300 test period into five mutually exclusive market regimes using only the benchmark index close price. Let , , and . The return cutoffs and are the 30th and 70th percentiles of , and the volatility cutoff is the 80th percentile of , all estimated once on the 2008–2014 training split and then frozen. Test days are labeled in the following priority order: HighVol if ; Bear if and ; Recovery if and ; Bull if ; and Neutral otherwise. This precedence makes every day belong to exactly one regime and prevents test-period information from entering the thresholds.
- Does the method achieve a better performance–communication trade-off? We finally analyze the trade-off between predictive performance and communication efficiency. For each federated method, we measure the per-round communication cost per selected client, including both download and upload payloads. The performance axis reports macro-averaged client RankIC on CSI300, the bubble size denotes the number of trainable parameters, and the vertical bars show the standard deviation over three seeds. A method is Pareto-efficient if no other method achieves higher RankIC with lower communication cost.
4.4.2. Sensitivity to Client Participation
4.4.3. Client-Scale Robustness
4.4.4. Contribution of Input Modalities
4.4.5. Projection Variants
5. Conclusions
Limitations
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| LLM | Large Language Model |
| FL | Federated Learning |
| PEFT | Parameter-Efficient Fine-Tuning |
| LoRA | Low-Rank Adaptation |
| SGD | Stochastic Gradient Descent |
| KL | Kullback–Leibler |
| JSD | Jensen–Shannon Divergence |
| IID | Independent and Identically Distributed |
| Non-IID | Non-Independent and Identically Distributed |
| FNSPID | Financial news–stock price dataset |
| Qlib | AI-oriented quantitative investment platform |
| CSI300 | China Securities Index 300 |
| CSI800 | China Securities Index 800 |
| IC | Information Coefficient |
| RankIC | Rank Information Coefficient |
| ICIR | Information Coefficient Information Ratio |
| RankICIR | Rank Information Coefficient Information Ratio |
| AR | Annualized Return |
| IR | Information Ratio |
| AV | Annualized Volatility |
| MDD | Maximum Drawdown |
| MAE | Mean Absolute Error |
| MSE | Mean Squared Error |
| MCC | Matthews Correlation Coefficient |
Appendix A. Additional Technical Details for the Method
| Symbol | Meaning |
|---|---|
| K/M | Number of clients/number of shared adapter prototypes |
| k/ | Communication round/selected client set in round k |
| Prototype adapter object m maintained by the server | |
| Client i’s simplex route over prototype adapters | |
| Local residual adapter learned by client i around its routed mixture | |
| P/q | Public random projection matrix/projected routing dimension |
| / | Normalized calibration gradient/projected calibration gradient |
| / | EMA prototype-update accumulator/projected prototype signature |
| Directional alignment score between client i and prototype m | |
| / | Client-to-prototype decomposed residual/aggregated prototype residual |
Appendix A.1. Adapter Instantiation
Appendix A.2. Server-Side Construction of Projected Prototype Signatures
Appendix A.3. Why Random Projection Preserves Routing Geometry
Appendix A.4. Derivation of the KL Mirror-Descent Routing Update
Appendix A.5. Least-Norm Decomposition and Exact Reconstruction
Appendix A.6. Round-Wise Training Procedure
- Server broadcast. The server sends to the participating clients.
- Projected routing at client i. Client i forms its current routed adaptercomputes the normalized calibration gradient , evaluates alignment scores , and updates the route to by the KL mirror-descent rule.
- Local residual learning. Client i formsruns E local SGD steps onand obtains the final residual . It uploads .
- Server reconstruction and aggregation. For each participating client, the server computes
Appendix B. Additional Experimental Details
Appendix B.1. Data Preprocessing
Appendix B.2. Prompt Construction
Appendix B.3. Federated Partitioning
Appendix B.4. Backbone and Adapter Configuration
Appendix B.5. Training Protocol
Appendix B.6. Hyperparameter Search
Appendix B.7. Fair Comparison Protocol
Appendix B.8. Efficiency Measurement
References
- Ou, H.H.; Chen, G.Y.; Lin, I.C. A self-sovereign identity blockchain framework for access control and transparency in financial institutions. Cryptography 2025, 9, 9. [Google Scholar] [CrossRef] [Scilit]
- Feng, R.; Jiang, S.; Liang, X.; Xia, M. Stgat: Spatial–temporal graph attention neural network for stock prediction. Appl. Sci. 2025, 15, 4315. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Bu, N.; Li, Z.; Zhang, Y.; Zhao, Z. AT-FinGPT: Financial risk prediction via an audio-text large language model. Financ. Res. Lett. 2025, 77, 106967. [Google Scholar] [CrossRef] [Scilit]
- Cheng, Z.; Lai, L.; Liu, Y.; Cheng, K.; Qi, X. Enhancing Financial Report Question-Answering: A Retrieval-Augmented Generation System with Reranking Analysis. arXiv 2026, arXiv:2603.16877. [Google Scholar] [CrossRef] [Scilit]
- Zheng, X.; Wu, L.; Yan, Z.; Tang, Y.; Zhao, H.; Zhong, C.; Chen, B.; Gong, J. Large Language Models Powered Context-Aware Motion Prediction in Autonomous Driving. In Proceedings of the 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2024; pp. 980–985. [Google Scholar]
- Cheng, K.; Qi, X.; Cheng, Z.; Lai, L.; Liu, X. Regime-Dependent Volatility Dynamics: Evidence from Time-Series Analysis. In Proceedings of the 2026 3rd International Conference on Applied Economics, Management Science and Social Development (AEMSS 2026); Atlantis Press: Dordrecht, The Netherlands, 2026; pp. 179–189. [Google Scholar] [CrossRef] [Scilit]
- Xu, Y.; Cohen, S.B. Stock movement prediction from tweets and historical prices. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Melbourne, Australia, 15–20 July 2018; pp. 1970–1979. [Google Scholar]
- Xu, W.; Liu, W.; Wang, L.; Xia, Y.; Bian, J.; Yin, J.; Liu, T.Y. Hist: A graph-based framework for stock trend forecasting via mining concept-oriented shared information. arXiv 2021, arXiv:2110.13716. [Google Scholar]
- Lin, H.; Zhou, D.; Liu, W.; Bian, J. Learning multiple stock trading patterns with temporal routing adaptor and optimal transport. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, Virtual, 14–18 August 2021; pp. 1017–1026. [Google Scholar]
- Li, T.; Liu, Z.; Shen, Y.; Wang, X.; Chen, H.; Huang, S. Master: Market-guided stock transformer for stock price forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–27 February 2024; Volume 38, pp. 162–170. [Google Scholar]
- Araci, D. Finbert: Financial sentiment analysis with pre-trained language models. arXiv 2019, arXiv:1908.10063. [Google Scholar]
- Yang, H.; Liu, X.Y.; Wang, C.D. Fingpt: Open-source financial large language models. arXiv 2023, arXiv:2306.06031. [Google Scholar]
- Xie, Q.; Han, W.; Zhang, X.; Lai, Y.; Peng, M.; Lopez-Lira, A.; Huang, J. PIXIU: A large language model, instruction data and evaluation benchmark for finance. In Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, LA, USA, 10–16 December 2023; pp. 33469–33484. [Google Scholar]
- Xie, Q.; Han, W.; Chen, Z.; Xiang, R.; Zhang, X.; He, Y.; Xiao, M.; Li, D.; Dai, Y.; Feng, D.; et al. Finben: A holistic financial benchmark for large language models. Adv. Neural Inf. Process. Syst. 2024, 37, 95716–95743. [Google Scholar] [CrossRef] [Scilit]
- Dong, Z.; Fan, X.; Peng, Z. Fnspid: A comprehensive financial news dataset in time series. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Barcelona, Spain, 25–29 August 2024; pp. 4918–4927. [Google Scholar]
- Liu, Y.; Cheng, Z.; Lai, L. Improving the Completeness and Comparability of Segment Disclosures: A Large Language Model Approach. arXiv 2026, arXiv:2605.23924. [Google Scholar] [CrossRef]
- McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; y Arcas, B.A. Communication-efficient learning of deep networks from decentralized data. In Proceedings of the Artificial Intelligence and Statistics, PMLR, Fort Lauderdale, FL, USA, 20–22 April 2017; pp. 1273–1282. [Google Scholar]
- Li, T.; Sahu, A.K.; Zaheer, M.; Sanjabi, M.; Talwalkar, A.; Smith, V. Federated optimization in heterogeneous networks. Proc. Mach. Learn. Syst. 2020, 2, 429–450. [Google Scholar]
- Karimireddy, S.P.; Kale, S.; Mohri, M.; Reddi, S.; Stich, S.; Suresh, A.T. Scaffold: Stochastic controlled averaging for federated learning. In Proceedings of the International Conference on Machine Learning, PMLR, Virtual, 13–18 July 2020; pp. 5132–5143. [Google Scholar]
- Arivazhagan, M.G.; Aggarwal, V.; Singh, A.K.; Choudhary, S. Federated learning with personalization layers. arXiv 2019, arXiv:1912.00818. [Google Scholar]
- Collins, L.; Hassani, H.; Mokhtari, A.; Shakkottai, S. Exploiting shared representations for personalized federated learning. In Proceedings of the International Conference on Machine Learning, PMLR, Virtual, 18–24 July 2021; pp. 2089–2099. [Google Scholar]
- Li, T.; Hu, S.; Beirami, A.; Smith, V. Ditto: Fair and robust federated learning through personalization. In Proceedings of the International Conference on Machine Learning, PMLR, Virtual, 18–24 July 2021; pp. 6357–6368. [Google Scholar]
- Zhang, J.; Vahidian, S.; Kuo, M.; Li, C.; Zhang, R.; Yu, T.; Wang, G.; Chen, Y. Towards building the federatedgpt: Federated instruction tuning. In Proceedings of the ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: Piscataway, NJ, USA, 2024; pp. 6915–6919. [Google Scholar]
- Wang, Z.; Shen, Z.; He, Y.; Sun, G.; Wang, H.; Lyu, L.; Li, A. Flora: Federated fine-tuning large language models with heterogeneous low-rank adaptations. Adv. Neural Inf. Process. Syst. 2024, 37, 22513–22533. [Google Scholar] [CrossRef] [Scilit]
- Bai, J.; Chen, D.; Qian, B.; Yao, L.; Li, Y. Federated fine-tuning of large language models under heterogeneous tasks and client resources. Adv. Neural Inf. Process. Syst. 2024, 37, 14457–14483. [Google Scholar] [CrossRef] [Scilit]
- Qi, J.; Luan, Z.; Huang, S.; Fung, C.; Yang, H.; Qian, D. Fdlora: Personalized federated learning of large language model via dual lora tuning. arXiv 2024, arXiv:2406.07925. [Google Scholar]
- Guo, P.; Zeng, S.; Wang, Y.; Fan, H.; Wang, F.; Qu, L. Selective Aggregation for Low-Rank Adaptation in Federated Learning. In Proceedings of the Thirteenth International Conference on Learning Representations, Singapore, 24–28 April 2025. [Google Scholar]
- Singhal, R.; Ponkshe, K.; Vepakomma, P. FedEx-LoRA: Exact aggregation for federated and efficient fine-tuning of large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, 27 July–1 August 2025; pp. 1316–1336. [Google Scholar]
- Xiao, C.; Liu, Y. A Multifrequency Data Fusion Deep Learning Model for Carbon Price Prediction. J. Forecast. 2025, 44, 436–458. [Google Scholar] [CrossRef] [Scilit]
- Zhang, F.; Fan, S.; Wang, H. Rethinking Activation Function: A Simple Path to Efficient and Accurate Time Series Forecasting. IEEE Trans. Knowl. Data Eng. 2026. early access. [Google Scholar] [CrossRef] [Scilit]
- Yan, J.; Zhou, N.; Cheng, Y.; Zhang, F.; Wang, H.; Wang, M.; Jin, B.; Li, M.; Lu, Q.; Zhang, W. Application of Machine-Vision-Driven Physics-Informed Neural Networks in Pantograph–Catenary System State Detection. Mech. Syst. Signal Process. 2026, 257, 114577. [Google Scholar] [CrossRef] [Scilit]
- Praveena, S.; Devi, S.P. Optimizing Retail Operations through Hybrid Machine Learning and Deep Learning Techniques in Demand Forecasting. Cybern. Syst. 2026, 57, 1–43. [Google Scholar] [CrossRef] [Scilit]
- Zhang, F.; Fan, S.; Wang, H. What If We Let Forecasting Forget? A Sparse Bottleneck for Cross-Variable Dependencies. In Proceedings of the Forty-Third International Conference on Machine Learning, Seoul, Republic of Korea, 6–11 July 2026. [Google Scholar]
- Sharma, N.; Jailia, M. Stock Trend Prediction Based on Hierarchical Trading Day Graph and Chaotic Spatio-Temporal Analysis. Cybern. Syst. 2026, 57, 146–188. [Google Scholar] [CrossRef] [Scilit]
- Yan, J.; Chen, B.; Zhang, F.; Cheng, Y.; Wang, H.; Wang, H.; Wang, M.; Li, T.; Zhang, W. Meta-Learning-Based Graph Convolutional Wavelet Network for Intelligent Dynamic Modeling of High-Speed Rail Subsystems. IEEE Trans. Veh. Technol. 2026. early access. [Google Scholar] [CrossRef] [Scilit]
- Yang, X.; Liu, W.; Zhou, D.; Bian, J.; Liu, T.Y. Qlib: An ai-oriented quantitative investment platform. arXiv 2020, arXiv:2009.11189. [Google Scholar]
- Zhang, Y.; Zhang, Y.; Liang, Y.; Mu, S.; Zhang, Z.; Chen, X. AD-VGF: An Improved Generative Adversarial Network for Credit Card Risk Identification and Management in Digital Economy. Cybern. Syst. 2025. advance online publication. [Google Scholar] [CrossRef] [Scilit]
- Cheng, Z.; Lai, L.; Liu, Y. Sustainable Hybrid Document-Routed Retrieval for Financial RAG: Resolving the Robustness-Precision Trade-off. arXiv 2026, arXiv:2603.26815. [Google Scholar] [CrossRef] [Scilit]
- Tian, B.; Liu, M.; Gao, H.a.; Li, P.; Zhao, H.; Zhou, G. Unsupervised Road Anomaly Detection with Language Anchors. In Proceedings of the 2023 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2023; pp. 7778–7785. [Google Scholar]
- Lian, S.; Cai, J.; Pan, D.; Chen, G.Y.; Xu, H.; Zhang, F.; Fan, G.; Pei, J.; Li, S. Anatomy-Aware Text-Visual Fusion with Dual-Perspective Prompts for Fine-Grained Lumbar Spine Segmentation. Int. J. Comput. Vis. 2026, 134, 241. [Google Scholar] [CrossRef] [Scilit]
- Jiang, Y.; Ferraro, F. SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models. arXiv 2026, arXiv:2601.03555. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Xiao, C.; Liu, Y. Confusion-resistant federated learning via diffusion-based data harmonization on non-IID data. Adv. Neural Inf. Process. Syst. 2024, 37, 137495–137520. [Google Scholar] [CrossRef] [Scilit]
- Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; Gelly, S. Parameter-efficient transfer learning for NLP. In Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA, 9–15 June 2019; pp. 2790–2799. [Google Scholar]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. Lora: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations, Virtual, 25–29 April 2022. [Google Scholar]
- Ye, R.; Wang, W.; Chai, J.; Li, D.; Li, Z.; Xu, Y.; Du, Y.; Wang, Y.; Chen, S. Openfedllm: Training large language models on decentralized private data via federated learning. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Barcelona, Spain, 25–29 August 2024; pp. 6137–6147. [Google Scholar]
- Sun, Y.; Li, Z.; Li, Y.; Ding, B. Improving lora in privacy-preserving federated learning. arXiv 2024, arXiv:2403.12313. [Google Scholar]
- Wu, F.; Li, Z.; Li, Y.; Ding, B.; Gao, J. Fedbiot: Llm local fine-tuning in federated learning without full model. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Barcelona, Spain, 25–29 August 2024; pp. 3345–3355. [Google Scholar]
- Xu, M.; Cai, D.; Wu, Y.; Li, X.; Wang, S. {FwdLLM}: Efficient federated finetuning of large language models with perturbed inferences. In Proceedings of the 2024 USENIX Annual Technical Conference (USENIX ATC 24), Santa Clara, CA, USA, 10–12 July 2024; pp. 579–596. [Google Scholar]
- Yang, Y.; Long, G.; Shen, T.; Jiang, J.; Blumenstein, M. Dual-personalizing adapter for federated foundation models. Adv. Neural Inf. Process. Syst. 2024, 37, 39409–39433. [Google Scholar] [CrossRef] [Scilit]
- Xiao, C.; Hou, L. Prototype-Aligned Federated Soft-Prompts for Continual Web Personalization. In Proceedings of the ACM Web Conference 2026, Dubai, United Arab Emirates, 13–17 April 2026; pp. 6743–6754. [Google Scholar]
- Jiang, Y.; Li, D.; Ferraro, F. DRP: Distilled Reasoning Pruning with Skill-Aware Step Decomposition for Efficient Large Reasoning Models. arXiv 2025, arXiv:2505.13975. [Google Scholar]
- Li, Y.; Ding, K.; Yang, C.; Chen, S.Y.; Tian, Y. Distilling Time Series Foundation Models for Efficient Forecasting. In Proceedings of the ICASSP 2026—2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: Piscataway, NJ, USA, 2026; pp. 4631–4635. [Google Scholar] [CrossRef] [Scilit]
- Xiao, C.; Xu, T.; Ma, S.; Jiang, Y.; Gao, H.; Wu, Y. Reversible Primitive–Composition Alignment for Continual Vision–Language Learning. In Proceedings of the Fourteenth International Conference on Learning Representations, Rio de Janeiro, Brazil, 23–27 April 2026. [Google Scholar]
- Zhou, H.; Tang, J.; Zhang, J.; Li, Y.; Xiao, C.; Hou, L.; Ke, Z.; Yao, J. Comem: Compositional Concept-Graph Memory for Vision–Language Adaptation. In Proceedings of the Fourteenth International Conference on Learning Representations, Rio de Janeiro, Brazil, 23–27 April 2026. [Google Scholar]
- Lin, N.; Xu, Y.; Yang, H.; Zhang, G.; Zhang, M.; Wang, S.; Hua, H.; Li, X. Dissociating the Neural Correlates of the Sociality and Plausibility Effects in Simple Conceptual Combination. Brain Struct. Funct. 2020, 225, 995–1008. [Google Scholar] [CrossRef] [Scilit]
- Zhang, G.; Xu, Y.; Zhang, M.; Wang, S.; Lin, N. The Brain Network in Support of Social Semantic Accumulation. Soc. Cogn. Affect. Neurosci. 2021, 16, 393–405. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Liu, T.; Zhao, H.; Zhou, G.; Zhang, Y.Q. Cerberus Transformer: Joint Semantic, Affordance and Attribute Parsing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2022; pp. 19649–19658. [Google Scholar]
- Chen, Y.; Cao, Z.; Ren, H.; Yang, C.; Li, W.; Wang, S.; Wang, Y.; Zhang, L.; Shao, Y.; Zhao, Z.; et al. RoboRouter: Training-Free Policy Routing for Robotic Manipulation. arXiv 2026, arXiv:2603.07892. [Google Scholar]
- Zhang, H.; Yang, S.; Liang, X.; Shang, C.; Jiang, Y.; Tao, C.; Xiong, J.; So, H.K.H.; Xie, R.; Chang, A.X.; et al. Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation. arXiv 2025, arXiv:2510.10925. [Google Scholar]
- Yang, H.; Liu, H.; Yuan, X.; Wu, K.; Ni, W.; Zhang, J.A.; Liu, R.P. Synergizing Intelligence and Privacy: A Review of Integrating Internet of Things, Large Language Models, and Federated Learning in Advanced Networked Systems. Appl. Sci. 2025, 15, 6587. [Google Scholar] [CrossRef] [Scilit]
- González-Quesada, J.C.; Trillo, J.R.; Porcel, C.; Pérez, I.J.; Cabrerizo, F.J. Modelling Large-Scale Group Decision-Making Through Grouping with Large Language Models. Future Internet 2025, 17, 381. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Zhao, C.; Xiao, C.; Duan, R.; Mo, W.; Gao, H.; Wang, W. Pi-CCA: Prompt-Invariant CCA Certificates for Replay-Free Continual Multimodal Learning. In Proceedings of the Fourteenth International Conference on Learning Representations, Rio de Janeiro, Brazil, 23–27 April 2026. [Google Scholar]
- Liao, B.; Zhao, Z.; Chen, L.; Li, H.; Cremers, D.; Liu, P. GlobalPointer: Large-Scale Plane Adjustment with Bi-Convex Relaxation. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2024; pp. 360–376. [Google Scholar]
- Zhao, Z.; Yang, H.; Liao, B.; Zeng, Y.; Yan, S.; Gu, Y.; Liu, P.; Zhou, Y.; Li, H.; Civera, J. Advances in Global Solvers for 3D Vision. arXiv 2026, arXiv:2602.14662. [Google Scholar]
- Lin, S.C.; Tian, F.; Wang, K.; Zhao, X.; Huang, J.; Xie, Q.; Borella, L.; White, M.; Wang, C.D.; Xiao, K.; et al. Open finllm leaderboard: Towards financial ai readiness. arXiv 2025, arXiv:2501.10963. [Google Scholar]
- Wu, H.; Zhang, W.; Shen, W.; Wang, J. Hybrid deep sequential modeling for social text-driven stock prediction. In Proceedings of the 27th ACM International Conference on Information and Knowledge Management, Torino, Italy, 22–26 October 2018; pp. 1627–1630. [Google Scholar]
- Hsu, T.M.H.; Qi, H.; Brown, M. Measuring the effects of non-identical data distribution for federated visual classification. arXiv 2019, arXiv:1909.06335. [Google Scholar]
- Chen, T.; Guestrin, C. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar]
- Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.Y. Lightgbm: A highly efficient gradient boosting decision tree. Adv. Neural Inf. Process. Syst. 2017, 30, 3146–3154. [Google Scholar]
- Prokhorenkova, L.; Gusev, G.; Vorobev, A.; Dorogush, A.V.; Gulin, A. CatBoost: Unbiased boosting with categorical features. Adv. Neural Inf. Process. Syst. 2018, 31, 6638–6648. [Google Scholar]
- Graves, A. Long short-term memory. In Supervised Sequence Labelling with Recurrent Neural Networks; Springer: Berlin/Heidelberg, Germany, 2012; pp. 37–45. [Google Scholar]
- Cho, K.; Van Merriënboer, B.; Gulçehre, Ç.; Bahdanau, D.; Bougares, F.; Schwenk, H.; Bengio, Y. Learning phrase representations using RNN encoder–decoder for statistical machine translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar, 25–29 October 2014; pp. 1724–1734. [Google Scholar]
- Bai, S.; Kolter, J.Z.; Koltun, V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv 2018, arXiv:1803.01271. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Veličković, P.; Cucurull, G.; Casanova, A.; Romero, A.; Liò, P.; Bengio, Y. Graph Attention Networks. In Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Feng, F.; Chen, H.; He, X.; Ding, J.; Sun, M.; Chua, T.S. Enhancing Stock Movement Prediction with Adversarial Training. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, International Joint Conferences on Artificial Intelligence Organization, Macao, China, 10–16 August 2019; pp. 5843–5849. [Google Scholar]
- Tong, H.; Li, J.; Wu, N.; Gong, M.; Zhang, D.; Zhang, Q. Ploutos: Towards Explainable Stock Movement Prediction with Financial Large Language Model. In Proceedings of the Companion Proceedings of the ACM on Web Conference 2025, Sydney, NSW, Australia, 28 April–2 May 2025; pp. 490–499. [Google Scholar]
- Xiao, M.; Jiang, Z.; Qian, L.; Chen, Z.; He, Y.; Xu, Y.; Jiang, Y.; Li, D.; Weng, R.L.; Peng, M.; et al. Enhancing financial time-series forecasting with retrieval-augmented large language models. arXiv 2025, arXiv:2502.05878. [Google Scholar]
- Xu, J.; Hong, C.; Huang, J.; Chen, L.Y.; Decouchant, J. AGIC: Approximate Gradient Inversion Attack on Federated Learning. arXiv 2022, arXiv:2204.13784. [Google Scholar]
- Bonawitz, K.; Ivanov, V.; Kreuter, B.; Marcedone, A.; McMahan, H.B.; Patel, S.; Ramage, D.; Segal, A.; Seth, K. Practical Secure Aggregation for Privacy-Preserving Machine Learning. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security; ACM: New York, NY, USA, 2017; pp. 1175–1191. [Google Scholar] [CrossRef] [Scilit]







| Method | Shared Adaptation | Client Personalization | Routing Signal | Server Update |
|---|---|---|---|---|
| FedDPA | One global adapter | Local adapter and test-time weighting | Instance-wise weighting | Global adapter |
| FDLoRA | One global LoRA branch | Personalized LoRA branch and fusion | Adaptive fusion | Global LoRA branch |
| FedSA-LoRA | Shared LoRA A factors | Local LoRA B factors | Selective factor sharing | Selected LoRA factors |
| FedEx-LoRA | Exact global LoRA update | Optional local state | None | Exact LoRA correction |
| MoE/adapter routing | Expert pool | Token- or instance-level expert choice | Learned gate logits | Expert parameters |
| Ours | Multiple adapter prototypes | Client route and local residual | Projected calibration-gradient alignment | Route-consistent prototype residuals |
| Method | Setting | Backbone | Input | Adapter/Model Head | Trainable Params | Rounds | Local Steps | Client Sampling | LR |
|---|---|---|---|---|---|---|---|---|---|
| Local-only | Federated, no aggregation | Llama-3.2-1B | Same as ours | LoRA rank + private head | 2.10 M | 100/50 | 3 | – | / |
| FedAvg | Federated | Llama-3.2-1B | Same as ours | LoRA rank + private head | 2.10 M | 100/50 | 3 | 25% | / |
| FedProx | Federated | Llama-3.2-1B | Same as ours | LoRA rank + private head | 2.10 M | 100/50 | 3 | 25% | / |
| SCAFFOLD | Federated | Llama-3.2-1B | Same as ours | LoRA rank + private head | 2.10 M | 100/50 | 3 | 25% | / |
| Ditto | Federated personalized | Llama-3.2-1B | Same as ours | Global LoRA + local model | 4.20 M | 100/50 | 3 | 25% | / |
| FLoRA | Federated PEFT | Llama-3.2-1B | Same as ours | LoRA, rank tuned within 10% params | 2.05 M | 100/50 | 3 | 25% | / |
| FedDPA | Federated PEFT | Llama-3.2-1B | Same as ours | Global/local adapters | 2.24 M | 100/50 | 3 | 25% | / |
| FedSA-LoRA | Federated PEFT | Llama-3.2-1B | Same as ours | Shared A, local B LoRA | 2.08 M | 100/50 | 3 | 25% | / |
| FedEx-LoRA | Federated PEFT | Llama-3.2-1B | Same as ours | Exact LoRA aggregation | 2.12 M | 100/50 | 3 | 25% | / |
| Ours | Federated PEFT | Llama-3.2-1B | Same as ours | prototypes + route + residual | 2.31 M | 100/50 | 3 | 25% | / |
| FinGPT/PIXIU/PloutosGPT/StockLLM | Non-federated reference | Original or matched backbone where applicable | Text or serialized features | Original method | 1 B–7 B | – | – | – | – |
| MASTER/XGBoost/LightGBM | Centralized structured reference | Structured model | Market covariates | Original model | 0.4 M–12 M | – | – | – | – |
| Category | Method | CSI300 | CSI800 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| IC ↑ | RankIC ↑ | AR ↑ | IR ↑ | IC ↑ | RankIC ↑ | AR ↑ | IR ↑ | ||
| Centralized forecasting | XGBoost | 0.051 | 0.050 | 0.230 | 1.90 | 0.040 | 0.047 | 0.080 | 0.60 |
| LSTM | 0.049 | 0.051 | 0.200 | 2.00 | 0.028 | 0.039 | 0.090 | 0.90 | |
| GRU | 0.052 | 0.052 | 0.190 | 1.50 | 0.039 | 0.044 | 0.070 | 0.60 | |
| TCN | 0.050 | 0.049 | 0.180 | 1.40 | 0.038 | 0.045 | 0.050 | 0.40 | |
| Transformer | 0.047 | 0.051 | 0.220 | 2.00 | 0.040 | 0.048 | 0.130 | 1.10 | |
| GAT | 0.054 | 0.041 | 0.190 | 1.30 | 0.043 | 0.042 | 0.100 | 0.70 | |
| MASTER | 0.064 | 0.076 | 0.270 | 2.40 | 0.052 | 0.066 | 0.280 | 2.30 | |
| Personalized FL | Local-only | 0.055 | 0.064 | 0.232 | 2.01 | 0.045 | 0.055 | 0.183 | 1.53 |
| FedAvg | 0.058 | 0.068 | 0.244 | 2.12 | 0.048 | 0.059 | 0.212 | 1.75 | |
| FedProx | 0.059 | 0.069 | 0.250 | 2.18 | 0.049 | 0.060 | 0.221 | 1.84 | |
| SCAFFOLD | 0.060 | 0.071 | 0.259 | 2.27 | 0.050 | 0.062 | 0.238 | 1.96 | |
| Ditto | 0.062 | 0.073 | 0.270 | 2.38 | 0.052 | 0.065 | 0.264 | 2.16 | |
| Federated PEFT | FLoRA | 0.063 | 0.075 | 0.279 | 2.48 | 0.053 | 0.067 | 0.279 | 2.31 |
| FlexLoRA | 0.064 | 0.076 | 0.284 | 2.51 | 0.053 | 0.067 | 0.283 | 2.34 | |
| FedDPA | 0.065 | 0.078 | 0.291 | 2.56 | 0.054 | 0.068 | 0.287 | 2.37 | |
| FedSA-LoRA | 0.066 | 0.080 | 0.298 | 2.62 | 0.055 | 0.070 | 0.293 | 2.42 | |
| FedEx-LoRA | 0.068 | 0.082 | 0.307 | 2.70 | 0.056 | 0.071 | 0.303 | 2.49 | |
| Ours | 0.072 | 0.087 | 0.331 | 2.91 | 0.059 | 0.075 | 0.333 | 2.73 | |
| Method | FNSPID | BigData22 | ACL18 | CIKM18 | ||||
|---|---|---|---|---|---|---|---|---|
| MAE ↓ | MSE ↓ | Acc. ↑ | MCC ↑ | Acc. ↑ | MCC ↑ | Acc. ↑ | MCC ↑ | |
| LSTM | 0.02493 | 17.00 | 51.00 | 0.010 | 53.00 | 0.060 | 53.00 | 0.020 |
| Transformer | 0.00544 | 0.50 | 52.74 | 0.041 | 55.16 | 0.083 | 54.87 | 0.041 |
| DTML | 0.01273 | 3.20 | 52.00 | 0.070 | 57.44 | 0.191 | 58.62 | 0.045 |
| StockNet | 0.01346 | 3.80 | 53.00 | 0.000 | 58.23 | 0.081 | 56.37 | 0.023 |
| SLOT | 0.01084 | 1.90 | 55.00 | 0.100 | 59.00 | 0.210 | 56.00 | 0.090 |
| FinGPT | 0.00687 | 0.81 | 54.72 | 0.064 | 56.12 | 0.109 | 55.86 | 0.053 |
| PIXIU/FinMA | 0.00672 | 0.78 | 51.00 | 0.020 | 56.28 | 0.104 | 53.24 | −0.031 |
| PloutosGPT | 0.00635 | 0.70 | 56.03 | 0.116 | 61.21 | 0.205 | 59.89 | 0.064 |
| StockLLM | 0.00619 | 0.67 | 56.58 | 0.119 | 60.76 | 0.218 | 60.41 | 0.098 |
| FedAvg | 0.00658 | 0.74 | 56.27 | 0.112 | 60.02 | 0.197 | 59.33 | 0.083 |
| FLoRA | 0.00621 | 0.66 | 56.83 | 0.123 | 60.94 | 0.212 | 60.12 | 0.094 |
| FedDPA | 0.00596 | 0.59 | 57.08 | 0.132 | 61.45 | 0.222 | 60.56 | 0.103 |
| FedSA-LoRA | 0.00571 | 0.52 | 57.42 | 0.137 | 61.96 | 0.231 | 60.88 | 0.112 |
| FedEx-LoRA | 0.00537 | 0.46 | 57.81 | 0.145 | 62.34 | 0.239 | 61.27 | 0.121 |
| Ours | 0.00482 | 0.38 | 58.63 | 0.162 | 63.19 | 0.258 | 62.18 | 0.138 |
| Dataset | Comparison | Mean Difference | 95% CI | p-Value | Significant at 0.05? |
|---|---|---|---|---|---|
| CSI300 | Ours–FedEx-LoRA | 0.005 | [0.0021, 0.0078] | 0.004 | Yes |
| CSI300 | Ours–FedSA-LoRA | 0.007 | [0.0039, 0.0102] | < | Yes |
| CSI800 | Ours–FedEx-LoRA | 0.004 | [0.0013, 0.0065] | 0.012 | Yes |
| FNSPID MAE | FedEx-LoRA–Ours | 0.00055 | [0.00021, 0.00088] | 0.006 | Yes |
| Variant | CSI300 | FNSPID | Open FinLLM Avg. | |||
|---|---|---|---|---|---|---|
| IC ↑ | RankIC ↑ | MAE ↓ | MSE ↓ | Acc. ↑ | MCC ↑ | |
| Ours full model | 0.0718 | 0.0871 | 0.00482 | 0.38 | 61.33 | 0.186 |
| w/o prototype pool () | 0.0659 (−0.0059) | 0.0792 (−0.0079) | 0.00548 (+0.00066) | 0.56 (+0.18) | 59.74 (−1.59) | 0.151 (−0.035) |
| w/o projected routing, fixed uniform route | 0.0676 (−0.0042) | 0.0810 (−0.0061) | 0.00524 (+0.00042) | 0.49 (+0.11) | 60.18 (−1.15) | 0.159 (−0.027) |
| w/o KL mirror descent | 0.0689 (−0.0029) | 0.0832 (−0.0039) | 0.00508 (+0.00026) | 0.45 (+0.07) | 60.56 (−0.77) | 0.168 (−0.018) |
| full-dimensional routing without projection | 0.0711 (−0.0007) | 0.0862 (−0.0009) | 0.00491 (+0.00009) | 0.40 (+0.02) | 61.04 (−0.29) | 0.181 (−0.005) |
| w/o local residual learning | 0.0647 (−0.0071) | 0.0778 (−0.0093) | 0.00556 (+0.00074) | 0.60 (+0.22) | 59.42 (−1.91) | 0.146 (−0.040) |
| w/o least-norm decomposition | 0.0668 (−0.0050) | 0.0803 (−0.0068) | 0.00536 (+0.00054) | 0.54 (+0.16) | 59.86 (−1.47) | 0.154 (−0.032) |
| w/o prototype-signature refresh | 0.0679 (−0.0039) | 0.0817 (−0.0054) | 0.00529 (+0.00047) | 0.51 (+0.13) | 60.05 (−1.28) | 0.157 (−0.029) |
| Calibration Batch Size | CSI300 RankIC | Route Entropy ↑ | Collapse Rate ↓ | Time/Round (s) |
|---|---|---|---|---|
| 4 | 0.0826 | 0.49 | 0.24 | 43.2 |
| 8 | 0.0849 | 0.57 | 0.16 | 45.1 |
| 16 | 0.0870 | 0.64 | 0.08 | 47.9 |
| 32 | 0.0871 | 0.67 | 0.05 | 52.8 |
| 64 | 0.0866 | 0.69 | 0.04 | 62.3 |
| Round | Route Entropy ↑ | Collapse Rate ↓ | Load Imbalance ↓ | Mean Route Change ↓ |
|---|---|---|---|---|
| 10 | 0.78 | 0.02 | 0.06 | 0.081 |
| 25 | 0.71 | 0.04 | 0.09 | 0.052 |
| 50 | 0.66 | 0.06 | 0.11 | 0.031 |
| 75 | 0.64 | 0.08 | 0.13 | 0.022 |
| 100 | 0.63 | 0.08 | 0.13 | 0.018 |
| Partition | RankIC | Gain over Uniform Route | Route Entropy | Prototype Load Imbalance | Corr. Heterogeneity–Gain |
|---|---|---|---|---|---|
| Random IID | 0.082 | 0.001 | 0.83 | 0.04 | 0.18 |
| Dirichlet | 0.084 | 0.003 | 0.75 | 0.08 | 0.36 |
| Dirichlet | 0.086 | 0.005 | 0.66 | 0.14 | 0.51 |
| Dirichlet | 0.088 | 0.007 | 0.57 | 0.21 | 0.63 |
| Sector-based | 0.087 | 0.006 | 0.62 | 0.18 | 0.58 |
| Method | Download MB | Upload MB | Total MB/Round | Time/Round (s) | CSI300 RankIC |
|---|---|---|---|---|---|
| FedAvg | 16.8 | 16.8 | 33.6 | 42.5 | 0.068 |
| FedDPA | 39.2 | 39.2 | 78.4 | 55.8 | 0.078 |
| FedSA-LoRA | 31.1 | 23.1 | 54.2 | 48.6 | 0.080 |
| FedEx-LoRA | 30.8 | 30.7 | 61.5 | 50.4 | 0.082 |
| Ours | 34.6 | 12.2 | 46.8 | 47.9 | 0.087 |
| Participation Ratio | RankIC | Rounds to 95% Best Validation RankIC | Final Route Entropy | Time/Round (s) |
|---|---|---|---|---|
| 10% | 0.083 | 74 | 0.69 | 42.1 |
| 25% | 0.087 | 58 | 0.64 | 47.9 |
| 50% | 0.088 | 52 | 0.61 | 59.4 |
| 100% | 0.088 | 45 | 0.58 | 88.7 |
| Number of Clients | Ours RankIC | FedEx-LoRA RankIC | Route Entropy | Prototype Load Imbalance |
|---|---|---|---|---|
| 20 | 0.087 | 0.082 | 0.64 | 0.14 |
| 50 | 0.085 | 0.081 | 0.67 | 0.11 |
| 100 | 0.083 | 0.079 | 0.70 | 0.09 |
| Variant | Text | Structured Covariates | FNSPID MAE | CSI300 RankIC |
|---|---|---|---|---|
| Text-only Ours | Yes | No | 0.00518 | – |
| Structured-as-text Ours | No | Serialized | 0.00541 | 0.081 |
| Text + structured-as-text Ours | Yes | Serialized | 0.00482 | 0.087 |
| Text + numeric-side-channel Ours | Yes | Numeric MLP fusion | 0.00476 | 0.088 |
| FedEx-LoRA, same input as ours | Yes | Serialized | 0.00537 | 0.082 |
| MASTER structured reference | No | Numeric | – | 0.076 |
| Routing Projection | RankIC | Routing Time/Round (s) | Additional Trainable Router Params |
|---|---|---|---|
| Rademacher random projection, | 0.087 | 1.9 | 0 |
| Gaussian random projection, | 0.0868 | 2.0 | 0 |
| Full-dimensional routing | 0.086 | 8.7 | 0 |
| Learned projection | 0.0873 | 3.4 | 0.26 M |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, B.; Ma, S.; Liu, Y. Routed Prototype Adapters for Federated Financial Return Prediction with Frozen LLMs. Electronics 2026, 15, 3900. https://doi.org/10.3390/electronics15173900
Li B, Ma S, Liu Y. Routed Prototype Adapters for Federated Financial Return Prediction with Frozen LLMs. Electronics. 2026; 15(17):3900. https://doi.org/10.3390/electronics15173900
Chicago/Turabian StyleLi, Bowen, Siyuan Ma, and Yang Liu. 2026. "Routed Prototype Adapters for Federated Financial Return Prediction with Frozen LLMs" Electronics 15, no. 17: 3900. https://doi.org/10.3390/electronics15173900
APA StyleLi, B., Ma, S., & Liu, Y. (2026). Routed Prototype Adapters for Federated Financial Return Prediction with Frozen LLMs. Electronics, 15(17), 3900. https://doi.org/10.3390/electronics15173900

