Merging Methods for Multilingual Knowledge Editing for Large Language Models: An Empirical Odyssey
Abstract
1. Introduction
- RQ1: How do vector merging methods perform in MKE?
- RQ2: Can TSVM effectively mitigate multilingual interference?
- RQ3: How do factors such as weight scale and rank ratio affect performance? We raise this question because both factors are fixed to defaults in prior locate-then-edit work, and their effect on MKE has not been measured.
- RQ4: Do merging methods benefit low-resource languages such as Thai and Vietnamese, and which merge operator best preserves their edits when many languages are edited jointly? We raise this question because low-resource languages score lowest in the aggregate results, and a single deployed model must serve them alongside high-resource languages.
- We evaluate six merging methods with two backbones and two base KE methods, and observe that vector summation with shared covariance achieves the strongest overall performance. We also show that TSVM can reduce interference under limited conditions, but in general, none of the tested merging methods effectively closes the gap between MKE and monolingual KE.
- We are the first to analyze the weight scaling factor in MKE, and find that the optimal scaling factor exceeds the default value of in most settings.
- We also investigate the effect of the rank compression ratio on TSVM performance and find that relatively low rank often leads to better results.
- We analyze performance by language-resource level and find that the choice of merging method has the largest effect on the relatively low-resource languages, where shared covariance summation and the orthogonalizing TSVM merge diverge most (RQ4).
2. Related Work
2.1. Knowledge Editing for Large Language Models
2.2. Task Vectors and Model Merging
3. Preliminaries
3.1. Large Language Models
3.2. Linear Associative Memory
3.3. Locate-Then-Edit Methods
4. Merging Methods for Multilingual Knowledge Editing
4.1. Problem Definition
4.2. Merging Functions
- Sum:
- Mean:
- TSVM [14]: First, we decompose each into , , using Singular Value Decomposition (SVD), then select the top-k highest singular values and their corresponding column and row vectors from and :The value k is determined as follows:where the factor r is a hyperparameter. It reflects the degree of rank compression: if r is equal to 1, that means no compression; if r is close to zero, the is highly compressed into a low-rank matrix. Subsequently, we concatenate resulted components across all languages as follows:To make orthogonal, we decompose to , , , by applying SVD, and set (similarly for ). Finally, we have the following:
- Sum-Cov: The same as Sum, but with s calculated by shared covariance. This is similar to the method used in the previous work [13] that adapted MEMIT to MKE. However, they apply the editing algorithm with one request (in multiple languages) at a time; by contrast, we apply the editing algorithm with massive requests (in multiple languages) simultaneously in one batch.
- Mean-Cov and TSVM-Cov: The same as Mean and TSVM, but with s calculated by shared covariance.
5. Experiments
5.1. Experimental Settings
5.1.1. Metrics
5.1.2. Dataset
5.1.3. Backbones and Base Methods
5.1.4. Monolingual Editing
5.1.5. Other Details
5.2. Experimental Results
5.2.1. Comparison of Averaged Accuracies of Merging Methods on MzsRE (RQ1, RQ2)
- Observation 1: Sum-Cov outperforms the other methods in most cases. Sum-Cov attains the best averaged accuracy in three of the four settings (, , and for Llama + MEMIT, Llama + AlphaEdit, and Qwen + MEMIT), exceeding the next-best merging method by , , and points, respectively. By contrast, plain summation without shared covariance (Sum) records almost zero accuracy. Mean, a magnitude-rescaled version of Sum, recovers to 46–49% but remains 11–13 points below Sum-Cov, and Mean-Cov is worse still (39–42%). Sharing covariance therefore helps when updates are summed but not when they are averaged.
- Observation 2: TSVM substantially improves overall performance, while TSVM-Cov fails to achieve higher scores than Sum-Cov in most cases. TSVM reaches 54–56% in the first three settings—far above Sum and Mean but still below Sum-Cov. However, as shown in Table 3, both TSVM and TSVM-Cov achieve higher accuracies than Sum-Cov with the Qwen backbone and AlphaEdit algorithm: TSVM attains versus for Sum-Cov (), and TSVM-Cov attains (). In this setting, Sum-Cov is also unstable—its Turkish score () is a clear outlier relative to the 50–60% range of the other languages.
- Observation 3: There is a substantial performance gap between monolingual and multilingual editing. The best multilingual method trails the monolingual upper bound (Mono) by , , , and points across the four settings, and no merging method closes this gap. Similarly, previous work [13] reports the performance gap between MKE and monolingual editing with experimental settings that differ from ours. They conduct editing and evaluation with one request at a time while we conduct experiments with all requests at once. We further investigate the possibility of mitigating multilingual interference with merging methods, but unfortunately, none of them effectively achieve the goal.
5.2.2. Effect of Weight Scale (RQ3)
5.2.3. Effect of Rank Ratio r (RQ3)
- Observation 1: In most cases, the curve of TSVM is unimodal while that of TSVM-Cov is multimodal. Specifically, the accuracy of TSVM-Cov drops sharply for rank ratios in the narrow range and then recovers, whereas TSVM stays unimodal throughout the evaluated range. We attribute this non-monotonic behavior to an interaction between the shared covariance and the low-rank truncation: the shared covariance solve already couples the per-language updates, and truncating it at a particular rank can discard directions that this coupling depends on, so accuracy degrades in a narrow band of r before higher ranks restore the missing directions. We analyze this shared covariance/low-rank interaction further in Section 6.4.
- Observation 2: Properly low rank ratio tends to achieve optimal performance. It is not surprising because of the low-rank nature of knowledge editing vectors [9]. In much prior work on model merging [14], they report rank reduction can boost merging performance. In the MKE, this result is the first analysis of the relation between the rank of and the performance.
5.2.4. Qualitative Error Analysis (RQ2)
6. Discussion
6.1. Effect of Vector Merging Framework (RQ1)
6.2. Effect of TSVM (RQ2)
6.3. Effect of Weight Scaling Factor (RQ3)
6.4. Effect of Rank Compression Ratio (RQ3)
6.5. Do Merging Methods Help Relatively Low-Resource Languages? (RQ4)
6.6. Limitations of This Work
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| LLM | Large Language Model |
| KE | Knowledge Editing |
| MKE | Multilingual Knowledge Editing |
| TSVM | Task Singular Vectors for Merging |
| RQ | Research Question |
| FFN | Feed-Forward Network |
| LAM | Linear Associative Memory |
| SVD | Singular Value Decomposition |
| MEMIT | Mass Editing Memory In Transformer |
| LRL | Low-Resource Language |
| HRL | High-Resource Language |
Appendix A. Evaluation Metric Definitions
- Efficacy is the accuracy on the editing requests themselves:
- Generalization is the accuracy on paraphrases of the edit requests:
- Specificity is the accuracy on questions that are unrelated to the editing requests:
- Portability is the accuracy on one-hop reasoning questions derived from the editing requests:
Appendix B. Computation of the Cross-Language Subspace Overlap in Figure 3
Appendix C. Conditioning of the Shared-Covariance Solve
- Condition number , the ratio of the largest to smallest singular value of .
- Null-space retention (AlphaEdit only), the fraction of edit-key energy that survives the null-space projection, .
- Edit-reproduction error. For each language ℓ, we form the matrix , i.e., the keys of language ℓ passed through the solve evaluated at their own keys; a faithful edit gives . We measure and average over languages. With the shared , this gives (the Sum-Cov solve); with the per-language , it gives (the per-language solve that TSVM merges). The last column of Table 7 is their ratio , the fidelity penalty incurred by sharing the covariance across languages.
Appendix D. Additional Qualitative Error Cases
| Language | New Object | Sum-Cov | TSVM | TSVM-Cov |
|---|---|---|---|---|
| English | 1952 | ✓ | ✓ | ✓ |
| Chinese | — | ✓ | ✓ | ✓ |
| Czech | 1952 | ✓ | ✓ | ✓ |
| Vietnamese | 1952 | × | ✓ | ✓ |
| Turkish | 1952 | ✓ | ✓ | × |
| French | 1952 | ✓ | ✓ | ✓ |
| Spanish | 1952 | × | × | ✓ |
| German | 1952 | × | × | ✓ |
| Russian | — | × | × | × |
| Dutch | 1952 | × | × | ✓ |
| Portuguese | 1952 | ✓ | ✓ | × |
| Thai | 1952 | × | × | ✓ |
| Language | New Object | Sum-Cov | TSVM | TSVM-Cov |
|---|---|---|---|---|
| English | Motown | ✓ | ✓ | ✓ |
| Chinese | Motown | ✓ | ✓ | ✓ |
| Czech | Motown | ✓ | ✓ | ✓ |
| Vietnamese | Motown | × | × | ✓ |
| Turkish | Motown | × | ✓ | ✓ |
| French | Mototown | × | × | × |
| Spanish | motown | ✓ | × | × |
| German | Motown | × | ✓ | ✓ |
| Russian | — | × | × | × |
| Dutch | Motown | × | ✓ | ✓ |
| Portuguese | Motown | ✓ | ✓ | ✓ |
| Thai | — | × | × | × |
| Language | New Object | Sum-Cov | TSVM | TSVM-Cov |
|---|---|---|---|---|
| English | 2005 | ✓ | ✓ | ✓ |
| Chinese | — | ✓ | × | × |
| Czech | 2005 | ✓ | ✓ | ✓ |
| Vietnamese | 2005 | × | ✓ | ✓ |
| Turkish | 2005 | × | ✓ | ✓ |
| French | 2005 | ✓ | × | ✓ |
| Spanish | 2005 | ✓ | ✓ | ✓ |
| German | 2005 | ✓ | ✓ | ✓ |
| Russian | — | × | × | × |
| Dutch | 2005 | ✓ | ✓ | ✓ |
| Portuguese | 2005 | × | ✓ | ✓ |
| Thai | 2548 | × | × | × |
| Language | New Object | Sum-Cov | TSVM | TSVM-Cov |
|---|---|---|---|---|
| Relatively low-resource | ||||
| Thai | Sherlock Holmes | ✓ | ✓ | ✓ |
| Vietnamese | Sherlock Holmes | × | ✓ | ✓ |
| Turkish | Sherlock Holmes | × | ✓ | × |
| Czech | Sherlock Holmes | ✓ | ✓ | ✓ |
| Dutch | Sherlock Holmes | × | ✓ | ✓ |
| High-resource | ||||
| English | Sherlock Holmes | ✓ | × | ✓ |
| Chinese | — | × | ✓ | × |
| French | Sherlock Holmes | ✓ | ✓ | ✓ |
| Spanish | Sherlock Holmes | ✓ | ✓ | ✓ |
| German | Sherlock Holmes | × | ✓ | ✓ |
| Russian | — | × | × | × |
| Portuguese | Sherlock Holmes | ✓ | ✓ | ✓ |
| Language | New Object | Sum-Cov | TSVM | TSVM-Cov |
|---|---|---|---|---|
| Relatively low-resource | ||||
| Thai | 1990 | ✓ | ✓ | ✓ |
| Vietnamese | 1990 | × | ✓ | ✓ |
| Turkish | 1990 | × | ✓ | ✓ |
| Czech | 1990 | × | ✓ | ✓ |
| Dutch | 1990 | ✓ | × | × |
| High-resource | ||||
| English | 1990 | ✓ | × | × |
| Chinese | — | × | × | × |
| French | 1990 | × | ✓ | ✓ |
| Spanish | 1990 | ✓ | ✓ | ✓ |
| German | 1990 | × | × | ✓ |
| Russian | — | × | × | × |
| Portuguese | 1990 | ✓ | ✓ | ✓ |
| Language | New Object | Sum-Cov | TSVM | TSVM-Cov |
|---|---|---|---|---|
| Relatively low-resource | ||||
| Thai | 1818 | × | × | × |
| Vietnamese | 1818 | × | ✓ | ✓ |
| Turkish | 1818 | × | ✓ | ✓ |
| Czech | 1818 | ✓ | ✓ | ✓ |
| Dutch | 1818 | × | × | × |
| High-resource | ||||
| English | 1818 | ✓ | ✓ | ✓ |
| Chinese | — | × | ✓ | × |
| French | 1818 | ✓ | ✓ | ✓ |
| Spanish | 1818 | × | × | × |
| German | 1818 | ✓ | ✓ | ✓ |
| Russian | — | × | × | × |
| Portuguese | 1818 | ✓ | ✓ | ✓ |
Appendix E. Exact-Match Results
| Methods | Edit & Test Languages | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| en | zh | cz | vi | tr | fr | es | de | ru | du | pt | th | avg | |
| Llama3.1-8B, MEMIT | |||||||||||||
| Sum | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| Mean | 5.89 | 0.21 | 4.82 | 2.25 | 4.25 | 3.57 | 3.46 | 7.04 | 0.79 | 4.43 | 4.64 | 0.04 | 3.45 |
| TSVM | 25.93 | 1.93 | 16.21 | 11.57 | 13.71 | 13.89 | 14.21 | 20.18 | 7.18 | 15.07 | 16.00 | 1.21 | 13.09 |
| Sum-Cov | 33.21 | 5.82 | 21.46 | 16.89 | 19.04 | 20.96 | 21.57 | 26.75 | 15.82 | 20.46 | 21.00 | 4.68 | 18.97 |
| Mean-Cov | 0.64 | 0.04 | 0.29 | 0.14 | 0.54 | 0.04 | 0.11 | 0.32 | 0.00 | 0.29 | 0.11 | 0.00 | 0.21 |
| TSVM-Cov | 14.29 | 1.07 | 9.04 | 4.93 | 8.00 | 6.96 | 7.04 | 12.04 | 4.25 | 8.14 | 8.25 | 0.50 | 7.04 |
| Mono | 40.61 | 15.54 | 29.71 | 26.25 | 27.29 | 30.93 | 31.50 | 35.54 | 25.46 | 25.25 | 29.68 | 8.00 | 27.15 |
| Llama3.1-8B, AlphaEdit | |||||||||||||
| Sum | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| Mean | 6.57 | 0.32 | 4.54 | 2.64 | 4.39 | 3.25 | 3.46 | 6.50 | 0.82 | 4.36 | 4.43 | 0.07 | 3.45 |
| TSVM | 23.14 | 1.39 | 14.64 | 10.50 | 12.54 | 12.79 | 12.29 | 18.21 | 5.46 | 13.50 | 13.54 | 0.89 | 11.57 |
| Sum-Cov | 31.21 | 4.82 | 19.86 | 15.00 | 17.04 | 18.07 | 18.96 | 25.21 | 12.54 | 18.54 | 18.64 | 2.68 | 16.88 |
| Mean-Cov | 0.68 | 0.04 | 0.29 | 0.14 | 0.54 | 0.04 | 0.11 | 0.39 | 0.00 | 0.29 | 0.11 | 0.00 | 0.22 |
| TSVM-Cov | 12.39 | 1.36 | 6.11 | 4.07 | 6.36 | 5.04 | 5.32 | 8.75 | 3.11 | 6.39 | 5.71 | 0.57 | 5.43 |
| Mono | 38.79 | 12.43 | 26.46 | 23.11 | 22.93 | 26.61 | 27.14 | 32.14 | 23.07 | 22.39 | 26.46 | 5.50 | 23.92 |
| Methods | Edit & Test Languages | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| en | zh | cz | vi | tr | fr | es | de | ru | du | pt | th | avg | |
| Qwen2.5-7B, MEMIT | |||||||||||||
| Sum | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| Mean | 7.29 | 1.39 | 3.14 | 1.64 | 2.61 | 2.50 | 3.43 | 4.71 | 1.29 | 4.57 | 3.64 | 0.32 | 3.04 |
| TSVM | 27.75 | 12.11 | 11.36 | 5.89 | 6.75 | 12.25 | 16.07 | 16.82 | 9.18 | 14.61 | 14.71 | 2.07 | 12.46 |
| Sum-Cov | 34.18 | 22.29 | 19.32 | 12.86 | 12.54 | 19.68 | 24.68 | 22.11 | 19.04 | 20.57 | 22.82 | 7.07 | 19.76 |
| Mean-Cov | 0.46 | 0.89 | 0.18 | 0.11 | 0.29 | 0.04 | 0.04 | 0.29 | 0.11 | 0.25 | 0.07 | 0.14 | 0.24 |
| TSVM-Cov | 27.46 | 13.32 | 10.54 | 7.82 | 6.32 | 11.50 | 15.07 | 15.36 | 9.25 | 12.89 | 13.32 | 2.21 | 12.09 |
| Mono | 44.50 | 37.25 | 31.54 | 25.68 | 29.18 | 34.71 | 37.29 | 31.68 | 25.71 | 32.46 | 34.39 | 15.18 | 31.63 |
| Qwen2.5-7B, AlphaEdit | |||||||||||||
| Sum | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| Mean | 10.68 | 1.50 | 4.96 | 2.43 | 3.11 | 4.00 | 5.11 | 7.25 | 1.68 | 6.75 | 4.96 | 0.57 | 4.42 |
| TSVM | 31.04 | 16.64 | 15.46 | 8.68 | 10.64 | 14.71 | 18.75 | 19.43 | 12.64 | 17.04 | 18.04 | 3.29 | 15.53 |
| Sum-Cov | 32.36 | 15.79 | 14.21 | 10.68 | 6.75 | 18.50 | 23.04 | 18.61 | 11.57 | 16.75 | 20.04 | 3.39 | 15.97 |
| Mean-Cov | 0.57 | 0.86 | 0.21 | 0.07 | 0.29 | 0.11 | 0.04 | 0.25 | 0.14 | 0.25 | 0.11 | 0.11 | 0.25 |
| TSVM-Cov | 32.11 | 14.43 | 13.46 | 8.89 | 9.29 | 17.46 | 19.96 | 19.50 | 9.54 | 16.43 | 19.11 | 2.61 | 15.23 |
| Mono | 46.00 | 27.89 | 36.75 | 30.68 | 33.89 | 40.86 | 41.79 | 35.32 | 33.29 | 37.21 | 38.79 | 24.93 | 35.62 |
References
- OpenAI. GPT-4 Technical Report; Technical Report; OpenAI: San Francisco, CA, USA, 2023. [Google Scholar]
- Llama Team, Meta AI. The Llama 3 Herd of Models; Technical Report; Meta AI: Menlo Park, CA, USA, 2024. [Google Scholar]
- Qwen Team, Alibaba Group. Qwen2 Technical Report; Technical Report; Alibaba Group: Hangzhou, China, 2024. [Google Scholar]
- Gemini Team, Google. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities; Technical Report; Google: Mountain View, CA, USA, 2025. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Thirty-First Annual Conference on Neural Information Processing Systems (NIPS), Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the Tenth International Conference on Learning Representations, Virtual, 24–29 April 2022. [Google Scholar]
- Mitchell, E.; Lin, C.; Bosselut, A.; Finn, C.; Manning, C.D. Fast Model Editing at Scale. In Proceedings of the Tenth International Conference on Learning Representations, Virtual, 24–29 April 2022. [Google Scholar]
- Yao, Y.; Wang, P.; Tian, B.; Cheng, S.; Li, Z.; Deng, S.; Chen, H.; Zhang, N. Editing large language models: Problems, methods, and opportunities. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 6–10 December 2023. [Google Scholar]
- Meng, K.; Bau, D.; Andonian, A.; Belinkov, Y. Locating and Editing Factual Associations in GPT. In Proceedings of the Thirty-Sixth Conference on Neural Information Processing Systems, Sydney, Australia, 6–12 December 2022. [Google Scholar]
- Meng, K.; Sharma, A.S.; Andonian, A.; Belinkov, Y.; Bau, D. Mass-Editing Memory in a Transformer. In Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Fang, J.; Jiang, H.; Wang, K.; Ma, Y.; Shi, J.; Wang, X.; He, X.; Chua, T.-S. AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models. In Proceedings of the Thirteenth International Conference on Learning Representations, Singapore, 24–28 April 2025. [Google Scholar]
- Wang, J.; Liang, Y.; Sun, Z.; Cao, Y.; Xu, J.; Meng, F. Cross-Lingual Knowledge Editing in Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), Bangkok, Thailand, 11–16 August 2024. [Google Scholar]
- Zhang, X.; Liang, Y.; Meng, F.; Zhang, S.; Chen, Y.; Xu, J.; Zhou, J. Multilingual Knowledge Editing with Language-Agnostic Factual Neurons. In Proceedings of the 31st International Conference on Computational Linguistics, Abu Dhabi, United Arab Emirates, 27–28 January 2025. [Google Scholar]
- Gargiulo, A.A.; Crisostomi, D.; Bucarelli, M.S.; Scardapane, S.; Silvestri, F.; Rodolà, E. Task Singular Vectors: Reducing Task Interference in Model Merging. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2025, Nashville, TN, USA, 11–15 June 2025. [Google Scholar]
- Wang, W.; Haddow, B.; Birch, A. Retrieval-augmented Multilingual Knowledge Editing. arXiv 2023, arXiv:2312.13040. [Google Scholar] [CrossRef]
- De Cao, N.; Aziz, W.; Titov, I. Editing Factual Knowledge in Language Models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic, 7–11 November 2021; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 6491–6506. [Google Scholar]
- Mitchell, E.; Lin, C.; Bosselut, A.; Manning, C.D.; Finn, C. Memory-Based Model Editing at Scale. In Proceedings of the 39th International Conference on Machine Learning, Baltimore, MD, USA, 17–23 July 2022; PMLR: New York, NY, USA, 2022; pp. 15817–15831. [Google Scholar]
- Zheng, C.; Li, L.; Dong, Q.; Fan, Y.; Wu, Z.; Xu, J.; Chang, B. Can We Edit Factual Knowledge by In-Context Learning? In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 6–10 December 2023. [Google Scholar]
- Geva, M.; Schuster, R.; Berant, J.; Levy, O. Transformer Feed-Forward Layers Are Key-Value Memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic, 7–11 November 2021. [Google Scholar]
- Li, X.; Li, S.; Song, S.; Yang, J.; Ma, J.; Yu, J. PMET: Precise Model Editing in a Transformer. In Proceedings of the 38th Annual AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–27 February 2024. [Google Scholar]
- Dai, D.; Dong, L.; Hao, Y.; Sui, Z.; Chang, B.; Wei, F. Knowledge Neurons in Pretrained Transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, 22–27 May 2022; Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 8493–8502. [Google Scholar]
- Khandelwal, A.; Singh, H.; Gu, H.; Chen, T.; Zhou, K. Cross-Lingual Multi-Hop Knowledge Editing. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, FL, USA, 12–16 November 2024; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 11995–12015. [Google Scholar]
- Wortsman, M.; Ilharco, G.; Gadre, S.Y.; Roelofs, R.; Gontijo Lopes, R.; Morcos, A.; Namkoong, H.; Farhadi, A.; Carmon, Y.; Kornblith, S.; et al. Model Soups: Averaging Weights of Multiple Fine-Tuned Models Improves Accuracy without Increasing Inference Time. In Proceedings of the 39th International Conference on Machine Learning, Baltimore, MD, USA, 17–23 July 2022. [Google Scholar]
- Ilharco, G.; Ribeiro, M.T.; Wortsman, M.; Schmidt, L.; Hajishirzi, H.; Farhadi, A. Editing Models with Task Arithmetic. In Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Yadav, P.; Tam, D.; Choshen, L.; Raffel, C.; Bansal, M. TIES-Merging: Resolving Interference When Merging Models. In Proceedings of the Thirty-Seventh Conference on Neural Information Processing Systems, New Orleans, LA, USA, 10–16 December 2023. [Google Scholar]
- Lang, S. Introduction to Linear Algebra; Springer Science & Business Media: Berlin/Heidelberg, Germany, 2012. [Google Scholar]
- Levy, O.; Seo, M.; Choi, E.; Zettlemoyer, L. Zero-shot relation extraction via reading comprehension. In Proceedings of the CoNLL 2017, Vancouver, BC, Canada, 3–4 August 2017. [Google Scholar]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of the 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, BC, Canada, 8–14 December 2019. [Google Scholar]
- Joshi, P.; Santy, S.; Budhiraja, A.; Bali, K.; Choudhury, M. The State and Fate of Linguistic Diversity and Inclusion in the NLP World. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 6282–6293. [Google Scholar]



| Method | Covariance | Merge Operation | Rank-Comp. | One-Line Rationale |
|---|---|---|---|---|
| Sum | Per-language | Summation | No | Naive additive merge; serves as the baseline. |
| Mean | Per-language | Averaging | No | Sum rescaled by to single-edit magnitude. |
| TSVM | Per-language | Low-rank SVD + orthogonal concat. | Yes | Tests whether low-rank structure reduces interference. |
| Sum-Cov | Shared | Summation | No | Shared covariance implicitly normalizes the update scale. |
| Mean-Cov | Shared | Averaging | No | Shared-covariance counterpart of Mean. |
| TSVM-Cov | Shared | Low-rank SVD + orthogonal concat. | Yes | Combines low-rank merging with shared covariance. |
| Methods | Edit & Test Languages | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| en | zh | cz | vi | tr | fr | es | de | ru | du | pt | th | avg | |
| Llama3.1-8B, MEMIT. | |||||||||||||
| Sum | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| Mean | 47.66 | 46.12 | 47.44 | 40.85 | 43.60 | 46.63 | 46.14 | 50.93 | 46.01 | 46.88 | 48.27 | 43.53 | 46.17 |
| TSVM | 62.72 | 51.73 | 56.72 | 50.80 | 52.86 | 57.12 | 56.44 | 61.78 | 53.40 | 56.80 | 57.68 | 47.24 | 55.44 |
| Sum-Cov | 66.41 | 55.23 | 60.01 | 54.76 | 56.89 | 62.03 | 60.94 | 65.28 | 60.37 | 60.73 | 60.92 | 51.83 | 59.62 |
| Mean-Cov | 39.72 | 43.28 | 39.20 | 33.95 | 36.91 | 38.16 | 37.87 | 40.84 | 41.75 | 38.45 | 38.90 | 41.56 | 39.22 |
| TSVM-Cov | 54.33 | 49.38 | 51.05 | 44.25 | 47.81 | 51.28 | 50.77 | 55.67 | 50.85 | 50.98 | 51.76 | 45.32 | 50.29 |
| Mono | 70.42 | 61.14 | 64.63 | 60.17 | 61.62 | 66.17 | 65.16 | 69.21 | 64.45 | 63.23 | 65.25 | 54.75 | 63.85 |
| Llama3.1-8B, AlphaEdit | |||||||||||||
| Sum | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| Mean | 48.10 | 46.62 | 47.40 | 41.43 | 44.04 | 46.74 | 46.53 | 50.97 | 46.21 | 47.13 | 48.07 | 43.82 | 46.42 |
| TSVM | 61.08 | 50.64 | 55.60 | 49.68 | 51.78 | 55.97 | 55.06 | 60.44 | 52.12 | 55.33 | 55.97 | 46.28 | 54.16 |
| Sum-Cov | 65.22 | 54.39 | 58.44 | 53.56 | 55.36 | 60.26 | 59.31 | 64.01 | 57.55 | 59.03 | 59.09 | 49.27 | 57.96 |
| Mean-Cov | 39.73 | 43.20 | 39.16 | 34.03 | 36.91 | 38.26 | 37.78 | 40.88 | 41.68 | 38.38 | 39.08 | 41.64 | 39.23 |
| TSVM-Cov | 52.29 | 49.27 | 48.16 | 43.23 | 46.06 | 48.99 | 48.53 | 53.13 | 49.02 | 48.95 | 49.25 | 44.68 | 48.46 |
| Mono | 69.52 | 59.03 | 62.74 | 58.49 | 59.43 | 64.50 | 63.71 | 67.92 | 63.21 | 61.42 | 63.66 | 52.76 | 62.20 |
| Methods | Edit & Test Languages | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| en | zh | cz | vi | tr | fr | es | de | ru | du | pt | th | avg | |
| Qwen2.5-7B, MEMIT | |||||||||||||
| Sum | 0.01 | 0.00 | 0.05 | 0.00 | 0.01 | 0.00 | 0.01 | 0.00 | 0.00 | 0.00 | 0.00 | 0.02 | 0.01 |
| Mean | 48.94 | 53.16 | 47.50 | 43.06 | 44.36 | 47.52 | 47.01 | 49.81 | 48.44 | 46.38 | 48.03 | 49.04 | 47.77 |
| TSVM | 63.25 | 61.46 | 53.48 | 48.20 | 47.60 | 56.47 | 57.74 | 58.43 | 55.87 | 54.99 | 56.42 | 52.38 | 55.52 |
| Sum-Cov | 66.90 | 66.69 | 59.77 | 53.60 | 53.58 | 60.88 | 62.94 | 62.59 | 63.61 | 59.84 | 61.95 | 59.66 | 61.00 |
| Mean-Cov | 40.38 | 49.45 | 40.87 | 37.19 | 39.10 | 40.45 | 39.31 | 40.51 | 45.17 | 38.49 | 40.69 | 47.38 | 41.58 |
| TSVM-Cov | 62.73 | 61.67 | 53.08 | 49.18 | 47.32 | 55.91 | 57.33 | 57.68 | 55.89 | 53.64 | 56.08 | 52.81 | 55.28 |
| Mono | 71.92 | 72.10 | 65.85 | 61.95 | 63.37 | 68.25 | 69.57 | 67.24 | 67.11 | 65.85 | 67.95 | 63.96 | 67.09 |
| Qwen2.5-7B, AlphaEdit | |||||||||||||
| Sum | 0.00 | 0.01 | 0.00 | 0.00 | 0.00 | 0.01 | 0.00 | 0.00 | 0.02 | 0.00 | 0.00 | 0.02 | 0.00 |
| Mean | 52.18 | 53.82 | 49.45 | 44.46 | 45.31 | 49.50 | 49.07 | 52.37 | 49.33 | 48.88 | 49.93 | 49.17 | 49.46 |
| TSVM | 65.07 | 64.13 | 57.10 | 51.26 | 51.12 | 58.16 | 59.92 | 60.58 | 59.05 | 57.64 | 59.04 | 55.01 | 58.17 |
| Sum-Cov | 63.07 | 59.50 | 50.53 | 46.68 | 39.93 | 56.00 | 58.96 | 55.05 | 54.92 | 52.09 | 55.94 | 51.13 | 53.65 |
| Mean-Cov | 41.05 | 50.05 | 41.36 | 37.48 | 39.59 | 41.09 | 39.77 | 40.97 | 45.43 | 38.83 | 41.16 | 47.54 | 42.03 |
| TSVM-Cov | 65.36 | 62.69 | 55.49 | 50.62 | 49.73 | 58.92 | 59.98 | 59.85 | 57.04 | 56.56 | 59.41 | 53.43 | 57.42 |
| Mono | 72.29 | 66.29 | 68.37 | 64.16 | 65.40 | 70.46 | 71.32 | 68.56 | 69.75 | 67.88 | 69.60 | 68.16 | 68.52 |
| Language | New Object | Sum-Cov | TSVM | TSVM-Cov |
|---|---|---|---|---|
| English | CBS | ✓ | ✓ | ✓ |
| Chinese | CBS | ✓ | ✓ | ✓ |
| Czech | CBS | × | × | × |
| Vietnamese | CBS | ✓ | ✓ | × |
| Turkish | CBS | × | × | × |
| French | CBS | × | ✓ | ✓ |
| Spanish | CBS | ✓ | ✓ | ✓ |
| German | CBS | × | × | ✓ |
| Russian | CBS | ✓ | ✓ | ✓ |
| Dutch | CBS | × | ✓ | × |
| Portuguese | CBS | ✓ | ✓ | ✓ |
| Thai | — | × | × | × |
| Language | New Object | Sum-Cov | TSVM | TSVM-Cov |
|---|---|---|---|---|
| Relatively low-resource | ||||
| Thai | 1956 | × | ✓ | ✓ |
| Vietnamese | 1956 | × | ✓ | ✓ |
| Turkish | 1956 | × | ✓ | ✓ |
| Czech | 1956 | ✓ | ✓ | ✓ |
| Dutch | 1956 | × | × | × |
| High-resource | ||||
| English | 1956 | ✓ | × | × |
| Chinese | — | ✓ | ✓ | ✓ |
| French | 1956 | × | ✓ | ✓ |
| Spanish | 1956 | ✓ | ✓ | ✓ |
| German | 1956 | ✓ | ✓ | ✓ |
| Russian | — | × | × | × |
| Portuguese | 1956 | ✓ | × | × |
| Backbone + Method | Merge | th | vi | tr | cz | du | LRL | HRL |
|---|---|---|---|---|---|---|---|---|
| Llama3.1-8B + MEMIT | TSVM | 47.2 | 50.8 | 52.9 | 56.7 | 56.8 | 52.9 | 57.3 |
| Sum-Cov | 51.8 | 54.8 | 56.9 | 60.0 | 60.7 | 56.8 | 61.6 | |
| TSVM-Cov | 45.3 | 44.3 | 47.8 | 51.0 | 51.0 | 47.9 | 52.0 | |
| Mono | 54.8 | 60.2 | 61.6 | 64.6 | 63.2 | 60.9 | 66.0 | |
| Llama3.1-8B + AlphaEdit | TSVM | 46.3 | 49.7 | 51.8 | 55.6 | 55.3 | 51.7 | 55.9 |
| Sum-Cov | 49.3 | 53.3 | 55.5 | 58.5 | 59.0 | 55.1 | 60.0 | |
| TSVM-Cov | 44.7 | 43.2 | 46.1 | 48.2 | 49.0 | 46.2 | 50.1 | |
| Mono | 52.8 | 58.5 | 59.4 | 62.7 | 61.4 | 59.0 | 64.5 | |
| Qwen2.5-7B + MEMIT | TSVM | 52.4 | 48.2 | 47.6 | 53.5 | 55.0 | 51.3 | 58.5 |
| Sum-Cov | 59.7 | 53.6 | 53.6 | 59.8 | 59.8 | 57.3 | 63.7 | |
| TSVM-Cov | 52.8 | 49.2 | 47.3 | 53.1 | 53.6 | 51.2 | 58.2 | |
| Mono | 64.0 | 61.9 | 63.4 | 65.9 | 65.8 | 64.2 | 69.2 | |
| Qwen2.5-7B + AlphaEdit | TSVM | 55.0 | 51.3 | 51.1 | 57.1 | 57.6 | 54.4 | 60.8 |
| Sum-Cov | 51.1 | 46.7 | 39.9 | 50.5 | 52.1 | 48.1 | 57.6 | |
| TSVM-Cov | 53.4 | 50.6 | 49.7 | 55.5 | 56.6 | 53.2 | 60.5 | |
| Mono | 68.2 | 64.2 | 65.4 | 68.4 | 67.9 | 66.8 | 69.8 |
| Backbone + Method | Kept | fidjoint | fidsolo | Joint/Solo | |
|---|---|---|---|---|---|
| Llama3.1-8B + MEMIT | – | 0.72 | 0.49 | 1.5 | |
| Llama3.1-8B + AlphaEdit | 0.91 | 0.77 | 0.56 | 1.4 | |
| Qwen2.5-7B + MEMIT | – | 0.63 | 0.34 | 1.9 | |
| Qwen2.5-7B + AlphaEdit | 0.60 | 0.39 | 0.09 | 4.3 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Lee, K.; Shin, K.-Y.; Lee, J.-H.; Suh, Y.-J. Merging Methods for Multilingual Knowledge Editing for Large Language Models: An Empirical Odyssey. Electronics 2026, 15, 2747. https://doi.org/10.3390/electronics15122747
Lee K, Shin K-Y, Lee J-H, Suh Y-J. Merging Methods for Multilingual Knowledge Editing for Large Language Models: An Empirical Odyssey. Electronics. 2026; 15(12):2747. https://doi.org/10.3390/electronics15122747
Chicago/Turabian StyleLee, Kunil, Ki-Young Shin, Jong-Hyeok Lee, and Young-Joo Suh. 2026. "Merging Methods for Multilingual Knowledge Editing for Large Language Models: An Empirical Odyssey" Electronics 15, no. 12: 2747. https://doi.org/10.3390/electronics15122747
APA StyleLee, K., Shin, K.-Y., Lee, J.-H., & Suh, Y.-J. (2026). Merging Methods for Multilingual Knowledge Editing for Large Language Models: An Empirical Odyssey. Electronics, 15(12), 2747. https://doi.org/10.3390/electronics15122747

