1. Introduction
Large language model (LLM)-based assistants are increasingly used to make complex information accessible through natural-language interaction. In sociological research, such systems can help analysts, decision makers, and non-specialist users interpret findings stored in relational databases, coded response fields, and reference tables. However, generative access to survey data creates risks that extend beyond linguistic fluency. Survey databases may contain demographic characteristics, geographical information, political or religious views, perceptions of social justice, health-related information, and free-text responses. When those data are transformed into readable narratives, the system may reveal sensitive attributes, enable re-identification through combinations of quasi-identifiers, reproduce harmful framing, or present an individual response as a population-level conclusion.
Unlike general web text, survey responses are collected from human participants for defined research purposes and under expectations of confidentiality. The fact that an LLM can produce a coherent description does not imply that every available attribute should be reproduced. Automated generation changes the accessibility of the data: information previously available only to authorised analysts can be reformulated and delivered directly through a conversational interface. This transformation creates an obligation to minimise unnecessary disclosure, preserve respondent autonomy, avoid unsupported generalisations, and keep released text consistent with the original purpose of data collection.
Several risks make the problem consequential. Direct identifiers may occur in source fields or free-text responses. Even after their removal, combinations such as exact age, locality, survey wave, and rare response content may increase indirect identification risk. Classical privacy models, including
k-anonymity and
ℓ-diversity, show that removing direct identifiers alone is insufficient when attribute combinations reveal sensitive information [
1,
2]. Language models may also memorise and reproduce fragments of training data [
3]. In addition, outputs may contain hallucinated claims, discriminatory framing, or unwarranted social inference [
4,
5].
This study proposes SocioTable-KZ, a privacy-aware bilingual pipeline for generating natural-language interpretations from interconnected sociological survey data. The expression privacy-aware is used deliberately: the method implements risk-reducing controls but does not claim differential privacy, cryptographic confidentiality, or a proven upper bound on re-identification risk. The pipeline combines protected relational-data preparation, deterministic JoinGraph serialisation, language-specific QLoRA adaptation, a domain-specific safety taxonomy, constrained rewriting, and controlled release.
The revised study makes five contributions:
A deterministic JoinGraph representation for minimised, auditable serialisation of linked survey tables;
A bilingual QLoRA generation procedure with independent Kazakh and Russian adapters and held-out automatic evaluation;
A documented two-stage safety annotation protocol with three expert annotators and a conservative Safe–Sensitive–Unsafe taxonomy;
Operational rules for identifier exclusion, quasi-identifier generalisation, human review, audit logging, and release control;
An explicit evidence map distinguishing empirically evaluated components from design choices that still require controlled ablation and formal privacy evaluation.
The remainder of the paper is organised as follows.
Section 2 reviews relational data-to-text generation, low-resource multilingual adaptation, and safety and privacy research.
Section 3 describes data provenance, consent, target construction, JoinGraph serialisation, model adaptation, safety annotation, and deployment controls.
Section 4 reports generation, training, safety, and operational-control results.
Section 5 discusses architectural implications and limitations.
Section 6 concludes the paper.
Author Contributions
Conceptualization, A.O. and M.M.; methodology, A.O., Z.Z. and A.Z.; software, A.O., Z.Z. and A.Z.; validation, A.O. and M.M.; formal analysis, A.O.; data curation, A.O., Z.Z. and A.Z.; writing—original draft preparation, A.O.; writing—review and editing, A.O. and M.M.; supervision, M.M.; project administration, A.O. All authors have read and agreed to the published version of the manuscript.
Funding
This computational research was funded by the Science Committee of the Ministry of Science and Higher Education of the Republic of Kazakhstan, grant BR24993001, “Creation of a large language model (LLM) to support the Kazakh language and technological progress”.
Institutional Review Board Statement
The study involved the secondary analysis of previously collected and anonymized non-biomedical sociological survey data. No new participants were recruited, no additional primary data were collected, and the authors had no access to direct personal identifiers. The Quality and Ethics Committee of the Institute of Philosophy, Political Science and Religious Studies determined that separate prospective ethics approval was not required for this secondary analysis (Protocol No. 5, 31 August 2025; official determination issued on 3 August 2026, Ref. No. 01-11/248-III). The data were processed in accordance with Article 17 of the Law of the Republic of Kazakhstan No. 94-V of 21 May 2013, “On Personal Data and Their Protection”.
Informed Consent Statement
Informed consent was obtained from all respondents included in the original sociological surveys. Before the substantive questionnaire, respondents received information about the organisation conducting the survey, the research purpose, voluntary participation, confidentiality, scientific and academic-publication use, and the right to refuse or discontinue. Consent was actively recorded through the mandatory question “Do you agree to participate voluntarily in this survey?” Only respondents who selected “Yes, I agree” proceeded; those who selected “No, I do not agree” were not included. A separate signed consent sheet was not used. The data-providing institution confirmed this procedure in writing on 30 July 2026.
Data Availability Statement
Restrictions apply to the respondent-level sociological survey data. The authors received an anonymised analytical dataset under institutional data-sharing conditions and are not authorised to redistribute it. Source code and aggregate supporting outputs may be shared after institutional de-identification and release review. Access to approved anonymised derivatives may be requested from the data provider and is subject to ethical, legal, and institutional approval.
Acknowledgments
The authors gratefully acknowledge the Institute of Philosophy, Political Science and Religious Studies of the Committee of Science, Ministry of Science and Higher Education of the Republic of Kazakhstan, for providing the anonymised analytical survey dataset and the official confirmation of the informed consent procedure and authorised scientific use of the data. The underlying sociological survey data were collected within project BR28713047 “Artificial Intelligence and the Ethics of Social Justice: A Conceptual Analysis of Opportunities and Risks in the Processes of Societal Modernization in Kazakhstan”.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Sweeney, L. k-Anonymity: A Model for Protecting Privacy. Int. J. Uncertain. Fuzziness Knowl.-Based Syst. 2002, 10, 557–570. [Google Scholar] [CrossRef]
- Machanavajjhala, A.; Kifer, D.; Gehrke, J.; Venkitasubramaniam, M. ℓ-Diversity: Privacy Beyond k-Anonymity. ACM Trans. Knowl. Discov. Data 2007, 1, 3. [Google Scholar] [CrossRef]
- Carlini, N.; Tramèr, F.; Wallace, E.; Jagielski, M.; Herbert-Voss, A.; Lee, K.; Roberts, A.; Brown, T.; Song, D.; Erlingsson, Ú.; et al. Extracting Training Data from Large Language Models. In Proceedings of the 30th USENIX Security Symposium (USENIX Security 21); USENIX Association: Berkeley, CA, USA, 2021; pp. 2633–2650. [Google Scholar]
- Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.; Madotto, A.; Fung, P. Survey of Hallucination in Natural Language Generation. ACM Comput. Surv. 2023, 55, 248. [Google Scholar] [CrossRef]
- Weidinger, L.; Mellor, J.; Rauh, M.; Griffin, C.; Uesato, J.; Huang, P.-S.; Cheng, M.; Glaese, M.; Balle, B.; Kasirzadeh, A.; et al. Ethical and Social Risks of Harm from Language Models. arXiv 2021, arXiv:2112.04359. [Google Scholar]
- Gatt, A.; Krahmer, E. Survey of the State of the Art in Natural Language Generation. J. Artif. Intell. Res. 2018, 61, 65–170. [Google Scholar] [CrossRef]
- Parikh, A.; Wang, X.; Gehrmann, S.; Faruqui, M.; Dhingra, B.; Yang, D.; Das, D. ToTTo: A Controlled Table-to-Text Generation Dataset. In Proceedings of EMNLP 2020; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 1173–1186. [Google Scholar] [CrossRef]
- Dušek, O.; Kasner, Z. Evaluating Semantic Accuracy of Data-to-Text Generation with Natural Language Inference. In Proceedings of INLG 2020; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 131–137. [Google Scholar] [CrossRef]
- Kasner, Z.; Dušek, O. Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation. In Proceedings of ACL 2024; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 12045–12072. [Google Scholar] [CrossRef]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the ICLR 2022, International Conference on Learning Representations, Virtual, 25–29 April 2022. [Google Scholar]
- Dettmers, T.; Pagnoni, A.; Holtzman, A.; Zettlemoyer, L. QLoRA: Efficient Finetuning of Quantized LLMs. In Advances in Neural Information Processing Systems; Neural Information Processing Systems Foundation, Inc.: La Jolla, CA, USA, 2023; Volume 36, pp. 10088–10115. [Google Scholar] [CrossRef]
- Togmanov, M.; Mukhituly, N.; Turmakhan, D.; Mansurov, J.; Goloburda, M.; Sakip, A.; Xie, Z.; Wang, Y.; Syzdykov, B.; Laiyk, N.; et al. KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan. In Proceedings of ACL 2025; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 14403–14416. [Google Scholar] [CrossRef]
- Inan, H.; Upasani, K.; Chi, J.; Rungta, R.; Iyer, K.; Mao, Y.; Tontchev, M.; Hu, Q.; Fuller, B.; Testuggine, D.; et al. Llama Guard: LLM-Based Input-Output Safeguard for Human-AI Conversations. arXiv 2023, arXiv:2312.06674. [Google Scholar]
- Zhang, Z.; Lei, L.; Wu, L.; Sun, R.; Huang, Y.; Long, C.; Liu, X.; Lei, X.; Tang, J.; Huang, M. SafetyBench: Evaluating the Safety of Large Language Models. In Proceedings of ACL 2024; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 15537–15553. [Google Scholar] [CrossRef]
- Wang, W.; Tu, Z.; Chen, C.; Yuan, Y.; Huang, J.-T.; Jiao, W.; Lyu, M. All Languages Matter: On the Multilingual Safety of LLMs. In Findings of ACL 2024; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 5865–5877. [Google Scholar] [CrossRef]
- Li, X.; Tramèr, F.; Liang, P.; Hashimoto, T. Large Language Models Can Be Strong Differentially Private Learners. In Proceedings of the ICLR 2022; International Conference on Learning Representations, Virtual, 25–29 April 2022. [Google Scholar]
- Jobin, A.; Ienca, M.; Vayena, E. The Global Landscape of AI Ethics Guidelines. Nat. Mach. Intell. 2019, 1, 389–399. [Google Scholar] [CrossRef]
- Floridi, L.; Cowls, J. A Unified Framework of Five Principles for AI in Society. Harv. Data Sci. Rev. 2019, 1. [Google Scholar] [CrossRef]
- Mittelstadt, B. Principles Alone Cannot Guarantee Ethical AI. Nat. Mach. Intell. 2019, 1, 501–507. [Google Scholar] [CrossRef]
- Tabassi, E. Artificial Intelligence Risk Management Framework (AI RMF 1.0); NIST AI 100-1; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2023. [CrossRef]
- Republic of Kazakhstan. Law No. 94-V “On Personal Data and Their Protection”; Parliament of the Republic of Kazakhstan: Astana, Kazakhstan, 2013.
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |