1. Introduction
Digital sustainability is concerned not only with the long-term accessibility and efficiency of technological systems but also with the capacity of digital environments to sustain social, cultural, and linguistic diversity. Language technologies now mediate access to education, information, public services, and cultural participation. Consequently, a language variety that cannot be reliably recognized, parsed, searched, or generated by digital systems faces a new form of marginalization: it may remain vibrant in everyday speech while becoming progressively less visible in the infrastructures through which contemporary knowledge circulates.
This problem is especially consequential for dialects. NLP systems are commonly trained on corpora dominated by standardized written registers, national news, institutional texts, and centrally produced educational materials. High performance on such sources may therefore coexist with substantial failure on conversational forms, regional morphosyntax, locally embedded vocabulary, and pragmatic particles. The resulting disparity is not simply a technical robustness problem. It affects which speakers can interact successfully with digital systems, which linguistic forms are preserved in searchable data, and which varieties are treated as legitimate in AI-supported educational environments.
The United Nations Sustainable Development Goal 4 calls for inclusive and equitable quality education, while Goal 10 emphasizes the reduction in inequalities [
1]. These commitments increasingly extend to digital environments, where access is shaped not only by connectivity but also by whether systems understand the language practices of their users. The United Nations Global Digital Compact similarly frames the desired digital future as inclusive, open, sustainable, fair, safe, and secure [
2]. From this perspective, linguistic inclusion is part of sustainable digital development rather than a peripheral concern.
The present study examines this issue through Turkish dialects. Turkish is morphologically rich and exhibits extensive regional variation in phonology, lexicon, inflection, word order, discourse particles, and pragmatic meaning. Yet most computational resources continue to privilege SWT. Earlier Turkish dialect-recognition studies demonstrated that regional speech can be classified with machine-learning methods [
3], but they generally focused on restricted speaker samples or acoustic features. Such work is valuable for identifying varieties, although classification alone does not show whether a system can interpret the linguistic structure and social meaning of regional speech.
International research has shown that NLP resources and performance are distributed unevenly across languages and varieties. Studies of linguistic diversity in NLP demonstrate that technological coverage remains concentrated in a limited set of well-resourced languages, while cross-linguistic evaluations identify systematic disparities in foundational and user-facing tasks [
4,
5]. Data statements have therefore been proposed to make population, variety, provenance, and intended-use assumptions explicit [
6]. Fairness-oriented scholarship further cautions that “bias” should not be inferred from a performance difference alone; the relevant harm, affected speakers, and sociolinguistic hierarchy must be specified [
7]. Community-centred work similarly stresses that language technology should expand speakers’ capabilities and respect local authority over linguistic knowledge [
8]. These perspectives establish the theoretical connection between technical coverage, representational justice, and the sustainability-oriented objectives of the present study.
However, digital preservation should not be equated with automatic standardization. Converting expressions from Turkish dialects into their SWT equivalents may facilitate retrieval, but it can also remove pragmatic force, local memory, interactional meaning, and identity. Kelly-Holmes [
9] warns against reducing living language to what is computationally measurable, while Erdocia et al. [
10] argue that language must be approached as a social practice rather than merely as a dataset. Digital linguistic sustainability therefore requires a balance between interoperability and continuity: systems should make dialect forms computationally accessible without treating them as defective approximations of a single norm.
1.1. A Three-Dimensional Framework for Digital Linguistic Sustainability
This study uses digital linguistic sustainability as an integrative framework comprising three interdependent dimensions.
First, technical robustness refers to a system’s capacity to process regional and standard forms consistently across tasks such as tokenization, part-of-speech tagging, dependency analysis, and semantic comparison. A technically sustainable system should not lose basic functionality when users depart from the dominant written register.
Second, cultural and representational continuity concerns whether digital systems preserve regionally embedded meanings, discourse functions, and linguistic identities. Representation is sustainable when speakers’ forms remain visible and interpretable rather than being erased through automatic correction or decontextualized data extraction.
Third, educational inclusion concerns the effects of language technologies on learners and teachers. AI-supported tools become pedagogically unsustainable when regional speech is repeatedly marked as wrong, unintelligible, or inferior. Conversely, culturally responsive systems can support dialect awareness, metalinguistic reflection, and equitable participation in education.
These dimensions are mutually dependent. Technical performance without cultural interpretation may produce efficient but exclusionary systems; cultural documentation without usable technological infrastructure may remain inaccessible; and educational deployment without representational safeguards may reinforce linguistic insecurity.
1.2. Sustainable and Inclusive Education as a Use Context
AI-supported education increasingly relies on language-sensitive functions, including automated writing feedback, speech recognition, intelligent tutoring, conversational assistance, reading support, and assessment. These applications are not linguistically neutral. When systems are optimized primarily for SWT, learners who use regional phonological, lexical, morphosyntactic, or pragmatic forms may receive less accurate transcription, inappropriate correction, or misleading evaluation. A technical performance gap can therefore become an educational participation gap.
From a sustainable-education perspective, inclusion requires more than access to digital infrastructure. Learners must also be able to interact with educational technologies without their habitual language being treated automatically as noise, error, or deficiency. This requirement connects the present study to SDG 4, particularly inclusive and equitable quality education, and to SDG 10, because linguistic variation can become an overlooked mechanism through which digital systems reproduce social and regional inequalities. Sustainable educational technology should distinguish between pedagogically relevant errors and legitimate linguistic variation, communicate uncertainty transparently, and allow teachers and learners to contest or contextualize automated judgments.
The study does not assume that all dialect forms should replace SWT in formal instruction. Rather, it argues for bidialectal and variation-aware educational design: systems may support acquisition of the standard variety while recognizing Turkish dialect forms accurately, explaining contrasts without stigmatization, and preserving the learner’s linguistic identity. This approach positions model robustness, culturally responsive feedback, and educator oversight as mutually reinforcing conditions of sustainable AI-supported education.
1.3. Research Gap and Contribution
Existing studies have generally examined either the technical recognition and normalization of dialects or the sociolinguistic consequences of standard-language dominance. Research on Turkish NLP has also concentrated mainly on standardized written resources, treebanks, morphology, and benchmark datasets [
11,
12,
13,
14]. The relationship among task-specific processing difficulties, speakers’ interpretations, and sustainable educational technology remains less fully developed. The present study brings these strands together through a linguistically grounded and sustainability-oriented perspective. It does not seek to develop or rank NLP architectures as a conventional computational benchmark; rather, it examines how transcribed Turkish dialect forms may be processed, represented, or overlooked in contemporary language technologies.
This study addresses that gap by combining a geographically distributed spoken corpus, task-specific NLP outputs, and qualitative interview material. Its contribution lies in relating recurrent processing observations to the linguistic properties of Turkish dialects and then considering their possible significance for cultural continuity, educational inclusion, and representational justice. The computational outputs are not treated as autonomous proof of educational or social effects, and the systems are not ranked as universally superior or inferior.
The sustainability contribution is therefore substantive rather than metaphorical. Regional language data constitute a cultural resource whose continued usability depends on digital infrastructures. A system that works well only for the dominant standard may remain operational while being socially unsustainable: it transfers the costs of technological exclusion to speakers whose linguistic practices are underrepresented. Conversely, a sustainable language technology should support long-term digital access, retain culturally situated meaning, and distribute the benefits of AI-supported education more equitably. This interpretation connects the study to SDG 4 (inclusive and equitable quality education), particularly Target 4.5 on disparities and vulnerable groups, and SDG 10 (reduced inequalities), particularly Target 10.2 on social inclusion. Participatory governance is treated as a cross-cutting principle that informs all three dimensions rather than as a separate dimension of the framework.
Table 1 summarizes the sustainability framework and the empirical indicators used in the study.
1.4. Research Questions
The study addresses the following questions:
What task-specific differences emerge when the examined NLP environments process SWT reference material and transcribed Turkish dialect forms?
Which morphological, syntactic, lexical, semantic, and pragmatic patterns recur in the processing observations?
How do linguists and dialect speakers interpret the representation of Turkish dialects in AI-based language technologies?
How can the task-specific linguistic observations and qualitative themes be interpreted together within a digital linguistic sustainability framework?
What potential design and governance principles follow from this linguistically grounded and interdisciplinary analysis?
The research questions are addressed through complementary analytical pathways. Questions 1 and 2 draw on the linguistic examination of task-specific outputs in the transcribed dialect data; Question 3 draws on semi-structured interviews and thematic interpretation; Question 4 is considered through interpretive integration; and Question 5 is addressed through the synthesis developed in the Discussion. This alignment establishes a direct connection between the research questions, the corresponding data sources, and the interpretive stages of the study.
2. Materials and Methods
This study adopts a linguistically grounded, exploratory mixed-methods design. The speech-data pathway identifies task-specific processing differences and recurrent linguistic patterns, whereas the interview pathway provides an interpretive perspective on representation, trust, educational inclusion, and technological justice. The two pathways are integrated only at the interpretation stage through the digital linguistic sustainability framework. Convergence is noted when both forms of material point to a related representational concern; divergence is retained where participant interpretations cannot be inferred from task-specific outputs alone. The findings are interpreted as task- and data-sensitive observations within a linguistically grounded exploratory mixed-methods framework. Mixed-methods integration is presented through a joint display following established guidance on integration at the interpretation and reporting stages [
15].
Figure 1 presents the workflow of the exploratory mixed-methods design.
The workflow distinguishes the speech-data and interview pathways before their interpretive integration. The speech-data pathway proceeds from Praat-assisted examination and IPA-based transcription to corpus preparation, annotation, task-specific NLP, and linguistic interpretation. The interview pathway proceeds through semi-structured interviews and thematic coding. Their convergence and divergence are considered only at the final interpretive stage.
2.1. Data Collection and Corpus Preparation
- (a)
Geographical Coverage, Participant Access, and Eligibility
Oral data were collected from 14 provinces representing Türkiye’s seven geographical regions: Kırklareli and Balıkesir (Marmara); Aydın and Denizli (Aegean); Adana and Mersin (Mediterranean); Konya and Kayseri (Central Anatolia); Erzurum and Ardahan (Eastern Anatolia); Şırnak and Mardin (Southeastern Anatolia); and Trabzon and Giresun (Black Sea). Two provinces were selected from each geographical region to ensure broad geographical coverage. The selection also considered the feasibility of reaching eligible speakers through academics, colleagues, and local contacts in the relevant provinces. These contacts facilitated communication with potential participants but did not take part in the recording, transcription, annotation, or analysis processes. The resulting geographical distribution provided a broad exploratory sample rather than exhaustive or statistically representative profiles of individual provincial dialects.
The speech corpus comprised 100 participants between the ages of 50 and 65, including 50 women and 50 men (mean age = 63). During recruitment, priority was initially given to speakers aged 60–65 because long-term local residence and sustained everyday language use were expected to support the retention of salient dialect features. Since this narrower age criterion could not be applied consistently across all 14 provinces, the range was extended to 50–65 years. Participants’ educational backgrounds predominantly consisted of primary schooling or non-literacy, with limited representation from secondary or higher education. Eligibility required participants to have been born and raised in the relevant province and not to have permanently resided elsewhere. These criteria supported sustained exposure to the local dialect, although no individual participant was treated as representing the complete dialect profile of a province or region.
Data collection was conducted online. Prospective participants received written information explaining the purpose of the research, the voluntary nature of participation, confidentiality, the scientific use of the data, and their right to decline or withdraw. Before any questions were asked or recording began, the procedure was explained verbally, and explicit verbal informed consent was obtained. Names, signatures, and other direct identifiers were not collected. The recordings and transcripts were stored using participant codes and labelled only with age, gender, and province information.
Recordings were examined with Praat and transcribed through a shared International Phonetic Alphabet (IPA)-based protocol. The transcription retained the phonetic detail required to document linguistically relevant realizations and preserve dialect forms during their conversion into textual NLP input. The examined NLP systems received textual material derived from the transcriptions rather than acoustic recordings. The resulting analysis therefore focused on the processing of transcribed dialect forms within text-based NLP environments.
Because no existing annotated resource adequately represented the corpus, the researchers prepared and annotated the textual data. Approximately 3500 structures were classified using shared linguistic criteria informed by Universal Dependencies, including UPOS, lemma, and dependency labels where relevant. A Monte Carlo-based randomization procedure was used to organize the annotation batches and reduce possible order and allocation effects. Two annotators independently reviewed the categorical annotations, and disagreements were resolved through adjudication. Cohen’s κ values of 0.78–0.84 indicate agreement between the human annotators rather than model performance.
Evaluation examples were selected from a speaker-grouped and province-stratified subset of the annotated corpus. Material from the same speaker was kept separate from the remaining corpus material, while the provincial distribution was retained in the evaluation subset. The task-specific observations reported in this article were drawn from this subset, thereby limiting overlap across speakers and maintaining geographical coverage.
- (b)
Textual Input and Task-Specific NLP
The analytical procedure consisted of four successive stages: audio examination, transcription, human annotation, and task-specific NLP. The spaCy- and Stanza-based configurations were used to examine morphosyntactic processing through POS-related and dependency-related outputs. BERTurk-based representations were used for an exploratory examination of contextual similarity between selected dialect sentences and their SWT counterparts. The systems retained their general-purpose Turkish configurations, allowing the study to examine how existing NLP environments respond to dialect forms without prior dialect-specific adaptation. Outputs were evaluated separately according to the linguistic task supported by each system and subsequently interpreted from a linguistic perspective.
The SWT reference material was prepared according to established orthographic and morphosyntactic conventions and provided a consistent interpretive baseline for examining the selected dialect forms. Its purpose was to clarify morphological, syntactic, lexical, and contextual differences between the dialect expressions and their SWT counterparts. Because the SWT reference material and the transcribed dialect data differ in modality and register, the observed contrasts were interpreted as task- and data-sensitive patterns. Possible effects of punctuation, disfluency, speaker age, and genre were also considered. A modality-matched spoken-SWT corpus would enable a more controlled computational comparison in future research.
Contextual similarity between selected dialect–SWT sentence pairs was explored through cosine-similarity scores obtained from BERTurk-based representations. The reported values were used to identify sentence pairs requiring closer linguistic examination, particularly where dialect morphology, lexical choice, discourse particles, or pragmatic meaning affected contextual alignment. No inferential statistical significance tests were conducted on these scores; accordingly, the values are not interpreted as demonstrating statistically significant differences among sentence pairs. A fully reproducible semantic evaluation would require a separately documented protocol specifying the model checkpoint, embedding layer, pooling strategy, pair-validation procedure, calibration settings, and complete analysis code. Within the exploratory scope of the present study, the scores are used solely as descriptive indicators to guide the close linguistic examination of selected sentence pairs.
2.2. Semi-Structured Interviews
In the second phase, focused semi-structured video interviews were conducted with 28 participants from two complementary groups: 14 linguists and 14 dialect speakers. These groups were included to bring together professional assessments and speakers’ experiences concerning the representation of Turkish dialects in AI-based language technologies.
The linguists were recruited through the researchers’ professional and academic networks. Participation was voluntary and based on their willingness to discuss the subject. No additional quotas were applied regarding age, gender, academic seniority, institutional affiliation, or geographical location.
The dialect-speaker group comprised 14 participants, with one speaker recruited from each province included in the study. These participants were selected independently of the speech corpus, and none had contributed to the recordings used in corpus construction. A criterion-based purposive sampling strategy was used to recruit locally rooted speakers with sustained experience of the relevant dialect [
16]. Priority was given to women aged 60–65 who had been born and raised in the relevant province and had not permanently resided elsewhere. Where an eligible woman could not be reached, a male participant within the same age range and with the same residence history was recruited. This selection produced a geographically distributed interview group whose members had sustained experience of local dialect use.
Each interview lasted approximately 5–8 min. The interviews followed a focused structure organized around four interrelated areas: (1) perceptions of dialect representation in educational settings; (2) experiences with and awareness of dialect processing in AI-based tools; (3) cultural belonging and perceptions of linguistic exclusion; and (4) trust in digital language technologies and patterns of use. These thematic areas provided a common framework for both participant groups, while follow-up prompts were adapted to the participants’ professional knowledge or lived linguistic experiences.
All interviews were conducted in accordance with the approved ethical protocol. Participants were informed about the purpose of the research, the voluntary nature of participation, confidentiality, and the intended scientific use of the material. Recording began after explicit verbal informed consent had been obtained. The recordings were subsequently transcribed, anonymized, and prepared for thematic content analysis. All interview quotations presented in English were translated from Turkish by the authors.
2.3. Thematic Content Analysis
The interview transcripts were examined through thematic content analysis. The coding procedure combined deductive categories derived from the digital linguistic sustainability framework with inductive refinement based on the participants’ accounts. Through iterative reading and comparison, the codebook was organized around four higher-order themes: (1) trust in language technologies; (2) cultural visibility and identity; (3) possible educational misalignment; and (4) technological justice and participation. Access and equality, linguistic diversity, and ethical responsibility were retained as cross-cutting codes that informed more than one thematic area.
One researcher conducted the qualitative coding manually through successive rounds of reading, initial coding, category refinement, re-coding, and analytic memoing. A stabilized codebook and an audit trail were maintained to document coding decisions and subsequent revisions. Accounts that differed from the dominant thematic tendency were retained during interpretation, enabling convergent and divergent perspectives to remain visible within the analysis.
The qualitative coding procedure was analytically separate from the linguistic reference-annotation process described in
Section 2.1. The linguistic annotations provided categorical reference material for the task-specific examination of model outputs, whereas thematic coding was used to interpret participants’ accounts of representation, trust, education, and technological justice. Consistency in the qualitative analysis was supported through iterative re-coding, the stabilized codebook, the audit trail, and consideration of accounts that did not fully converge with the principal interpretive patterns.
2.4. Interpretive Comparative Approach
In the final stage, the task-specific linguistic observations and qualitative themes were brought together through an interpretive joint display [
15]. The comparison examined three principal relationships: (1) recurrent morphosyntactic processing difficulties in relation to participants’ accounts of trust in language technologies; (2) standard-centred outputs in relation to perceptions of linguistic visibility and cultural representation; and (3) dialect-related educational concerns in relation to prospective principles for variation-aware technological design.
The joint display was used to identify points of convergence, divergence, and complementary interpretation between the two analytical pathways. Linguistic observations indicate the forms and contexts in which processing difficulties recur, while participant accounts clarify how technological representation may be experienced and interpreted. Their integration provides a basis for discussing the potential implications of the observed patterns for digital linguistic sustainability, educational inclusion, and representational justice while preserving the distinct evidential contribution of each dataset.
4. Discussion
The findings address Questions 4 and 5 by bringing the two analytical pathways into a common interpretive framework. The task-specific observations locate recurrent difficulties in the processing of dialect morphology, syntax, lexis, and pragmatics, while the interviews indicate how linguistic recognition relates to cultural visibility, trust, and possible educational experiences. Considered together, these findings show that digital linguistic sustainability involves more than the technical availability of a language in digital systems. It also concerns whether linguistic variation is processed with sufficient contextual sensitivity and whether speakers can participate in digital environments without their forms of expression being routinely reduced to SWT.
4.1. Technical Robustness as a Condition of Sustainability
The lower POS agreement values observed for the transcribed dialect material suggest that performance reported for SWT resources cannot be assumed to extend uniformly to dialect data. The recurrent difficulties involve compound verb forms, clipped or fused suffixes, flexible word order, dialect vocabulary, and discourse-sensitive particles. These features are particularly consequential in Turkish because grammatical relations and distinctions are frequently encoded within morphologically complex word forms [
13,
14].
The findings correspond with previous research documenting unequal technological support across languages and dialects [
4,
5]. They also support work in dialect NLP that emphasizes the importance of morphology-sensitive processing, transparent normalization procedures, and dedicated evaluation resources [
17,
18]. From this perspective, technical robustness refers not simply to high aggregate performance but to a system’s capacity to process legitimate linguistic variation without consistently treating it as noise, irregularity, or error.
The present findings also illustrate why model evaluation requires linguistic interpretation. A tagging or parsing difference may arise from the interaction of corpus composition, orthographic representation, tokenization, model coverage, and the grammatical characteristics of the form itself. Consequently, the significance of an output depends on the linguistic function affected and the context in which the system is used. Fairness-oriented evaluation can contribute to this interpretation by identifying the speakers concerned, the relevant use conditions, and the possible consequences of recurrent processing difficulties [
6,
7,
8].
4.2. Cultural Continuity and Computational Standardization
The semantic and pragmatic observations extend the discussion beyond formal accuracy. Dialect items such as gari, hele, and uşak carry meanings that depend on speaker stance, interpersonal relations, discourse sequence, affect, and local communicative conventions. Standard paraphrases may preserve the general propositional content of an utterance while weakening the pragmatic or cultural information conveyed by the dialect form.
This distinction is central to cultural continuity. Digital documentation can make dialects searchable, analysable, and accessible to future generations. However, documentation based primarily on standardization may detach linguistic forms from the contexts in which they acquire social meaning. The interview material supports this interpretation: participants associated recognition with the visibility and legitimacy of their ways of speaking, rather than viewing it solely as a matter of software functionality.
Digital linguistic sustainability therefore requires contextual representation as well as formal recognition. Provenance information, pragmatic function, speaker intention, and locally embedded meanings should be considered when dialect data are prepared, annotated, and evaluated. Normalization can remain useful for particular tasks, but it should coexist with representations that retain the original dialect form and its communicative function.
4.3. Educational Inclusion and Sustainable Learning Environments
The interview themes point to possible educational consequences of standard-centred language technologies. Automated correction, assessment, or tutoring systems may classify dialect features as errors when they are unable to distinguish legitimate variation from departures from a curricular standard. In such settings, learners may encounter a system limitation as an apparently authoritative judgment about their own speech.
The present study identifies this possibility as a design concern requiring direct educational investigation. Future research involving students, teachers, classrooms, and deployed applications would be needed to determine how automated language judgments affect participation, confidence, language awareness, and learning. This distinction allows the linguistic findings to inform educational inquiry without presenting prospective implications as measured classroom outcomes.
A variation-aware approach offers a productive direction for such research. Rather than replacing a dialect form automatically, an educational system could recognize the form, explain its relationship to SWT, and distinguish between contextual register choice and linguistic deficiency. This approach could support bidialectal awareness by helping learners understand how dialect and standard forms function across different communicative settings.
For example, a hypothetical variation-aware tutor encountering Gelecen mi? could provide the following feedback: “This is a recognized Turkish dialect form. In SWT, the same question is written Gelecek misin? Use the SWT form in a formal written assignment; the dialect form remains legitimate in its regional spoken context.” This feedback would teach register-sensitive standard usage without presenting the learner’s dialect as an error.
These considerations are relevant to SDG 4 and SDG 10 because inclusive digital education partly depends on whether language-sensitive technologies can accommodate legitimate linguistic diversity [
1,
19]. The connection lies in the study’s identification of a representational condition that may influence access and participation. Its educational consequences should be examined through future classroom-based and user-centred research.
4.4. An Integrated Model of Digital Linguistic Sustainability
The combined findings support an integrated model comprising three interdependent dimensions:
- (1)
Technical robustness: Language technologies should be evaluated on geographically and socially diverse data, with corpus composition, linguistic tasks, and evaluation conditions documented transparently.
- (2)
Cultural continuity: Data preparation and model evaluation should preserve contextual, pragmatic, and culturally embedded meanings alongside any task-specific normalization.
- (3)
Educational inclusion: Language technologies used in learning environments should distinguish legitimate linguistic variation from error and support equitable participation across different speaker groups.
These dimensions are mutually reinforcing. Technical recognition has limited cultural value if pragmatic meaning is lost; culturally rich resources have limited digital influence if contemporary systems cannot process them; and educational accessibility remains incomplete when only one form of Turkish is treated as linguistically legitimate.
Participation functions as a cross-cutting principle within this model. Dialect speakers, linguists, educators, and relevant user communities can contribute to corpus construction, annotation decisions, evaluation criteria, and assessments of appropriate technological use. Such participation helps connect technical development with the linguistic knowledge and communicative priorities of the communities represented.
4.5. Contribution to Sustainable Development
This study contributes to sustainability research by positioning linguistic diversity as part of the cultural and educational infrastructure of digital transformation. It connects task-specific processing observations with questions of contextual meaning, speaker interpretation, and equitable participation. In doing so, it brings a Turkish case into wider discussions of low-resource variation, inclusive AI, and culturally sustainable language technology.
The study also identifies a conceptual connection between SDG 4 and SDG 10. Language technologies may support more inclusive educational participation when they recognize legitimate variation and communicate the limits of automated analysis. Conversely, standard-centred systems may reproduce existing linguistic hierarchies when their outputs are treated as neutral or universally applicable.
Dialect recognition represents one enabling condition within this broader process. Sustainable development operates at institutional and societal levels and therefore depends on educational policy, technological governance, access, teacher practices, and community participation as well as model design. The present findings identify linguistic representation as one component of this wider structure and provide directions for its investigation in educational and social contexts.
4.6. Practical Implications
The findings suggest several priorities for Turkish language technologies. Model documentation should include dialect-sensitive evaluation materials and linguistically informed error analysis covering morphology, syntax, lexis, semantics, and pragmatics. Aggregate scores alone provide limited information when recurring errors affect culturally or communicatively significant forms.
Before automated feedback, assessment, or tutoring tools are introduced into educational settings, their outputs should be examined with speakers from different geographical and social backgrounds. Interfaces can communicate uncertainty, allow users and teachers to contest or correct outputs, and distinguish dialect forms from errors relative to a specific instructional target. Teacher guidance can further clarify the relationship between dialect variation, SWT, and register choice.
Community-informed resources are also important for long-term development. Dialect speakers, linguists, and educators can contribute to decisions about data selection, transcription, annotation, evaluation, and acceptable use. Openness should be balanced with informed consent, privacy, provenance, and community authority over linguistic data. The objective is not to construct a single fixed technological representation of each dialect, but to develop adaptable systems capable of recognizing variation and preserving contextual meaning.
4.7. Limitations and Directions for Further Research
The geographical coverage of 14 provinces provides a broad exploratory view of Turkish dialect diversity across the country’s seven geographical regions. The participant profile, comprising speakers aged 50–65 in the speech corpus and dialect speakers aged 60–65 in the interviews, was selected to foreground sustained experience with local speech practices. Consequently, the sample does not capture the perspectives of younger speakers, urban speakers whose repertoires reflect mobility and dialect contact, or migrant multilingual speakers. Further research involving these groups and additional provinces will complement the present profile by examining generational change, urban and migration-related variation, multilingual repertoires, and a wider geographical range.
The linguistic analysis focuses on recurrent processing patterns across the collected dialect material rather than constructing separate dialect profiles for individual provinces. This analytical orientation supports the study’s principal aim of identifying forms that may present difficulties for contemporary language technologies. Province-specific features, colloquial usage, multilingual contact, speaker-level variation, and genre-related differences offer productive dimensions for subsequent corpus-based research.
SWT material functions as an interpretive reference for examining how dialect forms are processed in relation to established orthographic and morphosyntactic conventions. A future study based on matched spoken SWT and dialect corpora could extend this comparison by controlling modality, register, disfluency, age, and genre more systematically.
The computational component is organized around the linguistic interpretation of task-specific outputs. The available records support the reported POS observations, dependency-related examples, error patterns, and exploratory contextual-similarity analysis. A dedicated computational benchmark could build on these findings through versioned pipelines, fixed model checkpoints, denominator-level reporting, UAS/LAS measures, confidence intervals, and openly documented analysis code. Such an extension would enable controlled replication while addressing a different methodological objective from the linguistically grounded inquiry pursued here.
The qualitative component enables a comparative interpretation of linguists’ and dialect speakers’ perspectives on linguistic recognition, trust, cultural visibility, and educational inclusion. Broader survey-based or longitudinal studies could extend these findings by examining how such perspectives vary across different demographic, generational, and geographical groups.
Together, these directions would extend the present linguistic framework through generational comparison, province-level analysis, controlled computational evaluation, and direct research in educational settings. Where ethical consent and data-protection requirements permit, the de-identified materials specified in the Data Availability Statement may also support such follow-up research.
The corpus developed for the present study also provides a foundation for a planned follow-up investigation with a more narrowly defined computational focus. That study will extend the existing material with a modality-matched spoken SWT reference, province-level linguistic comparisons, version-controlled processing pipelines, task-specific metrics, and more detailed regional analyses. These components require a separate research design and analytical framework beyond the linguistically grounded and interdisciplinary scope of the present article. The planned investigation will therefore build on the current findings by examining the identified processing patterns under more controlled and computationally reproducible conditions.
5. Conclusions
This study approaches the processing of Turkish dialects by AI-based language technologies through the framework of digital linguistic sustainability. The task-specific analysis reveals recurrent processing difficulties involving dialect morphology, flexible word order, lexical items, discourse particles, and pragmatically marked expressions. The interview findings complement these linguistic observations by relating digital recognition to cultural visibility, trust, and prospective educational use. Their interpretive integration brings together the technical, cultural, and educational dimensions of dialect representation.
Sustainable Turkish language technology requires more than expanding the volume of dialect data. Transparent corpus documentation, linguistically informed evaluation, preservation of contextual meaning, speaker participation, and a clear distinction between legitimate variation and processing error are equally important. In educational settings, these principles can guide the development of systems that recognize dialect forms, communicate uncertainty, and support awareness of the relationship between dialects and SWT. Research involving students, teachers, and deployed applications will be essential for evaluating these educational implications and examining their broader relevance to SDG 4 and SDG 10.
Turkish dialects constitute living repositories of linguistic knowledge, social memory, and cultural experience. By connecting technical robustness, cultural continuity, and educational inclusion, the study offers a linguistically grounded framework for incorporating this diversity into sustainable AI development. Its principal contribution lies in demonstrating that the digital future of Turkish depends not only on the technological processing of language but also on the capacity of language technologies to recognize and preserve the variation through which speakers express identity, belonging, and local knowledge.