Next Article in Journal
Intergenerational Trauma and Resilience in African American Families: A Dimensional Conceptual Analysis of Dyads and Triads
Previous Article in Journal
“Framed as a Criminal, Rather than as Artist”: A Narrative Study into Meaning-Making by UK Drill Artists
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Enhanced Integration of Multi-Disciplinary Inputs into a Narrative of an Ancient Migration, Based on Greater Chronological Precision Provided by a Novel Y-DNA Clock and Phylogenetic Branching

by
Desmond D. Mascarenhas
1,*,
Balaji Rajagapolan
2,
John W. Fox
3 and
Richard J. Johnson
4
1
Department of Research, Mayflower Organization for Research & Education, Sunnyvale, CA 94085, USA
2
Department of Civil, Environmental and Architectural Engineering, and CIRES, University of Colorado, Boulder, CO 80309, USA
3
Department of International Studies (Retired), American University of Sharjah, Sharjah P.O. Box 26666, United Arab Emirates
4
Department of Medicine, University of Colorado Anschutz Medical Campus, Aurora, CO 80045, USA
*
Author to whom correspondence should be addressed.
Genealogy 2026, 10(1), 14; https://doi.org/10.3390/genealogy10010014
Submission received: 18 November 2025 / Revised: 9 January 2026 / Accepted: 10 January 2026 / Published: 14 January 2026

Abstract

An accurate DNA clock can strengthen cross-disciplinary inputs in the study of genealogies and ancient migrations. New Y-chromosome sequence data gathered from a Lotli Pai Kaundinya (LPK) Brahmin cohort whose staged migration from the Pontic Steppe to the West Coast of India was previously reported, are used here to generate a more precise DNA clock. The formula distinguishes Y-mutation rates for transitions and transversions and corrects for dropped mutations in sequence reads. The formula is validated against a baptismal tree covering over four centuries (0–704 YBP interval), a published STR-based chronology for this same cohort (704–5200 YBP) and a comparison to Y-Full formation times for mutations older than 3000 YBP. Using this more precise clock, we support a proposed “founder effect” expansion in Khorasan during 4300–3800 YBP using a novel phylogenetic branching metric; and use archeological, numismatic, toponymic, climate reconstruction and ancient textual data to explore religious and professional dimensions of cultural kinship with other communities believed to have interacted with the LPK during their long migration. The availability of more precise dating facilitates the integration of such secondary data types, resulting in an enriched and more plausible migration narrative.

1. Introduction

Narratives about ancient migrations of people, language, and material culture often lack chronological specificity. The conjunction of archeological, linguistic, and genetic data types, for example, can be better used to understand the ancient world if they are securely matched in time. Narratives about ancient migration could thus be improved by techniques for more precise genetic dating. This project has two goals: the first is to build a more refined DNA clock, based on newly sequenced genomes from a well-characterized patrilineage whose migration from the Pontic Steppe to South Asia has been previously described (Mascarenhas et al. 2015). The second is to use the resulting chronological specificity to integrate a multiplicity of data types, thereby enriching and providing plausibility to the account.

1.1. Building a Better DNA Clock

The use of a particular patrilineage to study the general subject of ancient migrations clearly includes the limitation that one instance of a migration does not necessarily inform another. But it also obviates some drawbacks of other genetic approaches. Population-level studies are often plagued by sampling bias. Genetic data derived from ancient burial sites are a good case in point. Sampling bias is mitigated in large public databases such as Y-Full and FT-DNA (www.familytreedna.com) where any donor from any geographic location can contribute a sequence or genotype, but the volume of DNA files submitted to repositories still varies significantly from one country to the next. The recent proliferation of statistical analyses of autosomal DNA datasets gathered from vaguely defined geographic domains such as ‘Steppe’ and ‘Caucasus Hunter-Gatherer’ (CHG) provide conclusions that often temper enthusiasm. A modern human may learn that she is five or twelve percent CHG from an aggregate of migration events over millennia but may also wish to know what religious beliefs her own direct ancestors held three thousand years ago, how those beliefs differed from those of immediate neighbors, and the eventual consequences of that dissonance—such as being obliged to pick up sticks and migrate again.
In addressing the first goal, we sequenced multiple genomes of individuals from a previously described patrilineage of Lotli Pai Kaundinya (LPK) Brahmins (Mascarenhas et al. 2015). A primary goal of our study is to better understand this multi-stage LPK patrilineal migration, whose history covers a punctuated journey of about 2000 miles from the Pontic Steppe to the west coast of India over a span of about 5000 years. To this end, we develop a more accurate DNA clock than the one currently used to compute the age of mutations. A prime obstacle to using genomic sequences for that purpose is the phenomenon of “dropped SNPs” (single nucleotide polymorphisms) in sequence reads. Formation ages computed from low-quality reads in database deposits tend to result in large standard deviations of 25% or even more, thus limiting their usefulness in the precise dating of ancient events. Dropped SNPs tend to be ubiquitous in deposited Y-DNA sequence files, for several reasons: the human Y template is particularly challenging to sequence, the number of runs per sample (a major cost factor) varies between samples submitted to a database, technical precision can be influenced by the type of equipment used to generate sequence data, and variation in operator skill can influence sample preparation. One approach to “recapturing” missing SNPs is the post facto comparison of multiple kin genome sequences carrying an identical string of mutations above a certain node in the DNA tree. Provided the phylogenetic branches have been accurately mapped, a multiplicity of corroborative kin samples can generate confidence that all mutations within a given Y-DNA lineage over a specified chronological period have, in fact, been captured (Le and Durbin 2011). Conversely, because fewer kin genomes are publicly available for those mutations that have been formed, say, less than 3000 YBP, confidence decreases substantially as one moves out along any particular DNA branch towards the tree’s more recent mutations. To reach the same confidence level for recent SNPs would require extrapolation from genome sequences of multiple close relatives in order to ensure that no SNPs are missed after curation, and that is what we try to accomplish in this study. We find it a good rule of thumb to have at least 5–10 replicates of closely related genome sequences to catalog the mutations acquired until very modern times.
Once a complete and phylogenetically accurate mutation inventory has been assembled, a long-standing inconsistency in the computation of mutation age must be addressed. Transitions in the Y-chromosome occur about twice as frequently as transversions (Bonito et al. 2021) but this distinction has not always been brought into formulaic computations of mutation age. There other uncertainties: we do not know, for instance, how a period of rapid “founder” gene pool expansion might affect the accuracy of a DNA clock within a given branch during a period of accelerated phylogenetic branching. In this work, we use several innovations to try and address these obstacles.

1.2. Using Chronological Specificity to Integrate Other Data Types

The specific migration narrative for this patrilineage over a span of millenia is a rare example of the genetic “particular” potentially informing an otherwise vague “general” narrative about the southern migration of Indo-Iranic-language-speaking peoples from their original Yamnaya koine. A migration “event” recorded by history or by modern DNA footprints can be understood as an aggregate of hundreds, even thousands of instantiations of individuals (or families) moving from point A to point B, B to C, and so on, over multiple generations (mini-migrations). The choreography of this aggregate phenomenon presumably involves path-finding at the outset, which might involve ‘foraging’ events, wherein individuals with particular occupational skills venture into unfamiliar territory with the intention of laying down roots in suitably unexploited, favorable habitats—where suitability refers to the migrant’s professional work. A common example of this path-finding is to be found in the travel routes and contacts set up by traders. Since ancient caravans would have needed to underwrite costs in both directions one might, in that scenario, assume a bidirectional symmetry of commercial opportunities between A and B, for example.
Secondarily, once such a trail is defined, the enduring causes for net-directional journeys must be identified. A migrant might require an impetus: at a minimum, an asymmetrical professional opportunity (prospects at B better than A; prospects at C better than B) for the migrant individual or family. This might then be paired with changed circumstances such as progressive drought or religious persecution at the originating location. Or opportunities arising from earlier migration of kin, or a positive narrative about the destination. Thus, the migration is not of an individual or family alone, but of many conspiring variables. A migration narrative can be enriched by a description of such variables.
Reconstruction of any specific ancient journey is probably impossible, in the absence of a specific record providing an account of such an event. But outlining the general impetus and opportunity for repeated mini-migrations of Indo-Iranic language speakers of a given profession over a span of several centuries may be more tractable if the academic consensus over the location of an Indo-European language (IE) koine in the Yamnaya archeological horizon in the Pontic Steppe is well-founded (thus defining point “A”); and, furthermore, the linguistic and geographic coordinates of a well-documented professional ‘origin story’ for Saraswat Brahmins in ~3500 YBP exist. If deemed sufficiently plausible, these may anchor conjectures about migrations that may have preceded and followed that ancient pivot point (“B”, “C”).
Thinking about migrations in this holistic fashion vastly complicates the overall descriptive task. One must consider dimensions introduced by coexisting professions and cultural groups, and their objectives. One must consider the relationships between the migrant(s) and the community in which they were embedded. Thus, the analysis of a single instance of a migrating Y-haplogroup patrilineage, aided by improved chronological precision can, in theory, reduce the complexity of such interactions.
Here, we propose the co-migration of a distinctive genetic marker with a distinctive language group in the early stages of a migration saga (for which precise records are unavailable) and integrate that with data from subsequent millennia, for which more precise records are available. In the case of the LPK, ancient (from ~3500 YBP) cultural and textual histories are accompanied by a strict religious endogamy, at once allowing for secure tracking of a clearly defined Y-haplogroup subclade (R-Y7) over millennia, and also cross-indexing of that account with cultural beliefs, oral histories, symbols, sacerdotal texts and long-standing marriage customs. These, in turn, make the analysis of religious symbols in ancient coinage, toponyms, ethnonyms and numerous other data types more relevant. Our goal is to build a comprehensive, multi-dimensional migration narrative to show—to re-phrase an old archeological dictum—that these LPK people are (plausibly) not ‘pots’. Plausibility can in fact be strengthened if multiple data types point to the same interpretation within an accurate chronological frame.

2. Results

2.1. Geographic Distribution of Y-Haplogroups (FT-DNA Database)

Y-haplogroup R-Z93 has been suggested as a useful marker for patrilineal tracking of Indo-European-language (IE) speaking people (Pamjav et al. 2012). R-Z94 is a subclade of R-Z93 that serves as a particularly useful marker for the Indo-Iranic branch (Mascarenhas et al. 2015). Table 1 summarizes the relative geographic distribution of 2519 FT-DNA database submissions for three Y-haplogroups whose geographic provenance is distinct: R-Z283, R-Z94, and C-M216. R-Z283 is well represented in Western European areas whereas the other two Y-haplogroups are not, with one interesting exception: An R-Z94 subclade Y2619 is almost exclusively found in the Slavic countries Ukraine and Belarus. In the IE tree, some linguists place Balto-Slav as the most recent language group to split from Indo-Iranian, these constituting the two major ‘satum’ branches of IE (Kortlandt 2016). Table 1 shows that the provenance of R-Z94 tracks well with the distribution of Indo-Iranian languages, whereas C-M216 is widespread in the Andronovo region and East Asia, but much less so in the Indo-Iranian-language-speaking areas. As shown in Table 1B, the R-Y3 subclade of Z94 is particularly well-represented in Khorasan, South Asia and the Gulf States, where the R-Y7 branch of R-Y3 accounts for 80–90% of database samples.
The modern provenance of R-Z94 supports a geographical origin in the Pontic Steppe. A large public human genome database, Y-Full, provides additional perspective on R-Z94 Y-haplogroup distribution. As shown in Figure 1A, the Y2619 subclade of Z94 is almost exclusively found in speakers of Balto-Slav languages. It is absent south of the Caucasus, as is FGC56408, another branch of Z94 with distribution confined to the Andronovo and East Asian region. Though not shown in this figure, another significant branch of Z94, R-YP451, is found almost exclusively in the Circassian region of the northern Caucasus. All these branches radiate from the Yamnaya archeological horizon in the Pontic Steppe which has, by scholarly consensus, long been considered the original koine of IE-language speakers. Thus, R-Z94 may be a particularly informative mutation for studying aspects of the Yamnaya diaspora as it relates to the two major satem branches, Balto-Slav and Indo-Iranian. That diaspora is believed to have split these two language groups geographically between 5500 and 4300 YBP, the latter marking an ante quem for the archeologically defined phase of southward ‘kurgan’ people introgressions through the Caucasus. Four different approaches to the computation of the age of the split between Indo-Iranic from Balto-Slav have yielded estimates centered between 5500 and 4500 YBP, which could thus be considered the post quem and ante quem for departure of Indo-Iranian-speaking people from the Pontic Steppe (Blazec 2007; Nakhleh et al. 2005; Serva and Petroni 2008; Heggarty et al. 2023).

2.2. Computation of R-Z94 Subclade Branching Ages Illustrates the Problem of Dropped SNPs

Figure 1B illustrates how the problem of dropped SNPs is directly related to a lack of chronological specificity in the computation of mutation age. The large standards deviations in Y-Full’s computation of age are thought to be proportional to the incomplete reporting of mutations in the genome sequences, which are deposited in the database. Curation of these defects becomes harder as the number of replicates decreases (Le and Durbin 2011). At the present time, mutations whose formation is more recent than ~3000 YBP tend to suffer from a lack of sufficient replicates in the database which could be used to compare duplicates for missed SNPs with confidence. For example, as of this writing, only six genomes in the public Y-Full database carry R-Y16494, a mutation that formed about 3000 years ago.

2.3. Building and Validating a Linear Clock Based on Mutations Carried by the LPK Lineage; And Calculation of the Phylogenetic Branching Rate

The genomes of seven men in the clan (vangor) of the LPK cohort previously studied by STR analysis (Mascarenhas et al. 2015) were fully sequenced. Sequences of 40 other genomes, that were kindly provided by the Y-Full database (via YSEQ laboratory; YSEQ GmbH, Berlin, Germany), served as controls for this study. By use of these sequences as comparators, a complete set of SNPs carried by the LPK R-Y7 lineage was compiled, with expected dates of formation calculated from the present (nominally, 2000 CE) to about 5000 YBP. As detailed more fully in the Methods Section, a complete baptismal database going back to 1613 was created by accessing and translating several thousand handwritten parish records for Lotli town; these records were collated from churches and archives. In this manner, a complete family tree was constructed, and the exact times to common ancestors calculated for four LPK individuals going back to ~1583, when conversion of an ancestor to Christianity is believed to have taken place (Figure 1C). The branching dates for three other LPK donors came from oral and historical records of the clan’s most proximate geographical migration from Gujarat (740 CE; one individual) and the Madhva conversion to Vaishnava Hinduism in 1296 CE (two individuals). The logic for these assignments is discussed more fully in the previous study (Mascarenhas et al. 2015). Calculation of mutation rate intervals was performed by data fitting of mutation inventories to historical markers over a 704-year span. Since multiple pairwise comparisons could be made between the mutations carried by each of the seven LPK individuals, estimates of mutation rates (transitions and transversions) were derived with 12 intervals as inputs. The final plot of the cohort of LPK mutations for historical events going back to 704 YBP was linear (r2 = 0.995), as shown in Figure 2a, top panel. Transversion intervals calculated from this plot were 149.8 years and half that value (74.9 years) was assigned to transitions. Using those values, and calculating from actual transitions and transversions in the genome sequences, the values shown in Table 2 going back to R-Z94 were generated. These values are remarkably close to Y-Full estimates for older mutations, despite the significant differences in methodology. Validations were performed as follows. Longer time spans (>3000 YBP) were validated against the computed formation times provided by Y-Full for lineage mutations between R-Z94 and R-Y16494. That plot is linear with an r2 = 0.950 (Figure 2a, bottom panel). In addition, the span between 704 and 5200 YBP was validated against values generated from STR data computed from LPK donors in the previous study (Mascarenhas et al. 2015), and populations in the FT-DNA database, but using the improved methodology described in Methods, wherein the same 50-STR set was used in all comparisons. This plot was also linear (Figure 2a, middle panel) with an r2 = 0.976.

2.4. Pre-Vedic, Out-of-Yamanaya Migration of the LPK Lineage to Rig-i-Stan (~5000–3500 YBP)

This section summarizes the case laid out at length in Supplementary Text File ST1. Figure 3 traces the proposed migration of pre-Vedic LPK ancestors from the Yamnaya Indo-Iranic language koine to the Harakhuti Valley in Rig-i-Stan via a number of stages—notably a Y-haplogroup expansion stage in Khorasan, and following the west Caspian Transcaucasian and northern Persian journey route for co-traveling Y-haplogroups described in an earlier study (Mascarenhas et al. 2015).
As mentioned above, an ante-quem/post-quem for the date of departure of R-Y3 (an R-Z94 branch) from the Yamnaya between 5500 and 4500 YBP is supported by four different studies dating the linguistic split between Indo-Iranian and Balto-Slav (Blazec 2007; Nakhleh et al. 2005; Serva and Petroni 2008; Heggarty et al. 2023). The formation of the R-Z94 mutation itself has been dated by three different methods, cited in this study, to 5431, 5093 and 4885 YBP, providing another post quem date for the pre-Vedic LPK diaspora. In a complementary manner, the Transcaucasian segment of the proposed pre-Vedic LPK journey out of the Pontic step is informed by the appearance of introgressive kurgan cultures in Georgia and Dagestan between 4700 and 4300 YBP (Ghalichi et al. 2024; Kohl 1988; Moradi et al. 2023). Most calibrated radiocarbon dating of Velikent material falls in the 5200–4400 YBP range, consistent with this proposal (Kohl and Magomedov 2014). Based on the preponderance of bovine remains at Velikent (roughly ¾ bovid, ¼ caprid; no equids), and the marks left on cattle bones suggesting oxen-drawn carts, the Indo-Iranic-speaking communities associated with the pre-Vedic LPK may have been cattle drivers with a dairy product economic base, as was characteristic for the Yamnaya people (Kohl and Magomedov 2014; Maurer and Greenberg 2022). Ancient migrants generally sought environments that helped them maintain traditional lifestyles and professions. We therefore hypothesize that during their migration between Yamnaya and Rig-i-Stan (between ca. 4500–3500 YBP) the pre-Vedic LPK sought out habitats characterized by their preferred topographies: banks of rivers situated in intermontane valleys along the edge of arid plains (8–16 inches rainfall per annum). For cattle drivers, this points to transhumant practices with a possible sedentary component, or interaction. The intermontane valleys of Dagestan, parts of the Atrek River Valley, Khorasan-Razavi, Herat Valley and Harkhuti Valley in Rig-i-Stan fit this profile exactly (Figure 3).
In the mid-third millennium BCE, the intermontane valleys of Khorasan-Razavi and, presumably, adjacent regions of Herat, were unpopulated or lightly populated (Biscione and Vahdati 2020; Fouache et al. 2010). These regions were easily accessible from the Caspian littoral via the Atrek Valley or, alternatively, the ancient trade route going through Damghan. If founder effects were to follow, it reflects the choice of early R-Y3 settlers to settle in sparsely populated areas of Khorasan. It is clear that many such pockets existed in Khorasan-Razavi (Biscione and Vahdati 2020; Fouache et al. 2010). There was, importantly, regular traffic across these domains, as ancient caravan routes crisscrossed these regions. The possibility of a significant wine-making industry suggested by archeobotanical remains at Tepe Chalow showing an emphasis on grapes as crops (Vahdati et al. 2019) could, in turn, suggest a thriving hospitality business suited to the kind of pass-through activity most beneficial to the expansion of founder gene pools in thinly populated areas. Adjacent regions of Afghanistan in Herat Province (an intersection traversed by multiple major ancient trade routes) may have offered similar opportunities. These factors relate to the proposed expansion described in the next section. Based on all the above, we therefore propose that, prior to 4300 YBP, the patrilineal ancestors of the pre-Vedic R-Y7 LPK may have settled in the Khorasan geographical domain that includes Khorasan-Razavi and the Herat Valley.

2.5. Proposed R-Y7 Gene Pool Expansion Coinciding with BMAC/GKC (4300–3800)

Figure 4a summarizes the reasoning behind the proposed “sandwich effect” that may have led to a significant expansion of a “founder” R-Y7 LPK Y-haplogroup in a region of Khorasan near present-day Herat (Afghanistan) roughly between 4300 and 3800 YBP. A host of conspiring factors appears to explain the prosperity of the BMAC/GKC florescence during this time: the depopulation of major settlements in southern and southeastern parts of the Persian geographic basin (Tal-i Malyan, Jiroft, Tepe Yahya, Shar-i-Sukhta, Bampur, Mundigak) triggered by a 4.2 K aridity event known to have collapsed trade routes to Mesopotamian markets along the Persian Gulf that coincided with rapid population changes in the Lut/Shahdad region between 4500 and 3500 YBP (Figure 2b, bottom panel, and (Eskandaridamne 2021)) may have signaled a corresponding rise in the opportunity for the export of metals, precious stones and other goods through the northern Persian route less-affected by the drought. Together, these forces may have conspired to cause a net population displacement towards the BMAC/GKC.
As shown in Figure 4a, this population flux would almost certainly have been squeezed into a geographic and cultural corridor defined by desert and mountain barriers on the one hand, and a cultural kinship between population centers in southeastern Persia, Balochistan, and the Kopet Dagh—the Turan cultural corridor (see Supplementary Material Text File S1). One distinctive archeological artifact tying these regions together is the “stepped-cross” compartmented stamp seal described by (Salvatori 2000) whose limited find spots are shown in Figure 4a. This motif existed in this geographic domain at sites like Mundigak, well before Harappan times. Population movements caused by aridity or other factors may have resulted in population movements up and down this corridor. As shown in Figure 2b (top panel), a phylogenetic branching analysis shows the rate of branching over time in the R-Y7 and R-Y6 sister subclades originating from L657. This measure can serve as a proxy for the rate of reproduction in the carriers of these Y-haplogroups. The average phylogenetic branching rate exploded significantly during the BMAC/GKC period, ~5X above baseline for both sister clades, with the rate peaks for both being coincident around 3800 YBP, i.e., around the end of the BMAC/GKC expansion. Since the R-Y7 and R-Y6 lineage mutations and branch counts are genetically independent events, they serve as internal controls for each other. Significantly, the average mutation formation rate in the R-Y7 and R-Y6 lineages during this period is 124.0 and 119.4 years—compared to 129.5 years for R-Y7 computed by an independent procedure using STR-TMRCA for the interval R-Z94 to R-Y16494, covering over two millenia, and bracketing this era. In other words, the rate of mutation along the R-Y7 lineage does not appear to accelerate to anywhere near the degree seen for phylogenetic branching during the BMAC/GKC era. Phylogenetic branching, which suggests new reproduction, not just a population increase from an influx of old Y-haplogroups from the same lineage, is a potential marker for founder gene pool expansion in a lineage over a short time period. Such characteristics might reflect a slow and sustained multi-century introgression of new families whose daughters help grow the founder Y-haplogroup pools of local founder families, but whose own Y-haplogroups, at the same rate of reproduction would be far more likely to be lost by genetic drift, even with symmetrical reproduction rates.

2.6. Possible Linguistic Split of Iranic and Indic Dialects During BMAC/GKC Expansion

Linguistic studies calculate the split between Iranic and Indic languages between 4300 and 3800 YBP, a period exactly coinciding with the BMAC/GKC population expansion (Blazec 2007; Heggarty et al. 2023; Nakhleh et al. 2005; Serva and Petroni 2008). ‘Aryanization’ in the Vedic age (after 3700 YBP) caused by the dissemination of military and ideological technologies from IE-speaking kin in the Sintashta-Arkhaim Complex, must have occurred in parallel for the warrior and priestly sections of the two language groups, including the ancestors of the LPK in Rig-i-stan. Archeobotanical crop expansion data (Stevens et al. 2016) suggest a much amplified bilateral interaction between the BMAC/GKC on the one hand, and both the Andronovo region and Xinjiang province (Xia and Shang Dynasties; Figure 5B(a)) on the other, by the second quarter of the second millennium BCE, but definitely not prior to the BMAC/GKC period. The important cultural and technological diffusion period (3700–3200 YBP; Figure 4b) that follows the BMAC/GKC period is significant for many reasons, and is discussed more fully in Supplementary Material Text File ST1. It is worth noting that Indic and Iranic tribes would compile holy texts in Sanskrit (and later, Avestan) by as early as 3500 YBP, or even earlier, according to some scholars. As argued in Supplementary Material Text File ST1, the linguistic speciation of Indic from Iranic languages could not have taken place much later than the BMAC/GKC expansion period shown in Figure 4b, when population movement through the Turan cultural corridor was heightened by climate change, and other factors.

2.7. Styles, Cultures and Technologies: The 3700–3200 YBP Dissemination Period

Studies of pollen at Lake Almalou, east of Lake Urmia, show 3700 YBP as an early example of several notable peaks in the climate history of northwestern Iran (Djamali et al. 2009). As a proxy for agricultural prosperity and population growth, Almalou pollen peaks also coincided with, for example, the later rise in the Achaemenids and Sassanids. As shown in Figure 4b, the 3700–3200 YBP period is marked by the southward diffusion of the Sintashta-Arkhaim technocultural complex into the Persian Plateau; the diffusion of Xia and Shang dynasty pottery styles from China to NW Iran destinations such as Marlik and Luristan (Figure 5B(a)); the introgression of new crop species between Iran and China in both directions (Stevens et al. 2016); the dissemination of distinctive NW Iran pottery styles and crop species such as flaxseed and lentils to Ahar-Banas and other western Indian destinations through Gujarat (Sankalia 1963; Mascarenhas et al. 2015); and the introgression of new cultural elements that may have involved pre-Vedic Indo-Iranic language-speaking peoples entering the Indo-Gangetic basin from the northwest in the second millennium BCE—archeologically, the Gandhara Grave Culture, Cemetery H complex, Ochre-Colored-Pottery Culture, Copper Hoard horizons. It is likely that, just prior to this span of time (but no later than 3700 YBP), the LPK clan made its relatively short move from wherever it lived in the Khorasan region to the banks of the Harakhuti in Rig-i-Stan. A strong interconnection between southern Central Asia and SE Iran/W Pakistan/SW Afghanistan in the early second millennium BCE was marked by dissemination of highly distinctive BMAC/GKC prestige artifacts, such as miniature columns and stone disks, across the Turan domain—from Chalow and Bojnard in the northeast to Seistan in the southeast, and as far east as Mundigak in Rig-i-Stan and Quetta (Moradi et al. 2023; Biscione and Vahdati 2020). These dissemination tracks and timing, which followed much older disseminations of (Biscione and Vahdati 2020) distinctive pottery from Namazga to Quetta in the early third millennium, allow us to date the LPK’s move to Rig-i-Stan to the early second millennium. We estimate that Brahmanization of the LPK initially occurred in Rig-i-Stan between ca. 3700–3400 YBP. It is important to note that the Saraswat Brahmin LPK lineage may not have continued into the Indian sub-continent on their next stage of migration until the Painted Grey Ware period (3200–2800 YBP), an era commonly associated with the arrival of the Rigvedic people into the Punjab and Madhyadesa. This would be its first foray outside the R-Y7 core geographic area. By our DNA clock, the separation of LPK R-Y7 from the R-Y7 modal pool can be calculated to between 3146 and 2996 YBP, and by STR analysis (a completely different methodology) to 2971 YBP. Although the introgressions of other Indo-Iranic language-speaking peoples into the subcontinent from the Persian basin may have begun almost a millennium earlier than the PGW, these were almost certainly non-Vedic, non-Aryanized people. Nevertheless, those diffusions (Figure 5B(c)) may have provided an important cultural kinship trail to facilitate the LPK’s later journey. In particular, the Bhoja (Supplementary Material Text File ST2) and Kamboja (Supplementary Material Text File ST3) communities may have served this purpose after the introgression of the now-Brahmanized LPK’s into the Indian subcontinent. Even more generally, an unmistakable marker of early migration of Iranic tribes is found in the ubiquity of Proto-Elamite religious symbols stamped on the earliest coins of the subcontinent’s northwestern tribes between ~2500 and 1800 YBP (Figure 5B(b)). One must assume, therefore, that at least in the western belt of the subcontinent, a variety of cultural elements might seem familiar to in-coming tribes with Persian affinities, especially under the Saka Western Satrap dynasties. Figure 5B(c) summarizes some of the relevant introgressions of co-traveler communities into the Indian subcontinent. The earliest of these almost certainly led to founder effects (visible in toponyms today) as the estimated populations of the vast Indian subcontinent ranged from only 1–2 million in Harappan times to about 30 million by the start of the Common Era (Dyson 2018; Encyclopedia Britannica—for Harappan Population Estimate: U. of Groningen Angus Madison Project—for population of individual Asian countries from 1000–180 YBP; U. Texas at Austin—for habitable world land surface area); US Census Bureau—for world population estimates;). Many toponyms, theonyms, and ethnonyms from this period of northwestern influx are memorialized in the resulting founder settlements, and endure to this day.

2.8. Introgression into the Subcontinent of Communities Culturally and Professionally Associated with the Brahmanical LPK

The panels of Figure 5B(c) summarize introgressions relevant to LPK migration post Rig-i-Stan (left to right): (i) LPK Y-haplogroups located in Khorasan during Harappan times (>4400 YBP); (ii) introgression of Persian-influenced pre-Vedic IE-language speakers (possibly Yadava Bhoja) linked to the WP-BRW material culture (3700–3000 YBP) via an entry point through Gujarat; (iii) further movement into the Indo-Gangetic region where Bhojas may have co-located with Vedic LPK following the arrival of the latter into the Punjab (as part of the PGW culture) between 3200 and 2800 YBP; (iv) these communities were thereafter influenced by sequential northwestern tribal incursions of decreasing impact (Saka, Kushan, Huna); together, the net effect of these incursions is seen in the movement of a subset of northwest tribes associated with Proto-Elamite numismatic symbols such as the six-arched hill, or Yadava Bhoja lion-chakra, i.e., Vrishni symbols (Figure 5B(b)) southward towards Rajasthan, Malwa, and Gujarat. These western regions were also the places where Saka influence was greatest during the early Common Era; (v) the sacking of Vallabhi in Saurashtra by an Arab raid from Sindh in 1260 YBP effectively ended the Maitraka dynasty’s ability to employ LPK scribes, possibly triggering the LPK’s final journey by sea from Gujarat to a location situated downriver, and adjacent to both a pre-existing Bhoja kingdom in Chandor, Goa; and the expansion of a major west coast port, Gopakapattan, which would serve as the major entry point for the horse trade to Deccan kingdoms for many centuries, a core business enterprise of Kamboja syndicates. This final known leg of the LPK’s >2000-mile migration from the Pontic Steppe (Mascarenhas et al. 2015).

2.9. Catalytic Role of Climate

Using contemporary period connections between ENSO and precipitation over the Indian subcontinent and Central Asia, we reconstructed the average annual precipitation changes (in percentage) relative to the modern period average over these regions during four broad epochs covering 5.5 K YBP ~modern, which includes mid and late Holocene. This encompasses the ~4.2 ka mega drought event when migrations were thought to have been exacerbated (Figure 6). Figure 6A shows the percentage change in precipitation during pre-mega drought period (~5.5–4.5 K YBP), based on NINO 3.4 of ~−1.75 °C (NINO 3.4 is an index calculated from a key region in the equatorial Pacific Ocean used to monitor sea surface temperature anomalies) and hemispherical temperature gradient to be 0.4 °C. The gradient in rainfall between Indian subcontinent and central Asia is clear. The drier (~−5 to −20% lower precipitation than modern) regions run from Afghanistan to Turkey and the wetter (~5 to 20%) region spans Pakistan and western India. These gradients are modest but likely enabled a steady low-level southward move of the LPK from the Caspian region (Figure 3) to central Iran. However, note that southeast Iran and northern Arabian Peninsula shows higher rainfall reductions relative to present, thus forming a mobility barrier eastward to the Indus region along with the cultural barrier. Thus, any migration/movement would have to be northward to the BMAC/GKC region (Figure 4). During the period encompassing the mega drought ~4.5–3.5 K YBP in Figure 6B, this precipitation gradient is stronger and sharper. The wetter regions over Pakistan and northwestern India along Arabian Sea (mature phase of Harappan Civilization) show (25~75%) increase in precipitation and over a generally wetter (0~20%) Indian subcontinent. In contrast, precipitation appreciably declined (−25% to −75%) in Western and Central Asia, from Caucuses into Afghanistan. These gradients are consistent with simulations from other proxy records and GCMs (Dallmeyer et al. 2013; Wang et al. 2016). The precipitation reduction over southeastern Iran and northern Arabian Peninsula is even stronger thus strengthening the eastern mobility barrier mentioned above, favoring northward movement of LPK to BMAC/GKC and expansion of the Gene Pool (Figure 4). Given that the Harappan Civilization relied heavily on trade with Mesopotamia and maritime societies along the route for their flourishing, this epoch made things very difficult. In that, while Indian subcontinent was wetter the regions of their trading partners were undergoing severe and sustained dry period, thus, drying up markets for their products. Consequently, and ironically (despite wetter Indian subcontinent) this likely laid the grounds for the decline of Harappan society. The post mega drought period (3.5~2.5 ka BP and recent) (Figure 6C) based on NINO3.4 and temperature gradient of ~−0.5 °C and 0.1 °C, respectively, the precipitation change pattern seen in Figure 6B persisted though in reduced amounts. However, the gradient persisted between Central Asia and India. Furthermore, the rainfall over Pakistan and western India also considerably dropped (~5–25% of modern) relative to earlier periods, this includes southeastern Iran. Also, this increase persisted from western India to the Indo-Gangetic plains (east of Nepal) and over rest of India though with a small increase (5–15%) in precipitation. Furthermore, the reduction in precipitation in the BMAC/GKC region and much of Iran is weaker (~−10% of modern) with Arabian Peninsula and parts of Mesopotamia experiencing wetter (albeit modest) conditions (0~10%) relative to present.

3. Discussion

Reconstructing migrations of ancient peoples is a daunting task. It is not surprising, therefore, that narratives relating to the southward movement of Indo-Iranic language-speaking peoples from the presumed Yamnaya region koine have been mired in confusion and controversy. The main data types used to reconstruct ancient events (archeological, linguistic and genetic) carry enough chronological imprecision to make cross-disciplinary alignment challenging, often making narratives about ancient migrations unreliable. In this study, we aimed to add specificity to such narratives by building a better DNA clock; then using the chronological specificity specified by that clock to better integrate other data types into the migration account.

3.1. A Better DNA Clock and Founder Effects in Khorasan

Based on a “multiple-small-migration-event” model, clines statistically deduced from autosomal genomes may not help us track migrations that are dominantly male. A study of sex-specific inheritance in the late Neolithic/Bronze Age migration from the Pontic-Caspian Steppe, for example, showed a 5-14X male bias in the migrant pool (Goldberg et al. 2017). Carriers of Y-haplogroups would not be reflected in the autosomal DNA signal left by small numbers of male migrants, as that signal would have been quickly lost via genetic drift, thereby making the migration of these men “invisible” to the techniques currently in vogue for characterizing population movements. A recently published study (Lazaridis et al. 2025), for example, provides data on genetic clines associated with the Yamnaya people, but throws little light on the migration of Indo-Iranic language-speaking men traveling southeast through the Caucasus. Furthermore, the use of R-M417 in that study as a marker for expansion of male DNA towards the eastern Steppe is predicated on the assumption that most Indo-Iranic speakers went that way, an incorrect assumption, based on our results. Moreover, unlike the subclades of R-Z94 used in our study, R-M417 does not offer sufficient precision for making any kind of determination for linguistic groups in the first place. Conversely, with the aid of founder effects, Y-haplogroup signals might be fixed at certain locations along the migration route. We use a novel phylogenetic branching analysis to provide evidence for such a phenomenon in the case of R-Y7 in Khorasan.
The new methodology builds a more accurate mutation clock by mitigating the problem of dropped SNPs, and by scoring transitions and transversions differently in genomic sequences of closely related members within a single, patrilineal Y-haplogroup cohort. By virtue of the rich religious and historical traditions of this LPK cohort (Mascarenhas et al. 2015), we then use the new DNA clock to construct a logical and internally consistent 2000-mile LPK migration narrative going back about five thousand years.
In the rare case where the migrant(s) carry a distinctive Y-haplogroup, and the chosen destination is lightly populated, there is an opportunity for that distinctive Y-haplogroup to expand preferentially in that location—through so-called “founder effects”. For this phenomenon to occur, the migrant cohort must reproduce successfully in situ for some generations (to set up a sufficiently sized gene pool) and then many additional travelers must traverse that location at a rate that makes their own genetic contributions small enough to be lost through subsequent genetic drift.
Thus, the two main requirements for founder effects are: (a) sufficient genetic establishment of the founder Y-haplogroups; and (b) sufficient through-traffic to amplify the original genetic signal without causing long-term dilution. Specific data to support such founder effects is missing from earlier narratives for migration of Indo-Iranic speakers into Iran, Afghanistan, and South Asia. More specifically, the discriminative utility of the Y-haplogroups being tracked; the exact chronology of the migration events; plausible terminal locations for mini-migrations to relatively unpopulated areas, based on preferred habitats for the profession(s) of migrant(s); and, for required founder effects, evidence of significant cross-traffic to then amplify the founder Y-haplogroup signal, such as evidence for a rise in phylogenetic branching rate in the founder lineage—these have never before been provided for any competing theory of Indo-Iranic language-speaking human introgressions into the Persian/Afghan basin. For much of the twentieth century, Soviet archeologists promoted a narrative that conflated the dissemination of Sintashta military technology with a comprehensive settlement of Andronovo people and genes across the Persian Plateau and South Asian subcontinent. However, modern understandings cast great doubt upon this narrative. As previously stated, “within the entirety of the second millennium the only intrusive archeological culture that directly influences both Iran and North India is the BMAC” (Lamberg-Karlovsky 2002). The absence of Andronovo artifacts in South Asia and, even more damningly, the sheer math of the clearly large populations sizes on the Iranian plateau and Indus valley—many hundreds of thousands, based on settlement sizes—make it a tall order for a few hundred intrepid Andronovo nomads to influence populations in the Persian Basin and South Asia during the second millennium BCE in the profound genetic, linguistic and religious ways suggested. The out-of-Yamnaya migration phase of the pre-Vedic, Indo-Iranic language-speaking LPK migration to Rig-i-Stan, on the other hand, instantiates every aspect of the required dynamics.
By reconstructing suggestive archeological information from Velikent, we characterize pre-Vedic LPK ancestors as mobile transhumant cattle-drivers, specialized in mixed farming, specifically in arid intermontane river valleys (8–16” rain per annum) with proximity to severely arid plains. Along the route proposed by a previous study of LPK migration (Mascarenhas et al. 2015) such opportunities could be found at Velikent, Khorasan, and Rig-i-Stan. Agro-pastoralism is a proxy for productive and synergistic professional interactions between communities traditionally engaged in mobile pastoralism on the northern Steppe and southern neighbors involved in sedentary agriculture, especially during the crop expansion into the Steppe seen most dramatically at Tasbas and Ojakly during the second millennium BCE (Rouse et al. 2022; Spengler et al. 2014). Agro-pastoralism can also serve as a reference point for dating the first non-casual contacts between migrants from the north and settled farmers to the south. In semi-arid areas (8–16” rainfall per annum) especially, such as those found along the edges of the Karakum Desert with settled farmers, in the Kopet Dagh and BMAC oases (after 3800 YBP), and in Velikent in Dagestan (during 4700–4300 YBP), the synergies provided by agro-pastoralism would presumably have been motivating. It may have been a significant economic impetus for the staged migration of pre-Vedic, cattle-driving LPK ancestors. The well-documented archeological dates (Kohl 1988) for these important north–south interactions also provide termini post quem dates for the earliest significant southward migrations of Kurgan peoples on either side of the Caspian Sea—apparently separated by at least half a millennium and possibly much longer, at least using this economic marker. Such considerations add plausibility to the conclusions made with the new DNA clock.
The opportunity for founder effects involving the LPK Y-haplogroup, as described above, is to be found most clearly in Khorasan. Using a chronologically precise DNA clock, we advance a genetic argument for expansion of R-Y7 in Khorasan between 4300 and 3800 YBP, an interval that exactly coincides with an archeologically defined GKC/BMAC fluorescence. Next, we support the suggestion of disproportionate R-Y7/Y6 expansion with data showing a discontinuity in the phylogenetic branching rate for these two independent lineages, during the late GKC/BMAC period. Others have shown that an initially low frequency of a haplogroup can be amplified by genetic drift, even without positive selection (Kim et al. 2018). That such an event actually happened in this instance is further supported by the modern geographical footprint of the R-Y7 haplogroup shown in Table 1.
We also propose that the recorded aridity peak at 4.2 K YBP, which has been convincingly documented for regions between the Mediterranean and the Indus (Lawrence et al. 2021), could plausibly disturb established trade routes and drive a significant northward movement of people from third millennium settlements such a Shahr-i-Sukhta, Mundigak, Tepe Yahya, Konar Sandal, Bampur and Shahdad in southeastern Iran—all (archeologically) “abandoned” at about this time. Moreover, based on pottery styles and settlement sizes, a depopulation of large settlements in the Kpet Dagh such as Namazga and Altyn Depe may also point to a source of contemporaneous migration towards the Murghab (Kohl 1988; Kohl and Magomedov 2014).
The distinction between nucleated and dispersed settlement patterns in ancient times has been linked to cycles of climate change (Biscione and Vahdati 2020; Lawrence et al. 2021). During periods of denucleation, it is reasonable to believe that some percentage of the regional population will be obliged to relocate for work and survival. Thus, cyclic migration within a cultural kinship corridor (such as the Turan domain) could be triggered by an instance of climatic impetus, such as 4.2 K YBP aridity, followed by a reversal some centuries later, during a later phase of the long-term climatic cycle.
With archeobotanical correlates in support (Billings et al. 2022), further northward expansion of the proposed BMAC/GKC population rise during 4300–3800 YBP was limited by cultural and physical barriers, including the then-impassable Karalkum and Kyzylkum deserts. Meanwhile, logic suggests that rapid traffic growth along caravan trails crisscrossing Bactrian oases and other Khorasan settlements is likely to have occurred during this time period, thereby providing a key input to the founder-effect hypothesis. Conversely, aridity cycles may have led to reverse population shifts some centuries later from the BMAC/GKC southward to, say, Rig-i-Stan, during the so-called ‘expansion phase’ of BMAC/GKC prestige artifacts. Such artifacts have been documented in Seistan and Mundigak, especially in semi-arid locations (Biscione and Vahdati 2020)

3.2. The Catalytic Effects of Climate

Precipitation patterns were reconstructed from ENSO and Hemispherical Temperature Gradient Percent change (Molnar and Rajagopalan 2020) in annual precipitation relative to present day climatology (i.e., average precipitation during 1951–2015) during four broad epochs covering 5.5 K YBP to modern times were investigated. This encompasses the ~4.2 K YBP mega drought event when migrations were thought to have been exacerbated (Figure 6). During the period encompassing the mega drought ~4.5–3.5 K YBP in Figure 6B, this precipitation gradient is stronger and sharper. The wetter regions over Pakistan and northwestern India along Arabian Sea (mature phase of Harappan Civilization) show (25~75%) increase in precipitation and over a generally wetter (0~20%) Indian subcontinent. The climate reconstruction provides critical support for the proposed migration narrative in at least three major ways: (a) support for why the relative aridity of SE Iran during the BMAC/GKC florescence (Figure 6B) may have provided economic impetus for northward movement based on reduced food production; (b) support for why the relative aridity of SE Iran and the Arabian peninsula during the BMAC/GKC florescence (Figure 6B) may have provided economic impetus for northward movement of coastal communities based on the likely collapse of maritime trade through the Straits of Hormuz. Ancient maritime trade is widely believed to have involved small craft that traveled from point to point between littoral communities specialized in trans-shipment of particular goods, often engaging in value-added modification of those goods; and (c) support for why a collapse of maritime trade through the Straits of Hormuz may have, during the same time period, catastrophically affected a Harappan polity whose prosperity was predicated on exports to Mesopotamia markets. Conversely, this may also help explain the relative contemporaneous prosperity of the Khorasan-BMAC region as an alternative supplier of goods to the Mesopotamian consumer.

3.3. Using Chronological Specificity to Integrate Other Data Types

The reliability of models for ancient migration can benefit from cross-disciplinary linkages, especially in delineating interactions with non-elite layers of a migrant community, which make up the vast majority of any “migration” but are rarely featured by ancient chroniclers. In this study, cultural kinships were mapped circumstantially to the LPK in the post-Vedic period—thereby enriching their migration narrative. One was the geographic and cultural kinship of the LPK with pre-Vedic, Indo-Iranic language-speaking, Yadava Bhojas of Afghanistan [Supplementary Material Text File ST2]; another, with the Kamboja people [Supplementary Material Text File ST3]. Yet another example, also circumstantial, includes commoners with proto-Elamite antecedents as a social layer that co-locates with the LPK’s recent migration geography. We show that Proto-Elamite numismatic symbols were geographically co-extant with the LPK during their introgression into the subcontinent. Such symbols were often included by rulers and tribes to acknowledge the religious preferences of working classes.
From ancient religious records it appears that the Saraswat Brahmin LPK clan adopted a new priestly profession in Rig-i-Stan ca. 3500 YBP. As Brahmanical priests, the LPK now created synergistic professional relationships with new groups, notably with the Kamboja people—initially as providers of religious rituals to warriors who shared a nomadic pastoralist history with the LPK, then with the trade-related activities of Kambojas in Gujarat. This reflects the LPK’s professional adjustment to an increasingly secular world, beginning sometime in the early Common Era. [Supplementary MaterialText File ST3]. After some LPK were converted to Christianity in the late sixteenth century, convert LPK Brahmins became closely associated with Catholic priesthood as a profession, and especially with the Jesuit Order.
All these examples taken together illustrate—not just the significance of precise dating in describing migrations, but a case for the inclusion of data types that help us understand more about the people who took the long journey that we wish to document. We conclude that more precise dating can help enrich ancient migration narratives by adding confidence to our conclusions with the multiplicity of circumstantial links created with other data types.
The following limitations of our work should be acknowledged. R-Y7 as an Indo-Iranian marker is backed by circumstantial evidence. The assumption that drought is the major driver of the data depicted in Figure 6 is also, strictly speaking, circumstantial. Our methodology does not include a sensitivity analysis or any kind of quantitative linguistic/genetic modeling (e.g., no Bayesian phylogeny for splits), and the founder effects proposed here assume a sparse Khorasan population, when evidence for that fact is interpretive of archeological data. Cultural kinships proposed here sometimes rely on proximity without causal proof. Biases noted here for FT-DNA and Y-Full databases are not mitigated (e.g., bootstrap). The focus of this study on a single patrilineage can limit the generalization of its findings.

4. Materials and Methods

4.1. Samples and Subjects

The collection of DNA under informed consent from seven male LPK Saraswat Brahmins of Vangor 8 (as described in (Mascarenhas et al. 2015)) was performed as follows: volunteers registered for cheek swab DNA analysis at the YSEQ laboratory (YSEQ GmbH, Berlin, Germany), signed informed consent forms, received a kit and followed the requisite procedures specified by the YSEQ laboratory. Each volunteer ordered NGS whole-genome sequencing (WGS). Volunteers were reimbursed for the cost of the sequencing but received no other compensation. After testing was complete, volunteers provided investigators with access to their own results electronically. Permission to use data pertaining to Y-chromosome sequences for research and publication was provided by each subject.

4.2. Ethics Statement

This work was the second phase of a previously published study using the same subjects’ DNA for STR analysis (Mascarenhas et al. 2015). All subjects are college-educated. For this study, the subjects themselves purchased YSEQ kits for genomic analysis and forwarded mouth swabs for analysis to the YSEQ laboratory at their own cost, which was later reimbursed by the study budget. Each subject forwarded his own Y-chromosome BAM file via email to the Principal Investigator with written permission to use the data for this study. Per agreement with each subject, such copies of raw sequence data files were subsequently destroyed upon completion of the analysis. The protocol and computer analyses were approved by the Ethical Committee of the Mayflower Organization for Research and Education, Sunnyvale, CA.

4.3. DNA Analysis

Whole genome sequencing by the YSEQ laboratory was performed at 2 × 250 base runs using CeGaT technology. The resulting sequence data were mapped against a common reference. This widely used reference sequence comprises Sanger sequencing results of a R1b-U152 sample patched in a few sections with a sample of a haplogroup G man. The following hg38 regions of the Y chromosome suffer from frequent recombination events and are therefore not as useful for phylogenetic studies (Xu and Pang 2022):chrY:1..2781479 (pseudo autosomal region 1, PAR1); chrY:10072350..11686750 (synthetic assembled centromeric region, CEN); chrY:20054914..20351054 (DYZ19 125 bp repeat region); chrY:26637971..26673210 (post palindromic region, gradual start of Yq12 repetitive region); chrY:56887903..57217415 (pseudo autosomal region 2, PAR2). The 50 microsatellite short tandem repeats (STR) subset used in the refined analysis of the original data in (Mascarenhas et al. 2015) were DYS19, DYS385a, DYS385b, DYS388, DYS389a, DYS389b, DYS390, DYS391, DYS392, DYS393, DYS426, DYS436, DYS437, DYS438, DYS439, DYS442, DYS444, DYS446, DYS447, DYS448, DYS449, DYS454, DYS455, DYS456, DYS458, DYS459a, DYS459b, DYS460, DYS464a, DYS464b, DYS464c, DYS472, DYS481, DYS492, DYS511, DYS520, DYS531, DYS534, DYS537, DYS557, DYS565, DYS568, DYS570, DYS576, DYS590, DYS607, DYS640, YCAIIa, YCAIIb, Y-GATA-H4.

4.4. Baptismal Database

Over one thousand handwritten baptismal records were manually transcribed over a period of eight years. These records were from the parish church of Lotli, Salvador do Mundo, for records from 1850 to 1900 CE; and from Government Archives located in Panaji, Goa for records dated between1613 and 1850 CE. Church records for Lotli town for the period 1914–1950 were additionally obtained on microfilm from the Church of Latter Day Saints, Salt Lake City, UT. As these records are from poorly stored, often water- or insect-damaged books handwritten by parish priests over four centuries, not all baptismal records could be satisfactorily recovered or deciphered. Nevertheless, we estimate that at least 90% of all records were successfully tabulated. Records were then translated from the original Portuguese to English, with archaic Portuguese expressions and abbreviations decoded, where necessary, by the researcher carrying out the manual transcription. Each record typically lists three generations of data (child’s name and date of baptism—by tradition, typically within 3–7 days of birth, except in cases of serious illness; parents’ names, social class, and land ownership status and town of origin for spouses, grandparents’ names and godparents’ names). A comprehensive baptismal tree was constructed from these partially redundant records covering over 400 years of the family branches from each of the present-day male LPK DNA donors who participated in this study (save one, whose family was not converted to Christianity). Baptismal data shown in Figure 1 were appropriately redacted for recent generations, in order to protect the privacy of donors. As the family is well known in the town community, and at their request, in order to protect privacy, baptismal data additional to what is disclosed in this work is not authorized for release.

4.5. STR Analysis

More precise TMRCA calculations were performed using the original data and permissions gathered for the previous study, by the method specified there (Mascarenhas et al. 2015), except that instead of calculating TMRCA based on all available STRs, the new analysis recalculated all comparisons using only a common set of 50 STRs, which are listed above. Values for that identical STR set are available from the public database, FT-DNA (www.familytreedna.com; downloaded March 2025), and these used to calculate modal values for populations of other branches of the R-Z94 tree, allowing for an updated TMRCA computation for Z94, Y3, L-657, Y16494 branch points shown in Table 2, using modal LPK STR values as the common comparator.

4.6. Y-Full Genomic Data

A total of 40 reference genomic sequences for mutations between R-Z94 and R-Y16494 were kindly provided by the Y-Full public database collection through the kind offices of the YSEQ corporation. Formation times for these mutations, and the values used in those computations, were downloaded from the public Y-Full website (www.yfull.com) in March 2025.

4.7. Construction of a Y-DNA Clock

In order to maintain time intervals large enough to minimize short-term variance effects, the historically defined family branch point “0” (Madhava conversion, 1296 CE; (Mascarenhas et al. 2015) was selected as a reference point for calculation of genetic distances from transversions and transitions actually sequenced. Twelve overlapping intervals were used in the calculation using historical branch points. Comparisons were made iteratively between the historical set of baptismal and other historical branch points included in the 0–704 YBP calculation interval. A best fit of these data was achieved, and an interval of 149.8 years could then be assigned for transversions and half that (49.9 years) to transitions (Le and Durbin 2011). The correlation plot for historical reference dates and mutation dates calculated using the 149.8 year transversion value (and half value for transitions) is shown in the top panel of Figure 2a (r2 = 0.995). This clock formula was then validated against an STR-based calculation of TMRCA dates between 704 and 5200 YBP using a standardized set of 50 STRs as described above (Figure 2a, middle panel). Finally, it was validated against Y-Full formation dates from R-Z94 formation to 3000 YBP (Figure 2a, bottom panel). ‘Present day’ was arbitrarily set to 2000 CE in all cases. Mutation ages calculated by the new DNA clock for the entire mutational history of the LPK patrilineage going back to R-Z94 are shown in Table 2.

4.8. Phylogenetic Branching Index

A phylogenetic branching analysis was computed for R-Y7 and sister clade R-Y6 as a control, for successive mutation intervals on the R-Y7 tree branch from R-L657 to R-Y2428. This interval has a DNA clock chronology (average of Y-Full and this study) of 4393–3623 YBP from the end of the first to the last mutation interval. This set of intervals covers the BMAC/GKC expansion period. The average mutation formation rate during this period is 124.0 years—compared to 129.5 years computed by an independent procedure using STR-TMRCA for the interval R-Z94 to R-Y16494—which covers over two millenia, and brackets this era. Thus, the rate of mutation along the R-Y7 lineage does not appear to accelerate during the BMAC/GKC era. Phylogenetic branching rate was measured based on number of new branches sprouted off the same mutation interval (as shown on the current Y-Full tree). Branch lengths are not considered in this computation (Moody et al. 2022).

4.9. Paleoclimate Reconstructions

The El Nino Southern Oscillation (ENSO) is a robust ocean-atmosphere coupled phenomenon in the equatorial Pacific Ocean that is a major driver of global climate at seasonal and multi-year time scales. Cooler eastern and central equatorial Pacific Ocean Sea Surface Temperature (SST) anomalies (i.e., La Niña conditions) correlates strongly with wetter summer monsoon over Indian subcontinent, which includes the Indus Valley Civilization region and, drier conditions over central Asia (Gill et al. 2015; Feng et al. 2022). An opposite climate effect occurs with warmer eastern and central equatorial SSTs (i.e., El Niño conditions). Other regional and local climate factors notwithstanding, one can use the SST index (average SST anomalies in the NINO3.4 region) that quantify the ENSO to capture the key aspects (i.e., ‘signal’) of precipitation variability in these regions. This approach forms the basis of our baseline paleoclimate reconstruction of precipitation. Statistical space-time models (Ossandon et al. 2024; Gill et al. 2017) are developed on the SST data over equatorial Pacific Ocean covering the modern historical period. Researchers have drilled marine core sediments at several locations in the equatorial Pacific and using Geochemical analyses have reconstructed SSTs, covering the Holocene period (see Gill et al. 2017 for details). These core location data combined with the statistical models enable the reconstruction of SSTs over the entire equatorial Pacific region (Ossandon et al. 2024; Gill et al. 2017). Monthly precipitation data is available for the entire globe on a 5° × 5° grid (Molnar and Rajagopalan 2020) for the modern historical period, from which annual precipitation (May–April average) is computed and consequently, the change in precipitation (as percentage) relative to the modern period (1950~present) average. Marcott et al. (Marcott et al. 2013) compiled the temperature gradient between the northern and southern hemispheres during the mid-Holocene period (4~5 K BP) as ~0.8 °C and ~0.1 °C in late Holocene. This gradient is also known to impact the monsoons over the Indian subcontinent (Gill et al. 2017; Zhao and Harrison 2012) and North Africa (Molnar and Rajagopalan 2020). Linear regression models are fitted between annual precipitation change and the ENSO index and hemispherical temperature gradient, at each grid, covering the Central Asia and Indian subcontinent. With the reconstructed the ENSO index and the regressions, precipitation changes are reconstructed at various time periods during the Holocene. This approach was used in the reconstruction of mid-Holocene precipitation change over Sahel region (Molnar and Rajagopalan 2020). All the data, modern period SSTs, Holocene SSTs from marine sediment cores, global precipitation and hemispherical temperature averages are freely available and described in the references above. During the mid-Holocene (4~6 ka BP), SSTs in the east central equatorial Pacific region (NINO3.4) were much (~1~2 °C) cooler (Ossandon et al. 2024; Gill et al. 2017) than today. Moreover, Global Circulation Models (GCMs) also confirm this cooling with reduced El Niño occurrences (Osman et al. 2021). This cooling along with the stronger hemispherical temperature gradient enhanced the monsoonal circulation resulting in wetter Indian subcontinent, especially the Indus Valley region and drier central Asia (Feng et al. 2022). Nominal standard errors obtained from the regression model used in the paleoclimate reconstructions of Figure 6 are shown in Supplementary Material Text File ST4. Uncertainties exist in estimates of ENSO and hemispherical temperature gradient, the two predictors used in the regression model. These nominal errors are akin to those offered in Gill et al. (2015, 2017). A Bayesian framework (e.g., Ossandon et al. 2024) can enable a more robust treatment of these and other uncertainties.

4.10. Statistical Methods

Probability values (p values) were computed using Student’s T-test and expressed relative to indicated controls, except where otherwise noted. Group size was as noted in each case.

5. Conclusions

An accurate DNA clock can strengthen cross-disciplinary inputs in the study of genealogies and ancient migrations. The availability of more precise dating facilitates the integration of secondary data types, resulting in an enriched and more plausible migration narrative.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/genealogy10010014/s1. (Haber et al. 2016; Handa 2007; Kumar 2019; Mehrjoo et al. 2019; Pathak et al. 2018; Pinto 2023; Thapar 1984; Warmington and Thapar 2015) are cited in Supplementary Materials File.

Author Contributions

Conceptualization, D.D.M. and B.R.; Methodology, D.D.M. and B.R.; Formal analysis, D.D.M. and B.R.; Investigation, D.D.M. and B.R.; Resources, D.D.M., B.R. and R.J.J.; Writing—original draft, D.D.M.; Writing—review & editing, D.D.M., B.R., J.W.F. and R.J.J.; Project administration, D.D.M. and R.J.J. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki, and approved by the Mayflower Research Institute Ethical Committee (protocol code MRIEC-20190408 and date of approval 2019-04-08).

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

Data may be obtained, subject to privacy and disclosure constraints, by emailing Desmond Mascarenhas at the Mayflower Research Institute (desmond2@mayflowerworld.org). Because the combination of baptismal tree and genomic sequences can make subjects precisely and individually identifiable, no additional data of either type will be released beyond what is disclosed in the article.

Conflicts of Interest

The authors declare no conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
LPKLotli Pai Kaundinya
Y-DNAY-chromosome DNA
YBPYears Before Present
STRShort Tandem Repeats
TMRCATime to Most Recent Common Ancestor
BMACBactria Margiana Archeological Complex
GKCGreater Khorasan Culture

References

  1. Billings, Traci N., Barbara Cerasetti, Luca Forni, Roberto Arciero, Rita Dal Martello, Marialetizia Carra, Lynne M. Rouse, Nicole Boivin, and Robert N. Spengler. 2022. Agriculture in the Karakum: An archaeobotanical analysis from Togolok 1, southern Turkmenistan (ca. 2300–1700 B.C.). Frontiers in Ecology and Evolution 10: 995490. [Google Scholar] [CrossRef] [Scilit]
  2. Biscione, Raffaele, and Ali A. Vahdati. 2020. The BMAC presence in eastern Iran: State of affairs in December 2018—Towards the Greater Khorasan Civilization? In The World of the Oxus Civilization, 1st ed. London: Routledge, Ch.19. p. 24. [Google Scholar]
  3. Blazec, Vaclav. 2007. From August Schleicher to Sergei Starostin: On the development of the tree-diagram models of the Indo-European languages. Journal of Indoeuropean Studies 35: 82–109. [Google Scholar]
  4. Bonito, Maria, Eugenia D’aTanasio, Francesco Ravasini, Selene Cariati, Andrea Finocchio, Andrea Novelletto, Beniamino Trombetta, and Fulvio Cruciani. 2021. New insights into the evolution of human Y chromosome palindromes through mutation and gene conversion. Human Molecular Genetics 30: 2272–85. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Dallmeyer, Anne, Martin Claussen, Yongbo Wang, and Ulrike Herzschuh. 2013. Spatial variability of Holocene changes in the annual precipitation pattern: A model-data synthesis for the Asian monsoon region. Climate Dynamics 40: 2919–36. [Google Scholar] [CrossRef] [Scilit]
  6. Djamali, Morteza, Jacques-Louis de Beaulieu, Valérie Andrieu-Ponel, Manuel Berberian, Naomi F. Miller, Emmanuel Gandouin, Hamid Lahijani, Majid Shah-Hosseini, Philippe Ponel, Mojtaba Salimian, and et al. 2009. A late Holocene pollen record from Lake Almalou in NW Iran: Evidence for changing land-use in relation to some historical events during the last 3700 years. Journal of Archaeological Science 36: 1364–75. [Google Scholar] [CrossRef] [Scilit]
  7. Dyson, Tim. 2018. A Population History of India: From the First Modern People to the Present Day. Oxford: Oxford University Press. [Google Scholar]
  8. Eskandaridamne, Nasir. 2021. A Landscape Archaeology in Shahdad Plain (Dasht-e Lut) from the 5th to the 2nd Millennium BC. Ph.D. dissertation, Université de Lyon and University of Tehera, Tehran, Iran. [Google Scholar]
  9. Feng, Fan, Yong Zhao, Anning Huang, Yang Li, and Xin Zhou. 2022. Different Seasonal Precipitation Anomaly Patterns in Central Asia Associated with Two Types of El Niño During 1891–2016. Frontiers in Earth Science 10: 771362. [Google Scholar] [CrossRef] [Scilit]
  10. Fouache, Éric, Henri-Paul Francfort, Julio Bendezu-Sarmiento, Ali Akbar Vahdati, and Johanna Lhuillier. 2010. The Horst of Sabzevar and regional water resources from the Bronze Age to the present day (Northeastern Iran). Geodinamica Acta 23: 287–94. [Google Scholar] [CrossRef] [Scilit]
  11. Ghalichi, Ayshin, Sabine Reinhold, Adam B. Rohrlach, Alexey A. Kalmykov, Ainash Childebayeva, He Yu, Franziska Aron, Lena Semerau, Katrin Bastert-Lamprichs, Andrey B. Belinskiy, and et al. 2024. The rise and transformation of Bronze Age pastoralists in the Caucasus. Nature 635: 917–25. [Google Scholar] [CrossRef] [Scilit]
  12. Gill, Emily C., Balaji Rajagopalan, and Peter Molnar. 2015. Subseasonal variations in spatial signatures of ENSO on the Indian summer monsoon from 1901 to 2009. Journal of Geophysical Research: Atmospheres 120: 8165–85. [Google Scholar] [CrossRef] [Scilit]
  13. Gill, Emily C., Balaji Rajagopalan, Peter H. Molnar, Yochanan Kushnir, and Thomas M. Marchitto. 2017. Reconstruction of Indian summer monsoon winds and precipitation over the past 10,000 years using equatorial pacific SST proxy records. Paleoceanography and Paleoclimatology 32: 195–216. [Google Scholar] [CrossRef] [Scilit]
  14. Goldberg, Amy, Torsten Günther, Noah A. Rosenberg, and Mattias Jakobsson. 2017. Ancient X chromosomes reveal contrasting sex bias in Neolithic and Bronze Age Eurasian migrations. Proceedings of the National Academy of Sciences USA 114: 2657–62. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Haber, Marc, Massimo Mezzavilla, Yali Xue, David Comas, Paolo Gasparini, Pierre Zalloua, and Chris Tyler-Smith. 2016. Genetic evidence for an origin of the Armenians from Bronze Age mixing of multiple populations. European Journal of Human Genetics 24: 931–36. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Handa, Devendra. 2007. Tribal Coins of Ancient India. New Delhi: Aryan Books. [Google Scholar]
  17. Heggarty, Paul, Cormac Anderson, Matthew Scarborough, Benedict King, Remco Bouckaert, Lechosław Jocz, Martin Joachim Kümmel, Thomas Jügel, Britta Irslinger, Roland Pooth, and et al. 2023. Language trees with sampled ancestors support a hybrid model for the origin of Indo-European languages. Science 381: eabg0818. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Kim, Bernard Y., Christian D. Huber, and Kirk E. Lohmueller. 2018. Deleterious variation shapes the genomic landscape of introgression. PLoS Genetics 14: e1007741. [Google Scholar] [CrossRef] [Scilit]
  19. Kohl, Philip L. 1988. The Northern “Frontier” of the Ancient Near East: Transcaucasia and Central Asia Compared. American Journal of Archaeology 92: 591–96. [Google Scholar] [CrossRef] [Scilit]
  20. Kohl, Philip L., and Rabadan G. Magomedov. 2014. Early Bronze developments on the West Caspian Coastal Plain. Paléorient 40: 93–114. [Google Scholar] [CrossRef] [Scilit]
  21. Kortlandt, Frederik. 2016. Balto-Slavic and Indo-Iranian. Baltistica 51: 355–64. [Google Scholar] [CrossRef] [Scilit]
  22. Kumar, Vinay. 2019. Black and Red Ware Culture: A Reappraisal. Heritage: Journal of Multidisciplinary Studies in Archaeology 7: 397–404. [Google Scholar]
  23. Lamberg-Karlovsky, Carl C. 2002. Archaeology and Language: The Indo-Iranians. Current Anthropology 43: 63–88. [Google Scholar] [CrossRef] [Scilit]
  24. Lawrence, Dan, Alessio Palmisano, and Michelle W. de Gruchy. 2021. Collapse and continuity: A multi-proxy reconstruction of settlement organization and population trajectories in the Northern Fertile Crescent during the 4.2kya Rapid Climate Change event. PLoS ONE 16: e0244871. [Google Scholar] [CrossRef] [Scilit]
  25. Lazaridis, Iosif, Nick Patterson, David Anthony, Leonid Vyazov, Romain Fournier, Harald Ringbauer, Iñigo Olalde, Alexander A. Khokhlov, Egor P. Kitov, Natalia I. Shishlina, and et al. 2025. The genetic origin of the Indo-Europeans. Nature 639: 132–42. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Le, Si Quang, and Richard Durbin. 2011. SNP detection and genotyping from low-coverage sequencing data on multiple diploid samples. Genome Research 21: 952–60. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Marcott, Shaun A., Jeremy D. Shakun, Peter U. Clark, and Alan C. Mix. 2013. A Reconstruction of Regional and Global Temperature for the Past 11,300 Years. Science 339: 1198–201. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Mascarenhas, Desmond D., Anupuma Raina, Christopher E. Aston, and Dharambir K. Sanghera. 2015. Genetic and Cultural Reconstruction of the Migration of an Ancient Lineage. BioMed Research International 2015: 651415. [Google Scholar] [CrossRef] [Scilit]
  29. Maurer, Gwendoline, and Raphael Greenberg. 2022. Cattle drivers from the north? Animal economy of a diasporic Kura-Araxes community at Tel Bet Yerah. Levant 54: 309–30. [Google Scholar] [CrossRef] [Scilit]
  30. Mehrjoo, Zohreh, Zohreh Fattahi, Maryam Beheshtian, Marzieh Mohseni, Hossein Poustchi, Fariba Ardalani, Khadijeh Jalalvand, Sanaz Arzhangi, Zahra Mohammadi, Shahrouz Khoshbakht, and et al. 2019. Distinct genetic variation and heterogeneity of the Iranian population. PLoS Genetics 15: e1008385. [Google Scholar] [CrossRef] [Scilit]
  31. Molnar, Peter, and Balaji Rajagopalan. 2020. Mid-Holocene Sahara-Sahel Precipitation From the Vantage of Present-Day Climate. Geophysical Research Letters 47: e2020gl088171. [Google Scholar] [CrossRef] [Scilit]
  32. Moody, Edmund Rr, Tara A Mahendrarajah, Nina Dombrowski, James W. Clark, Celine Petitjean, Pierre Offre, Gergely J. Szöllősi, Anja Spang, and Tom A. Williams. 2022. An estimate of the deepest branches of the tree of life from ancient vertically evolving genes. eLife 11: e66695. [Google Scholar] [CrossRef] [Scilit]
  33. Moradi, Hossein, Hamed Tahmasebi Zave, Ali Akbar Eshghi, and Hossein Sarhaddi-Dadian. 2023. The Interactions between Sistan and Great Khorasan Culture (GKC) in the Second Half of the Third Millennium BC. Journal of Sistan and Baluchistan Studies 3: 49–63. [Google Scholar]
  34. Nakhleh, Luay, Donald A. Ringe, and Tandy Warnow. 2005. Perfect Phylogenetic Networks: A New Methodology for Reconstructing the Evolutionary History of Natural Languages. Language 81: 382–420. [Google Scholar] [CrossRef] [Scilit]
  35. Osman, Matthew B., Jessica E. Tierney, Jiang Zhu, Robert Tardif, Gregory J. Hakim, Jonathan King, and Christopher J. Poulsen. 2021. Globally resolved surface temperatures since the Last Glacial Maximum. Nature 599: 239–44. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Ossandon, Álvaro, Javier Gual, Balaji Rajagopalan, William Kleiber, and Thomas Marchitto. 2024. Spatial and temporal Bayesian hierarchical model over large domains with application to paleo sea surface temperature reconstruction. Paleoceanography and Paleoclimate 39: e2024pa004844. [Google Scholar] [CrossRef] [Scilit]
  37. Pamjav, Horolma, Tibor Fehér, Endre Németh, and Zsolt Pádár. 2012. Brief communication: New Y-chromosome binary markers improve phylogenetic resolution within haplogroup R1a1. American Journal of Physical Anthropology 149: 611–15. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Pathak, Ajai K., Anurag Kadian, Alena Kushniarevich, Francesco Montinaro, Mayukh Mondal, Linda Ongaro, Manvendra Singh, Pramod Kumar, Niraj Rai, Jüri Parik, and et al. 2018. The Genetic Ancestry of Modern Indus Valley Populations from Northwest India. American Journal of Human Genetics 103: 918–29. [Google Scholar] [CrossRef] [Scilit]
  39. Pinto, Celsa. 2023. Concise History of Goa, 1st ed. Saligao: Goa-1556, pp. 24–27. [Google Scholar]
  40. Rouse, Lynne M., Paula N. Doumani Dupuy, and Elizabeth Baker Brite. 2022. The Agro-pastoralism debate in Central Eurasia: Arguments in favor of a nuanced perspective on socio-economy in archaeological context. Journal of Anthropological Archaeology 67: 101438. [Google Scholar] [CrossRef] [Scilit]
  41. Salvatori, Sandro. 2000. Bactria and Margiana Seals: A New Assessment of Their Chronological Position and a Typological Survey. East and West 50: 97–145. [Google Scholar]
  42. Sankalia, Hasmukh D. 1963. New Light on the Indo-Iranian or Western Asiatic Relations between 1700 BC-1200 BC. Artibus Asiae 26: 312. [Google Scholar] [CrossRef] [Scilit]
  43. Serva, Maurizio, and Filippo Petroni. 2008. Indo-European languages tree by Levenshtein distance. Europhysics Letters 81: 68005. [Google Scholar] [CrossRef] [Scilit]
  44. Spengler, Robert N., Michael D. Frachetti, and Paula N. Doumani. 2014. Late Bronze Age agriculture at Tasbas in the Dzhungar Mountains of eastern Kazakhstan. Quaternary International 348: 147–57. [Google Scholar] [CrossRef] [Scilit]
  45. Stevens, Chris J., Charlene Murphy, Rebecca Roberts, Leilani Lucas, Fabio Silva, and Dorian Q Fuller. 2016. Between China and South Asia: A Middle Asian corridor of crop dispersal and agricultural innovation in the Bronze Age. The Holocene 26: 1541–55. [Google Scholar] [CrossRef] [Scilit]
  46. Thapar, Romila. 1984. From Lineage to State: Social Formations in the Mid-First Millennium B.C. in the Ganga Valley. Oxford: Oxford University Press. [Google Scholar]
  47. Vahdati, Ali Akbar, Raffaele Biscione, Riccardo La Farina, Marjan Mashkour, Margareta Tengberg, Homa Fathi, and Azadeh Fatemeh Mohaseb. 2019. Preliminary report on the first season of excavations at Tepe Chalow: New GKC (BMAC) finds in the plain of Jajarm, NE Iran. The Iranian Plateau during the Bronze Age. Development of Urbanisation, Production and Trade 1: 179–200. [Google Scholar]
  48. Wang, Jianjun, Liguang Sun, Liqi Chen, Libin Xu, Yuhong Wang, and Xinming Wang. 2016. The abrupt climate change near 4400 yr BP on the cultural transition in Yuchisi, China and its global linkage. Scientific Reports 6: 27723. [Google Scholar] [CrossRef] [Scilit]
  49. Warmington, Eric Herbert, and Romila Thapar. 2015. Barygaza. In Oxford Research Dictionary. Oxford: Oxford University Press. [Google Scholar] [CrossRef] [Scilit]
  50. Xu, Yong, and Qianqian Pang. 2022. Repetitive DNA Sequences in the Human Y Chromosome and Male Infertility. Frontiers in Cell and Developmental Biology 10: 831338. [Google Scholar] [CrossRef] [Scilit]
  51. Zhao, Yan, and S. P. Harrison. 2012. Mid-Holocene monsoons: A multi-model analysis of the inter-hemispheric differences in the responses to orbital forcing and ocean feedbacks. Climate Dynamics 39: 1457–87. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Summary of strategy for DNA clock construction. (A) Examples of relative geographic provenance for discriminating Y-haplogroups in FT-DNA database, showing discriminant distribution of R-Z94 subclades for associating Indo-Iranian (R-Y3) and Balto-Slav (R-Y2619) language carriers and negative control for Andronovo (R-FGC56408). Locations of Yamnaya koine, R-Y7 pool in Khorasan, and R-Y7 LPK patrilineage are also shown. (B) Genomes in Y-Full database (March 2025) showing large standard deviations for computed age of mutations below R-Z94 in the R-Y7 lineage, until R-Y16494, below the LPK branch point, wherein the number of genomes available is n = 6. See text for significance of these numbers. (C) Example archival records of baptismal dates used to construct the Vangor 8 clan >400-year-old complete family tree, used in selecting men for genomic sequencing. Branch points correspond to those shown in Table 2.
Figure 1. Summary of strategy for DNA clock construction. (A) Examples of relative geographic provenance for discriminating Y-haplogroups in FT-DNA database, showing discriminant distribution of R-Z94 subclades for associating Indo-Iranian (R-Y3) and Balto-Slav (R-Y2619) language carriers and negative control for Andronovo (R-FGC56408). Locations of Yamnaya koine, R-Y7 pool in Khorasan, and R-Y7 LPK patrilineage are also shown. (B) Genomes in Y-Full database (March 2025) showing large standard deviations for computed age of mutations below R-Z94 in the R-Y7 lineage, until R-Y16494, below the LPK branch point, wherein the number of genomes available is n = 6. See text for significance of these numbers. (C) Example archival records of baptismal dates used to construct the Vangor 8 clan >400-year-old complete family tree, used in selecting men for genomic sequencing. Branch points correspond to those shown in Table 2.
Genealogy 10 00014 g001
Figure 2. Chronology of Migration. (a) Precision of DNA mutation clock constructed as described in Methods, and validated by baptismal and historical records (top panel), STR analysis (middle panel) and Y-Full estimates (bottom panel); (b) coincident events during GKC/BMAC fluorescence period (4300–3800 YBP): phylogenetic branching rate for independent sister clades R-Y7 and R-Y6 (top panel) and depopulation of Shahdad region, i.e., Lut–Data from (Eskandaridamne 2021) (bottom panel). YBP = years before present. Map lines delineate study areas and do not necessarily reflect accepted geopolitical boundaries for the periods studied. * historical reference point.
Figure 2. Chronology of Migration. (a) Precision of DNA mutation clock constructed as described in Methods, and validated by baptismal and historical records (top panel), STR analysis (middle panel) and Y-Full estimates (bottom panel); (b) coincident events during GKC/BMAC fluorescence period (4300–3800 YBP): phylogenetic branching rate for independent sister clades R-Y7 and R-Y6 (top panel) and depopulation of Shahdad region, i.e., Lut–Data from (Eskandaridamne 2021) (bottom panel). YBP = years before present. Map lines delineate study areas and do not necessarily reflect accepted geopolitical boundaries for the periods studied. * historical reference point.
Genealogy 10 00014 g002
Figure 3. Topographical map of the proposed migration route followed by Pre-Vedic ancestors of the LPK patrilineage from the Pontic Steppe to Rig-i-Stan. The route follows the clan’s preferred habitat, deduced as described in the text. Stage “3” indicates two possible parallel routes—via the Atrek River Valley or through the ancient Hissar trade route, both options terminating in Khorasan-Razavi/Herat. Map lines delineate study areas and do not necessarily reflect accepted geopolitical boundaries for the periods studied.
Figure 3. Topographical map of the proposed migration route followed by Pre-Vedic ancestors of the LPK patrilineage from the Pontic Steppe to Rig-i-Stan. The route follows the clan’s preferred habitat, deduced as described in the text. Stage “3” indicates two possible parallel routes—via the Atrek River Valley or through the ancient Hissar trade route, both options terminating in Khorasan-Razavi/Herat. Map lines delineate study areas and do not necessarily reflect accepted geopolitical boundaries for the periods studied.
Genealogy 10 00014 g003
Figure 4. Migration of LPK. (a) Proposed sandwich effect during GKC/BMAC florescence caused by climate (drought), possible population movements (white arrows show one direction) in “corridor” constricted by Turan cultural and geographical barriers (deserts and mountains). Location of discrete autosomal pools GP1-4, of which ancient GP3 (Turkmen) and GP4 (SE Iran/Baloch) pools may act as additional barriers to constrain geographic opportunities for founder genetic effects during the proposed GKC/BMAC zone for R-Y7/R-Y6 phylogenetic expansion during 4300–3800 YBP florescence; (b) 3700–3200 YBP dissemination era in the ancient world: multidirectional exchanges involving the Persian Plateau, China, Sintashta (Urals), Andronovo and South Asia. White Arrows show presumed dissemination routes for technologies, culture, and (possibly) peoples. YBP = years before present. Map lines delineate study areas and do not necessarily reflect accepted geopolitical boundaries for the periods studied.
Figure 4. Migration of LPK. (a) Proposed sandwich effect during GKC/BMAC florescence caused by climate (drought), possible population movements (white arrows show one direction) in “corridor” constricted by Turan cultural and geographical barriers (deserts and mountains). Location of discrete autosomal pools GP1-4, of which ancient GP3 (Turkmen) and GP4 (SE Iran/Baloch) pools may act as additional barriers to constrain geographic opportunities for founder genetic effects during the proposed GKC/BMAC zone for R-Y7/R-Y6 phylogenetic expansion during 4300–3800 YBP florescence; (b) 3700–3200 YBP dissemination era in the ancient world: multidirectional exchanges involving the Persian Plateau, China, Sintashta (Urals), Andronovo and South Asia. White Arrows show presumed dissemination routes for technologies, culture, and (possibly) peoples. YBP = years before present. Map lines delineate study areas and do not necessarily reflect accepted geopolitical boundaries for the periods studied.
Genealogy 10 00014 g004
Figure 5. Cultural Fellow Travelers. (A) Chronology of pottery styles migrating across Karakum/Kyzylkum barrier from China to NW Iran. (B) Religious motifs in seals and coinage. (a) 6-arch hill, Elam (5100–4900 YBP); (b) left seal/right drawing, 6-arch hill, Elam (5100–4800 YBP); (c) 6-arch hill, Elam (4100 YBP); (d) 6-arch hill, NW India (2200–1800 YBP); (e) 6-arch hill and Vasudeva-Sankarshana, Afghanistan (2200–2000 YBP); (f) 6-arch hill, N. India (2200–1800 YBP); (g) Coinage series of Mathura and Saurashtra Satraps (selected examples): Vrishni-Bhoja icons (Mathura), Vrishni-Bhoja (Saurashtra, 34 CE), Vrishni-Bhoja (Saurashtra, 78 CE), 6-arch hill (Saurashtra, 130 CE), 6-arch hill (Saurashtra, 181 CE), 3-arch hill (Saurashtra, 222 CE). (C) Proposed Introgressions of pre-Vedic, Indo-Iranic-speaking tribes into the Indian subcontinent. Pr-X = in Khorasan; Pr-Su = Ahar-Banas, Pr-Mh = Maharashtra; Pr-Mg = Magadhi; LPK = Lotli Pai Kaundinya; Map lines delineate study areas and do not necessarily reflect accepted geopolitical boundaries for the periods studied.
Figure 5. Cultural Fellow Travelers. (A) Chronology of pottery styles migrating across Karakum/Kyzylkum barrier from China to NW Iran. (B) Religious motifs in seals and coinage. (a) 6-arch hill, Elam (5100–4900 YBP); (b) left seal/right drawing, 6-arch hill, Elam (5100–4800 YBP); (c) 6-arch hill, Elam (4100 YBP); (d) 6-arch hill, NW India (2200–1800 YBP); (e) 6-arch hill and Vasudeva-Sankarshana, Afghanistan (2200–2000 YBP); (f) 6-arch hill, N. India (2200–1800 YBP); (g) Coinage series of Mathura and Saurashtra Satraps (selected examples): Vrishni-Bhoja icons (Mathura), Vrishni-Bhoja (Saurashtra, 34 CE), Vrishni-Bhoja (Saurashtra, 78 CE), 6-arch hill (Saurashtra, 130 CE), 6-arch hill (Saurashtra, 181 CE), 3-arch hill (Saurashtra, 222 CE). (C) Proposed Introgressions of pre-Vedic, Indo-Iranic-speaking tribes into the Indian subcontinent. Pr-X = in Khorasan; Pr-Su = Ahar-Banas, Pr-Mh = Maharashtra; Pr-Mg = Magadhi; LPK = Lotli Pai Kaundinya; Map lines delineate study areas and do not necessarily reflect accepted geopolitical boundaries for the periods studied.
Genealogy 10 00014 g005
Figure 6. Paleoclimate Reconstructions. Precipitation patterns reconstructed from the El Niño-Southern Oscillation (ENSO) and Hemispherical Temperature Gradient (Molnar and Rajagopalan 2020). Percent change in annual precipitation relative to present day climatology (i.e., average precipitation during 1951–2015) during four broad epochs covering 5.5 K YBP to modern times; YBP = years before present. Map lines delineate study areas and do not necessarily reflect accepted geopolitical boundaries for the periods studied.
Figure 6. Paleoclimate Reconstructions. Precipitation patterns reconstructed from the El Niño-Southern Oscillation (ENSO) and Hemispherical Temperature Gradient (Molnar and Rajagopalan 2020). Percent change in annual precipitation relative to present day climatology (i.e., average precipitation during 1951–2015) during four broad epochs covering 5.5 K YBP to modern times; YBP = years before present. Map lines delineate study areas and do not necessarily reflect accepted geopolitical boundaries for the periods studied.
Genealogy 10 00014 g006
Table 1. Geographic distribution of Y-haplogroups.
Table 1. Geographic distribution of Y-haplogroups.
A. Geographic distribution of selected Y-haplogroups (Database: FT-DNA).
CLADES
Geographic RegionnR-Z283R-Z94C-M216
Scandinavian Countries389100.00.00.0
Baltic Countries58197.81.70.5
Danubian Countries13998.60.70.7
Ukraine-Belarus (Slavic)19668.930.11.0
Other W. European (not Russia)67489.63.61.9
Western Russia (European)8991.09.00.0
Volga-Uralic (C. Russian)15596.11.32.6
N. Andronovo Region (E. Russian)1681.30.018.8
S. Andronovo Region (C. Asian)1303.10.896.2
East Asian170.05.994.1
Northern Caucasus (Russian Fed)3122.671.06.5
Anatolia-Transcaucasia2321.769.68.7
Khorasan and South Asian370.075.724.3
Arabian Peninsula420.095.24.8
Totals (n)25192092212182
B. Distribution of R-Y7 branch among R-Y3 Y-haplogroups (Database: FT-DNA).
R-Z94 SUBCLADES
Geographic RegionnNon-Y3Y3 (Pct. R-Y7)
Transcaucasia, Anatolia, Mesopotamia4888.511.5 (66.7)
Gulf States4815.884.2 (90.0)
Khorasan (UZB, TKM, IRN, AFG), South Asia5518.481.6 (77.8)
Table 2. Single Nucleotide Polymorphisms (SNPs) in sequenced R-Y7 LPK genomes: Calculated ages of branch points (YBP); * R-Y16494 mutation is not present in LPK patrilineage, but marks ante quem/post quem branch point of LPK from modal R-Y7.
Table 2. Single Nucleotide Polymorphisms (SNPs) in sequenced R-Y7 LPK genomes: Calculated ages of branch points (YBP); * R-Y16494 mutation is not present in LPK patrilineage, but marks ante quem/post quem branch point of LPK from modal R-Y7.
Branch
Point
Genome IDRef.G#SNP IDHg38 CoordAncDerY-Full Est. Age (YBP)Calculated Age (YBP)
16MultipleY-Full>10R-Z94Y:18881562TC4885 ± 9035093
15MultipleY-Full>10R-Y3Y:15689178TC4885 ± 9035018
14MultipleY-Full>10R-Y2Y:15421488GT4439 ± 9014943
13MultipleY-Full>10R-Y27Y:8571844TA4360 ± 9074794
12MultipleY-Full>10L657Y:12039158GA4360 ± 9074644
11MultipleY-Full>10R-M605Y:6942895AG4217 ± 9154569
10MultipleY-Full>10R-Y9Y:17271472CA3977 ± 8184494
9MultipleY-Full>10R-Y7Y:17363983AC3977 ± 8184344
8MultipleY-Full>10R-Y30Y:15971354CT3751 ± 7614194
7MultipleY-Full>10R-Y29Y: 3452217TC3751 ± 7614120
6MultipleY-Full>10R-Y944Y:9035232TG3524 ± 7534045
5MultipleY-Full>10R-Y2439Y:14116601noneAins3573 ± 2443895
4MultipleY-Full>10R-Y2428Y:8238039AG3355 ± 3793820
3MultipleY-Full6FT64014Y:6500298CTn.a.3745
7R-MF748927Y:20163458GAn.a.3146–3670
7rs201523158Y:20139777CGn.a.3146–3670
7rs201762000Y:20138939AGn.a.3146–3670
7R-MF742206Y:20074941AGn.a.3146–3670
7R-BY3881Y:20073631TAn.a.3146–3670
7rs200387433Y:26656957TAn.a.3146–3670
2MultipleY-Full7R-Y16494*Y:8203613AT3175 ± 2673071
7novel SNPY:11278736GAn.a.1348–2996
7novel SNPY:11275457CTn.a.1348–2996
7novel SNPY:11275456GAn.a.1348–2996
7novel SNPY:11221674CAn.a.1348–2996
7R-A24217Y:4209699CTn.a.1348–2996
7R-A24220Y:9285164TCn.a.1348–2996
7R-A24221Y:9302034TCn.a.1348–2996
7R-A24223Y:13415761CGn.a.1348–2996
7R-A24224Y:13418827AGn.a.1348–2996
7R-A24227Y:14215293CAn.a.1348–2996
7R-A24228Y:15592582TCn.a.1348–2996
7R-A24229Y:15635147TAn.a.1348–2996
7R-A24233Y:19179106GTn.a.1348–2996
7R-A24236Y:19539105TCn.a.1348–2996
7R-A24240Y:22252471AGn.a.1348–2996
7R-A24230Y:16700505GAn.a.1348–2996
7R-A24230Y:16700505GAn.a.1348–2996
1YSEQ19576LKR-VgU.2 * n.a.1348
6R-A24218Y:6840776TGn.a.749–1348
6R-A24225Y:13468068TCn.a.749–1348
6R-A24226Y:13534001CAn.a.749–1348
6R-A24232Y:17056333GAn.a.749–1348
6R-A24238Y:19951374TGn.a.749–1348
0YSEQ11320LKR-Vg12.1 * n.a.749
0YSEQ19273LKR-Vg15.1 * n.a.749
4R-A24244Y:7738920CAn.a.375–749
4R-A24247Y:14828363AGn.a.375–749
4R-A24249Y:21638606GTn.a.375–749
−1YSEQ19574LKR-Vg-EM1 n.a.375
3R-A24246Y:12829319GTn.a.225–375
−2YSEQ19577LKR-Vg8.3 * n.a.225
2R-A24243Y:6829098AGn.a.75–225
2R-A24245Y:8360939TCn.a.75–225
2R-A24248Y:15911707TCn.a.75–225
−3YSEQ19578LKR-Vg8.2 *1 n.a.75
(2000 CE)YSEQ5926LKR-Vg8.1 *1 n.a.0
n.a. = not available from Y-Full.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Mascarenhas, D.D.; Rajagapolan, B.; Fox, J.W.; Johnson, R.J. Enhanced Integration of Multi-Disciplinary Inputs into a Narrative of an Ancient Migration, Based on Greater Chronological Precision Provided by a Novel Y-DNA Clock and Phylogenetic Branching. Genealogy 2026, 10, 14. https://doi.org/10.3390/genealogy10010014

AMA Style

Mascarenhas DD, Rajagapolan B, Fox JW, Johnson RJ. Enhanced Integration of Multi-Disciplinary Inputs into a Narrative of an Ancient Migration, Based on Greater Chronological Precision Provided by a Novel Y-DNA Clock and Phylogenetic Branching. Genealogy. 2026; 10(1):14. https://doi.org/10.3390/genealogy10010014

Chicago/Turabian Style

Mascarenhas, Desmond D., Balaji Rajagapolan, John W. Fox, and Richard J. Johnson. 2026. "Enhanced Integration of Multi-Disciplinary Inputs into a Narrative of an Ancient Migration, Based on Greater Chronological Precision Provided by a Novel Y-DNA Clock and Phylogenetic Branching" Genealogy 10, no. 1: 14. https://doi.org/10.3390/genealogy10010014

APA Style

Mascarenhas, D. D., Rajagapolan, B., Fox, J. W., & Johnson, R. J. (2026). Enhanced Integration of Multi-Disciplinary Inputs into a Narrative of an Ancient Migration, Based on Greater Chronological Precision Provided by a Novel Y-DNA Clock and Phylogenetic Branching. Genealogy, 10(1), 14. https://doi.org/10.3390/genealogy10010014

Article Metrics

Back to TopTop