Evidence page · genetics and linguistics
A community of five million people descends from about a hundred.
The Vysya of Andhra Pradesh and Telangana carry one of the most extreme founder events recorded in any human population — more extreme than the Ashkenazi Jewish bottleneck. They are also a middle-ranking merchant caste, not an oppressed one, and that is the part worth pausing on. This page sets out when South Asian communities stopped exchanging marriage partners, what the genetics can and cannot say about why, and what the same period did to the Dravidian languages.
The Vysya founder event.
the numbers
~100
individuals in the founding group
3,000–2,000
years ago, the bottleneck
~5M
descendants today
100×
rate of BChE deficiency vs other groups
For an estimated three thousand years the Vysya have experienced negligible gene flow from neighbouring groups. The consequence is medical: a roughly hundredfold elevated rate of butyrylcholinesterase deficiency, a metabolic disorder causing severe reactions to common anaesthetics — a recessive variant that drifted to high frequency because there was too little variation for it to be diluted. For comparison, the Ashkenazi Jewish founding population is estimated at around 350 individuals. The Vysya bottleneck is tighter.
But the Vysya are not an oppressed caste.
and that is the finding, not an aside
A survey of 263 South Asian groups found 81 with strong founder events, including 14 with census populations above a million. The named list runs the length of the ranking: Gujjar, Baniya, Pattapu Kapu, Vadde, Yadav, Kshatriya Agnikula, Naga, Kumhar, Reddy, Kallar, Brahmin Manipuri, Arunthathiyar and Vysya — Brahmin to Dalit, merchant to cultivator to scavenger.
Endogamy is not something done to the bottom of the order. It is the architecture of the whole thing. The Vysya sit in the middle, and their genetic isolation is as complete as anyone's. Groups at the top enforced it on themselves as rigorously as it was enforced on those below — which is what a system of mutually sealed compartments looks like from the inside.
When the mixing stopped.
a window that opened and closed
The Vysya are early, but they are not exceptional in kind. The wider pattern is a demographic transformation with a beginning and an end:
| Period | What was happening |
|---|---|
| c. 4,200–1,900 BP | Major mixture between populations across India. Well after the establishment of agriculture. Groups with unmixed ancestry were still living in India during this window, and mixture reached even relatively isolated groups such as the Palliyar and Bhil. |
| c. 3,500–2,000 BP | Founder events cluster. Communities begin sealing. The Vysya bottleneck falls at the early end. |
| after c. 2,000 BP | Mixture becomes rare. Today every group in mainland India is admixed — but the admixture is old, and almost nothing has been added since. |
Why it stopped: what the genetics does and does not say.
the limit of the method
A bottleneck is a measurement of who had children with whom. It carries no information about the rules, beliefs or coercion that produced the pattern. The genetic literature is careful about this and so is this page.
What can be said is that the closure coincides with the period in which the vocabulary of exclusion multiplies in the Sanskrit textual record. Counted across 18.8 million words, śūdra rises from 1.8 per 100,000 words in the Saṃhitās to 221.5 in the law books; caṇḍāla from 0.1 to 35.4; the word aspṛśya, "not-to-be-touched," is absent from the Saṃhitās and Brāhmaṇas entirely and enters in the Sūtra literature. Two independent records — one biological, one textual — show a system hardening in the same window. That is a correlation of considerable strength and it is not a demonstration of cause. See the vocabulary curve.
The Indus question.
what the dates do and do not settle
It is tempting: a merchant community, a civilisation known for long-distance trade. The Indus urban phase ends around 1900 BCE — about 3,900 years ago — and the Vysya founder event dates to 3,000–2,000 years ago. That gap alone settles nothing: a founder event dates the end of mixing, not a population's origin, and nine hundred years is thirty generations, ample for a slow migration with admixture along the way. The objection that would settle it is about ancestry composition relative to neighbouring communities — and as set out below, that comparison has not been published.
Worse for the idea, Vaiśya is a varṇa term from the Sanskrit fourfold scheme — a slot in a classification, not the name of a lineage. Communities acquire such labels; the label is not evidence of descent. What the Vysya do carry, like every mainland Indian population, is Indus-derived ancestry through the Indus Periphery Cline — but so does everyone, and it says nothing about occupation.
Testing the migration hypothesis properly.
the timing objection is weak; the ancestry test is not
A reasonable reconstruction runs: the Indus cities decline around 1800 BCE; trading families move gradually east and south over many generations, mixing as they go; they reach the Andhra coast, establish themselves, and only then seal into endogamy around 1000 BCE. The timing objection to this is weak. Eight hundred years is more than thirty generations, and a founder event dates the end of mixing, not the origin of a population. Nothing in the chronology forbids a slow migration.
But the hypothesis makes a prediction that can be checked. If a community descends from northwestern migrants who arrived and then sealed, it should carry more northwestern ancestry than its neighbours — a higher West Eurasian or Iranian-related fraction, preserved by the very isolation that produced the bottleneck.
The comparison the hypothesis needs is narrow and specific: do the Vysya carry more northwestern-related ancestry than their immediate Andhra neighbours — Reddy, Kapu, Kamma, Velama, the local tribal populations? That is a within-region question, and it is the only version of the comparison that means anything.
The structure analyses available compare across a global or all-India panel — Pathans, Punjabis, Gujaratis, Kashmiri Pandits, Sri Lankan Tamils and Telugu speakers together. In that frame the West Eurasian component is highest in the northwestern populations, and Telugu populations including the Vysya sit toward the southern end. That tells you the Vysya are a South Indian population, which nobody disputes. It does not tell you whether they differ from the Andhra communities beside them, because those communities were not the comparison set.
An earlier version of this section treated that as a test of the migration hypothesis. It is not one — comparing a Telugu community against Pathans establishes nothing about its origin relative to other Telugu communities. The within-Andhra comparison is a straightforward analysis on data that largely exists. It has not been published.
What the regional comparison would have to show
sort by date or by founder size
| Observation | Reading |
|---|---|
| Vysya carry more northwestern-related ancestry than Reddy, Kapu and Kamma | Consistent with a founding group of migrant origin whose isolation preserved the signal |
| Vysya carry the same as their neighbours | Consistent with a founding group drawn from the local population |
| Vysya carry less | Argues against migrant origin |
The founder-event table is the page. Sorting it by date against size is the comparison the prose makes in words.
None of these has been reported. What is reported at the community level is descriptive rather than quantified — Reddy and Kamma cluster together with a broadly similar ancestry balance; Kapu shows more internal regional variation than tightly endogamous communities like Kamma or Velama; coastal, Rayalaseema and Telangana profiles differ, with inland and highland populations sometimes carrying higher southern-component proportions. That is a landscape with real structure in it, and nobody has published the numbers.
So the migration hypothesis is currently unfalsified rather than supported. It makes a clear prediction, the data to test it is largely collected, and the test has not been carried out. That is a different situation from a hypothesis that has been checked and failed, and this page does not conflate the two.
Other Telugu groups that sealed.
the Vysya are not alone in the region
The founder-event survey names several communities from the Telugu-speaking region, which bears directly on whether this was one community's history or a regional process:
| Community | Region | Traditional position |
|---|---|---|
| Vysya | Telangana | Merchant |
| Reddy | Telangana | Landholding; village headmen and tax collectors, in kingdom administration from the seventh century |
| Pattapu Kapu | Andhra Pradesh | Landholding agriculturist |
| Kshatriya Agnikula | Andhra Pradesh | Warrior-status |
| Vadde | Andhra Pradesh | Stone-working and earth-working |
Five communities from one linguistic region, spanning merchant, landholding, warrior and labouring positions, all showing strong founder events. This is not the isolation of one unusual group. It is a regional pattern in which many communities sealed, and the seals did not correlate with rank.
Did these communities seal at the same time, and did some seal and later re-open? Per-community founder dates and ancestry profiles for these groups are not published at the resolution the question needs. The survey reports drift strength and differentiation; individual bottleneck dates for most groups are not given.
What is reported is that regional variation exists — coastal Andhra communities show different profiles from Rayalaseema and Telangana ones, with inland and highland populations sometimes carrying higher southern-component proportions. That is consistent with different sealing dates and different degrees of subsequent exchange, and it is not evidence for any particular sequence. The sampling to resolve it would be straightforward and has not been done.
Who imposed the languages.
a claim worth stating precisely
Languages spread with power, not by consensus, and the population of any modern linguistic state was never uniformly composed of that language's speakers. The administrative record supports the general shape of this. The Telugu term Reddi — earlier Raddi, Rattoḍi, Rattakuḍi, connected to Rāṣṭrakūṭa — denoted village headmen responsible for organising cultivation and collecting taxes, and from the seventh century members of these families held important administrative posts in kingdoms. The Kapu were landholders who also served as military commanders under Vijayanagara.
Supported: the communities that carried regional administration and land control in the Telugu country were non-Brahmin landholding and commercial groups. Administration is conducted in a language, and linguistic standardisation runs through revenue, court and temple record-keeping — their institutions. These groups are a plausible vector for how one variety became the standard and spread across a population that did not uniformly speak it. That is a real mechanism and the administrative record is consistent with it.
The distinction that matters: spreading and standardising a language is not the same as originating it, and the two get run together. Dravidian is millennia older than any state formation in the Telugu country, so no community can have been "first speakers." But being the group that made a language official, prestigious and administratively necessary is a large historical fact in its own right — and it is the thing that determines which varieties survive and which are absorbed.
Not established: which specific communities did this in which region, or when. The Reddi–Rāṣṭrakūṭa etymology and the seventh-century administrative posts are documented; a general account of Telugu standardisation naming its agents is not something this page has sourced, and it would need to be argued from inscriptions rather than inferred from caste position.
The Dravidian family and its branches.
Kolipakam et al. 2018
A Bayesian phylogenetic analysis of first-hand lexical data from native speakers dates the Dravidian family to approximately 4,500 years old — about 80 varieties, 220 million speakers. And its internal structure is not what the geography suggests.
| Branch | Languages |
|---|---|
| South I | Tamil, Malayalam, Kannada, Tulu, Koḍagu, Toda, Kota, Badga, Beṭṭa Kurumba, Yeruva |
| South II · Central · North | Telugu, Gondi, Koya, Kuwi, Kolami, Ollari Gadba, Parji, Brahui, Kuṛux, Malto |
The best-supported model returns a main two-way split with Tamil, Malayalam and Kannada on one side and Telugu on the other — grouped with Gondi and Koya, and on the same side of the primary division as Brahui in Balochistan and Kuṛux in Jharkhand.
So the intuition is right in its main part: Malayalam and Tamil are the shallowest split; Kannada is deeper but within the same branch; Telugu is not in that branch at all. Whether Telugu's separation is the family's deepest division depends on which constraints the model is given — see below.
The authors' own caution matters here. They report "considerable uncertainty with regard to the relationships between the main branches," and a differently constrained analysis returns a different higher-order topology. The four branches are secure; the order in which they separated is not.
How firm is 4,500 years?
the interval the headline dropped
"Approximately 4,500 years old" is the point estimate that reached every press release. The paper reports a root with a mean of 4,650 years and a median of 4,433, and a 95% highest posterior density interval of roughly 3,000–6,500 years — consistent across the models tested except the stochastic Dollo model, which performs markedly worse than the others. That is an interval three and a half thousand years wide. And what the paper says about the number is: “we cannot exclude the possibility that the root of the Dravidian language family is significantly older than 4500 years.”
The authors say so plainly: the uncertainty on the root age is large. The number that travelled was the midpoint.
What that interval contains is the whole of the question:
| If the root is… | Then Dravidian diversified… |
|---|---|
| 6,500 BP (4500 BCE) | Long before the Indus cities. The family was already splitting a millennium and a half before the Mature Harappan phase begins. |
| 4,500 BP (2500 BCE) | During the Mature Indus phase. The primary split is contemporary with the cities at their height. |
| 3,000 BP (1000 BCE) | Well after the Indus decline. The family fragments in the post-urban period, not before it. |
All three sit inside the published interval. The evidence as it stands does not choose between "Dravidian split during the Indus period" and "Dravidian split because of its collapse." Anyone asserting either with confidence is reading the midpoint as though it were the finding.
The estimate that disagrees, and by a factor of three
The 4,500-year figure is not the only published estimate, and the paper says so itself: “Contrary to other work by Pagel et al., our findings suggest this younger age rather than their Proto-Dravidian estimate of around 13 000 years ago. Future work should investigate the disparity in clock rates here.” Two further estimates sit between them — Fuller places the differentiation of the Dravidian languages around 6,000 years ago on reconstructed crop vocabulary, and Southworth puts the diversification of the North, Central and South branches at 4,500–4,000 years ago, tied to the Southern Neolithic. A field whose published estimates run from about 4,000 to about 13,000 years has not settled this, and a page that reports only the youngest of them is not reporting the state of the question.
And the calibration has a known structural bias
The dates are anchored on first written attestations — Tamil's earliest lithic inscription of about 254 BCE is used to constrain the South I group to be at least 2,250 years old. Methodologists describe this as sound in principle: a protolanguage is necessarily older than its attested descendant. But it has a consequence.
The calibration itself is confirmed against the paper: the authors “included a calibration so that this group could not be younger than 2250 years, because Tamil was recorded first in 254 BCE”. A constraint of that form sets a floor under the inferred age of that group — it forces the estimate to be at least as old as the attestation, and cannot pull it younger.
The deeper issue is that writing is not speech. Tamil was inscribed in 254 BCE; it was spoken for an unknown period before anyone cut it into stone. Calibrating on the first surviving inscription measures when a literate state began keeping records, which is a fact about administration, not about language age.
The branching order is also unstable — a correction
The best-supported model returns a main two-way split with Telugu (South II) opposite Tamil, Malayalam and Kannada (South I). But under a different constraint regime — one that adds a ninth calibration constraining South I and South II to be monophyletic — the analysis recovers the traditional topology instead: in the authors’ words, “the North group splits off first, and South I + South II diverges last.” Verified against the paper.
Under that reading Telugu's separation from Tamil is the most recent major division rather than the deepest. The four branches are secure. The order in which they separated is not, and the authors state as much: they find “considerable uncertainty with regard to the relationships between the main branches”, and report that “the current dataset has low resolution for the higher-order subgrouping, despite recovering the four subgroups with reasonable confidence.” Both verified against the paper.
So the honest position on Telugu is narrower than the striking version. Telugu is certainly not within South I — it does not belong with Tamil, Malayalam and Kannada, and the intuition that it is a different order of separation is correct. Whether it is the family's deepest division or its shallowest major one is not resolved by this data.
What the eastward hypothesis would need
A scenario in which an early-branching group moved east and later south — through the Gangetic plain and down into the peninsula — is not excluded, and the modern distribution of Kuṛux and Malto in Jharkhand and Bengal is exactly the kind of residue such a movement would leave. But the linguistic dating cannot test it, because a root age does not locate anything in space.
What would test it is the same thing missing everywhere else on this question: ancient DNA and dated archaeology from the eastern and central corridors, at the resolution that exists for the northwest. That survey has not been done, and until it is, the eastward route and the peninsular route remain equally unfalsified.
A test of that tree, from the dictionary.
and it disagrees, for an instructive reason
The claim can be checked independently. Counting shared etymological root numbers across 26 Dravidian languages and computing overlap gives this:
| Mean overlap | South I | South II | Central | North |
|---|---|---|---|---|
| South I | 0.267 | 0.154 | 0.129 | 0.099 |
| South II | 0.154 | 0.252 | 0.187 | 0.114 |
| Central | 0.129 | 0.187 | 0.322 | 0.125 |
| North | 0.099 | 0.114 | 0.125 | 0.238 |
At branch level it works: every branch is most similar to itself, and similarity falls off with distance. But at language level it breaks.
| Telugu's nearest neighbours by shared roots | Overlap | Branch |
|---|---|---|
| Kannaḍa | 0.547 | South I — wrong branch |
| Tamil | 0.451 | South I — wrong branch |
| Tulu | 0.422 | South I — wrong branch |
| Malayalam | 0.376 | South I — wrong branch |
| Koṇḍa | 0.216 | South II — its own branch |
Two artefacts, both worth naming. First, documentation: Telugu, Kannada, Tamil, Malayalam and Tulu are the large literary languages with 2,000–2,800 recorded roots each, while Koṇḍa, Pengo and Manḍa have 400–900. A small recorded inventory depresses overlap automatically, regardless of relatedness. Second, contact: Telugu and Kannada have been neighbours in intense exchange for more than a millennium, and borrowed words look exactly like inherited ones in a raw count.
Shared vocabulary measures contact and documentation effort as much as it measures descent. The phylogenetic method uses cognate coding and historical calibration precisely to get past that — which is why its answer is the one to trust here, and why this cross-check is published as a disagreement rather than buried.
Iranian-related ancestry.
what is established, and the part that is not
The component usually called "Iranian farmer-related" is the largest single element in the ancestry of the Indus world and of everyone descended from it:
- The Indus Periphery Cline individuals carry ~45–82% Iranian farmer-related and ~11–50% AASI ancestry, with negligible Anatolian farmer admixture.
- ASI — the southern ancestral component — is a minimum of ~55% Indus Periphery Cline ancestry. ANI is ~70%.
- Every mainland Indian population carries it, north and south. It predates the steppe arrival by millennia and is the shared inheritance underneath the north–south gradient.
- The lineage is deep: it shares a common ancestor with ancient Iranians more than 12,000 years ago. It is not evidence of recent migration from Iran.
Which specific communities in each part of South India carry the highest Iranian-related proportions is not established at that resolution in the published literature. What exists is coarser: a study of West Maharashtra found the Brahmin caste there carries higher Ancient Iranian and Steppe contributions than the Kunbi Maratha caste, and the general pattern across the subcontinent is that these proportions correlate with traditional rank.
A community-by-community, region-by-region table for South India would be an invention, and this page will not produce one. The sampling that would support it has not been published. That is a gap in the record, and naming it is more useful than filling it.
Founder events and ancestry proportions are genome-wide. Haplogroup frequencies are not: a Y-haplogroup traces one unbroken father-to-son line, one ancestor out of more than a thousand in the same generation. A group can carry a lineage at high frequency and hold little of the ancestry that lineage arrived with, or the reverse. Nothing on this page converts one into the other, and where both appear they are labelled.
The honest limits.
- Sources: the Vysya founder size, date and BChE rate, and the 81-of-263 founder-event survey, from the published founder-event literature — snippet level; the survey is not named on this page, and is recorded as an unidentified source. The platform’s own endogamy-clock page reports the Vysya bottleneck with error bars (144 ± 27 generations, roughly ± 780 years); the point figures used here should be read with that spread; the 4,200–1,900 BP mixture window from Moorjani et al. 2013 (AJHG); Dravidian family age and branch structure from Kolipakam et al. 2018 (Royal Society Open Science 5:171504, doi:10.1098/rsos.171504) — abstract, indexed full-text excerpts and citation statements read here; the topology result and the authors’ statement of uncertainty are verified, the 95% interval is not and has been withdrawn; Indus Periphery proportions from Narasimhan et al. 2019 (Science) — snippet level. The overlap matrix is computed here from shared etymological root numbers;
dedr_roots.jsonships with the site. - The 4,500-year Dravidian date is a point estimate — mean 4,650, median 4,433 — inside a 95% interval of roughly 3,000–6,500 years, spanning the pre-Indus, Mature Indus and post-Indus periods alike, and the authors state they cannot exclude a significantly older root. Published estimates elsewhere run from about 4,000 years (Southworth) to about 13,000 (Pagel et al.). Calibration is anchored on first written attestation, which sets a floor rather than measuring language age. The interval is corroborated at secondary level; the paper has not been opened in full.
- The branch order is unresolved. The best-supported model puts Telugu opposite Tamil at the primary split; a differently constrained model recovers the traditional topology. The four branches are secure; their sequence is not.
- Nothing here connects the Vysya to Indus-era traders. The chronology does not exclude it; the test that would settle it needs a within-Andhra ancestry comparison that has not been published.
- Ancestry components are artefacts of the model and reference panels chosen, not natural categories, and are used only for relative comparison between populations analysed together.
- Community-level Iranian-ancestry figures for South India are not published at that resolution and none are given.
The claims, their sources, and what is still open
This page carries a reversed correction — a withdrawn interval that was later restored. The record of that belongs on the page.