Skip to content
Mela Keela

Evidence page · genetics and linguistics

A community of five million people descends from about a hundred.

The Vysya of Andhra Pradesh and Telangana carry one of the most extreme founder events recorded in any human population — more extreme than the Ashkenazi Jewish bottleneck. They are also a middle-ranking merchant caste, not an oppressed one, and that is the part worth pausing on. This page sets out when South Asian communities stopped exchanging marriage partners, what the genetics can and cannot say about why, and what the same period did to the Dravidian languages.

The Vysya founder event.

the numbers

~100

individuals in the founding group

3,000–2,000

years ago, the bottleneck

~5M

descendants today

100×

rate of BChE deficiency vs other groups

For an estimated three thousand years the Vysya have experienced negligible gene flow from neighbouring groups. The consequence is medical: a roughly hundredfold elevated rate of butyrylcholinesterase deficiency, a metabolic disorder causing severe reactions to common anaesthetics — a recessive variant that drifted to high frequency because there was too little variation for it to be diluted. For comparison, the Ashkenazi Jewish founding population is estimated at around 350 individuals. The Vysya bottleneck is tighter.

But the Vysya are not an oppressed caste.

and that is the finding, not an aside

Endogamy runs across the hierarchy, not down it

A survey of 263 South Asian groups found 81 with strong founder events, including 14 with census populations above a million. The named list runs the length of the ranking: Gujjar, Baniya, Pattapu Kapu, Vadde, Yadav, Kshatriya Agnikula, Naga, Kumhar, Reddy, Kallar, Brahmin Manipuri, Arunthathiyar and Vysya — Brahmin to Dalit, merchant to cultivator to scavenger.

Endogamy is not something done to the bottom of the order. It is the architecture of the whole thing. The Vysya sit in the middle, and their genetic isolation is as complete as anyone's. Groups at the top enforced it on themselves as rigorously as it was enforced on those below — which is what a system of mutually sealed compartments looks like from the inside.

When the mixing stopped.

a window that opened and closed

The Vysya are early, but they are not exceptional in kind. The wider pattern is a demographic transformation with a beginning and an end:

PeriodWhat was happening
c. 4,200–1,900 BPMajor mixture between populations across India. Well after the establishment of agriculture. Groups with unmixed ancestry were still living in India during this window, and mixture reached even relatively isolated groups such as the Palliyar and Bhil.
c. 3,500–2,000 BPFounder events cluster. Communities begin sealing. The Vysya bottleneck falls at the early end.
after c. 2,000 BPMixture becomes rare. Today every group in mainland India is admixed — but the admixture is old, and almost nothing has been added since.
India shifted from a place where mixing between groups was ordinary to a place where it was rare — and the shift has a date.

Why it stopped: what the genetics does and does not say.

the limit of the method

Genetics dates the event. It does not explain it.

A bottleneck is a measurement of who had children with whom. It carries no information about the rules, beliefs or coercion that produced the pattern. The genetic literature is careful about this and so is this page.

What can be said is that the closure coincides with the period in which the vocabulary of exclusion multiplies in the Sanskrit textual record. Counted across 18.8 million words, śūdra rises from 1.8 per 100,000 words in the Saṃhitās to 221.5 in the law books; caṇḍāla from 0.1 to 35.4; the word aspṛśya, "not-to-be-touched," is absent from the Saṃhitās and Brāhmaṇas entirely and enters in the Sūtra literature. Two independent records — one biological, one textual — show a system hardening in the same window. That is a correlation of considerable strength and it is not a demonstration of cause. See the vocabulary curve.

The Indus question.

what the dates do and do not settle

Direct descent from Indus traders is not supported

It is tempting: a merchant community, a civilisation known for long-distance trade. The Indus urban phase ends around 1900 BCE — about 3,900 years ago — and the Vysya founder event dates to 3,000–2,000 years ago. That gap alone settles nothing: a founder event dates the end of mixing, not a population's origin, and nine hundred years is thirty generations, ample for a slow migration with admixture along the way. The objection that would settle it is about ancestry composition relative to neighbouring communities — and as set out below, that comparison has not been published.

Worse for the idea, Vaiśya is a varṇa term from the Sanskrit fourfold scheme — a slot in a classification, not the name of a lineage. Communities acquire such labels; the label is not evidence of descent. What the Vysya do carry, like every mainland Indian population, is Indus-derived ancestry through the Indus Periphery Cline — but so does everyone, and it says nothing about occupation.

Testing the migration hypothesis properly.

the timing objection is weak; the ancestry test is not

A reasonable reconstruction runs: the Indus cities decline around 1800 BCE; trading families move gradually east and south over many generations, mixing as they go; they reach the Andhra coast, establish themselves, and only then seal into endogamy around 1000 BCE. The timing objection to this is weak. Eight hundred years is more than thirty generations, and a founder event dates the end of mixing, not the origin of a population. Nothing in the chronology forbids a slow migration.

But the hypothesis makes a prediction that can be checked. If a community descends from northwestern migrants who arrived and then sealed, it should carry more northwestern ancestry than its neighbours — a higher West Eurasian or Iranian-related fraction, preserved by the very isolation that produced the bottleneck.

The test is available in principle. It has not been run.

The comparison the hypothesis needs is narrow and specific: do the Vysya carry more northwestern-related ancestry than their immediate Andhra neighbours — Reddy, Kapu, Kamma, Velama, the local tribal populations? That is a within-region question, and it is the only version of the comparison that means anything.

The structure analyses available compare across a global or all-India panel — Pathans, Punjabis, Gujaratis, Kashmiri Pandits, Sri Lankan Tamils and Telugu speakers together. In that frame the West Eurasian component is highest in the northwestern populations, and Telugu populations including the Vysya sit toward the southern end. That tells you the Vysya are a South Indian population, which nobody disputes. It does not tell you whether they differ from the Andhra communities beside them, because those communities were not the comparison set.

An earlier version of this section treated that as a test of the migration hypothesis. It is not one — comparing a Telugu community against Pathans establishes nothing about its origin relative to other Telugu communities. The within-Andhra comparison is a straightforward analysis on data that largely exists. It has not been published.

What the regional comparison would have to show

Inspect

sort by date or by founder size

ObservationReading
Vysya carry more northwestern-related ancestry than Reddy, Kapu and KammaConsistent with a founding group of migrant origin whose isolation preserved the signal
Vysya carry the same as their neighboursConsistent with a founding group drawn from the local population
Vysya carry lessArgues against migrant origin

The founder-event table is the page. Sorting it by date against size is the comparison the prose makes in words.

None of these has been reported. What is reported at the community level is descriptive rather than quantified — Reddy and Kamma cluster together with a broadly similar ancestry balance; Kapu shows more internal regional variation than tightly endogamous communities like Kamma or Velama; coastal, Rayalaseema and Telangana profiles differ, with inland and highland populations sometimes carrying higher southern-component proportions. That is a landscape with real structure in it, and nobody has published the numbers.

So the migration hypothesis is currently unfalsified rather than supported. It makes a clear prediction, the data to test it is largely collected, and the test has not been carried out. That is a different situation from a hypothesis that has been checked and failed, and this page does not conflate the two.

Other Telugu groups that sealed.

the Vysya are not alone in the region

The founder-event survey names several communities from the Telugu-speaking region, which bears directly on whether this was one community's history or a regional process:

CommunityRegionTraditional position
VysyaTelanganaMerchant
ReddyTelanganaLandholding; village headmen and tax collectors, in kingdom administration from the seventh century
Pattapu KapuAndhra PradeshLandholding agriculturist
Kshatriya AgnikulaAndhra PradeshWarrior-status
VaddeAndhra PradeshStone-working and earth-working

Five communities from one linguistic region, spanning merchant, landholding, warrior and labouring positions, all showing strong founder events. This is not the isolation of one unusual group. It is a regional pattern in which many communities sealed, and the seals did not correlate with rank.

The question this raises, and cannot yet answer

Did these communities seal at the same time, and did some seal and later re-open? Per-community founder dates and ancestry profiles for these groups are not published at the resolution the question needs. The survey reports drift strength and differentiation; individual bottleneck dates for most groups are not given.

What is reported is that regional variation exists — coastal Andhra communities show different profiles from Rayalaseema and Telangana ones, with inland and highland populations sometimes carrying higher southern-component proportions. That is consistent with different sealing dates and different degrees of subsequent exchange, and it is not evidence for any particular sequence. The sampling to resolve it would be straightforward and has not been done.

Who imposed the languages.

a claim worth stating precisely

Languages spread with power, not by consensus, and the population of any modern linguistic state was never uniformly composed of that language's speakers. The administrative record supports the general shape of this. The Telugu term Reddi — earlier Raddi, Rattoḍi, Rattakuḍi, connected to Rāṣṭrakūṭa — denoted village headmen responsible for organising cultivation and collecting taxes, and from the seventh century members of these families held important administrative posts in kingdoms. The Kapu were landholders who also served as military commanders under Vijayanagara.

What follows, and what does not

Supported: the communities that carried regional administration and land control in the Telugu country were non-Brahmin landholding and commercial groups. Administration is conducted in a language, and linguistic standardisation runs through revenue, court and temple record-keeping — their institutions. These groups are a plausible vector for how one variety became the standard and spread across a population that did not uniformly speak it. That is a real mechanism and the administrative record is consistent with it.

The distinction that matters: spreading and standardising a language is not the same as originating it, and the two get run together. Dravidian is millennia older than any state formation in the Telugu country, so no community can have been "first speakers." But being the group that made a language official, prestigious and administratively necessary is a large historical fact in its own right — and it is the thing that determines which varieties survive and which are absorbed.

Not established: which specific communities did this in which region, or when. The Reddi–Rāṣṭrakūṭa etymology and the seventh-century administrative posts are documented; a general account of Telugu standardisation naming its agents is not something this page has sourced, and it would need to be argued from inscriptions rather than inferred from caste position.

The Dravidian family and its branches.

Kolipakam et al. 2018

A Bayesian phylogenetic analysis of first-hand lexical data from native speakers dates the Dravidian family to approximately 4,500 years old — about 80 varieties, 220 million speakers. And its internal structure is not what the geography suggests.

BranchLanguages
South ITamil, Malayalam, Kannada, Tulu, Koḍagu, Toda, Kota, Badga, Beṭṭa Kurumba, Yeruva
South II · Central · NorthTelugu, Gondi, Koya, Kuwi, Kolami, Ollari Gadba, Parji, Brahui, Kuṛux, Malto
Telugu is not in the Tamil–Kannada branch

The best-supported model returns a main two-way split with Tamil, Malayalam and Kannada on one side and Telugu on the other — grouped with Gondi and Koya, and on the same side of the primary division as Brahui in Balochistan and Kuṛux in Jharkhand.

So the intuition is right in its main part: Malayalam and Tamil are the shallowest split; Kannada is deeper but within the same branch; Telugu is not in that branch at all. Whether Telugu's separation is the family's deepest division depends on which constraints the model is given — see below.

The authors' own caution matters here. They report "considerable uncertainty with regard to the relationships between the main branches," and a differently constrained analysis returns a different higher-order topology. The four branches are secure; the order in which they separated is not.

How firm is 4,500 years?

the interval the headline dropped

The 95% interval runs from about 3,000 to 6,500 years ago — and a much older root cannot be excluded

"Approximately 4,500 years old" is the point estimate that reached every press release. The paper reports a root with a mean of 4,650 years and a median of 4,433, and a 95% highest posterior density interval of roughly 3,000–6,500 years — consistent across the models tested except the stochastic Dollo model, which performs markedly worse than the others. That is an interval three and a half thousand years wide. And what the paper says about the number is: “we cannot exclude the possibility that the root of the Dravidian language family is significantly older than 4500 years.”

The authors say so plainly: the uncertainty on the root age is large. The number that travelled was the midpoint.

What that interval contains is the whole of the question:

If the root is…Then Dravidian diversified…
6,500 BP (4500 BCE)Long before the Indus cities. The family was already splitting a millennium and a half before the Mature Harappan phase begins.
4,500 BP (2500 BCE)During the Mature Indus phase. The primary split is contemporary with the cities at their height.
3,000 BP (1000 BCE)Well after the Indus decline. The family fragments in the post-urban period, not before it.

All three sit inside the published interval. The evidence as it stands does not choose between "Dravidian split during the Indus period" and "Dravidian split because of its collapse." Anyone asserting either with confidence is reading the midpoint as though it were the finding.

The estimate that disagrees, and by a factor of three

The 4,500-year figure is not the only published estimate, and the paper says so itself: “Contrary to other work by Pagel et al., our findings suggest this younger age rather than their Proto-Dravidian estimate of around 13 000 years ago. Future work should investigate the disparity in clock rates here.” Two further estimates sit between them — Fuller places the differentiation of the Dravidian languages around 6,000 years ago on reconstructed crop vocabulary, and Southworth puts the diversification of the North, Central and South branches at 4,500–4,000 years ago, tied to the Southern Neolithic. A field whose published estimates run from about 4,000 to about 13,000 years has not settled this, and a page that reports only the youngest of them is not reporting the state of the question.

The competing estimateThe uncertainty is not only an interval around one estimate. A competing estimate of ~13,000 years, which the cited paper itself names and argues against, is the strongest counterevidence to this page’s framing, and it appears above.

And the calibration has a known structural bias

The dates are anchored on first written attestations — Tamil's earliest lithic inscription of about 254 BCE is used to constrain the South I group to be at least 2,250 years old. Methodologists describe this as sound in principle: a protolanguage is necessarily older than its attested descendant. But it has a consequence.

A calibration point sets a floor, and the floor anchors the estimate

The calibration itself is confirmed against the paper: the authors “included a calibration so that this group could not be younger than 2250 years, because Tamil was recorded first in 254 BCE”. A constraint of that form sets a floor under the inferred age of that group — it forces the estimate to be at least as old as the attestation, and cannot pull it younger.

The deeper issue is that writing is not speech. Tamil was inscribed in 254 BCE; it was spoken for an unknown period before anyone cut it into stone. Calibrating on the first surviving inscription measures when a literate state began keeping records, which is a fact about administration, not about language age.

The branching order is also unstable — a correction

Whether Telugu is deeply split from Tamil depends on which constraints are imposed

The best-supported model returns a main two-way split with Telugu (South II) opposite Tamil, Malayalam and Kannada (South I). But under a different constraint regime — one that adds a ninth calibration constraining South I and South II to be monophyletic — the analysis recovers the traditional topology instead: in the authors’ words, “the North group splits off first, and South I + South II diverges last.” Verified against the paper.

Under that reading Telugu's separation from Tamil is the most recent major division rather than the deepest. The four branches are secure. The order in which they separated is not, and the authors state as much: they find “considerable uncertainty with regard to the relationships between the main branches”, and report that “the current dataset has low resolution for the higher-order subgrouping, despite recovering the four subgroups with reasonable confidence.” Both verified against the paper.

So the honest position on Telugu is narrower than the striking version. Telugu is certainly not within South I — it does not belong with Tamil, Malayalam and Kannada, and the intuition that it is a different order of separation is correct. Whether it is the family's deepest division or its shallowest major one is not resolved by this data.

What the eastward hypothesis would need

A scenario in which an early-branching group moved east and later south — through the Gangetic plain and down into the peninsula — is not excluded, and the modern distribution of Kuṛux and Malto in Jharkhand and Bengal is exactly the kind of residue such a movement would leave. But the linguistic dating cannot test it, because a root age does not locate anything in space.

What would test it is the same thing missing everywhere else on this question: ancient DNA and dated archaeology from the eastern and central corridors, at the resolution that exists for the northwest. That survey has not been done, and until it is, the eastward route and the peninsular route remain equally unfalsified.

A test of that tree, from the dictionary.

and it disagrees, for an instructive reason

The claim can be checked independently. Counting shared etymological root numbers across 26 Dravidian languages and computing overlap gives this:

Mean overlapSouth ISouth IICentralNorth
South I0.2670.1540.1290.099
South II0.1540.2520.1870.114
Central0.1290.1870.3220.125
North0.0990.1140.1250.238

At branch level it works: every branch is most similar to itself, and similarity falls off with distance. But at language level it breaks.

Telugu's nearest neighbours by shared rootsOverlapBranch
Kannaḍa0.547South I — wrong branch
Tamil0.451South I — wrong branch
Tulu0.422South I — wrong branch
Malayalam0.376South I — wrong branch
Koṇḍa0.216South II — its own branch
Why raw lexical overlap gets it wrong

Two artefacts, both worth naming. First, documentation: Telugu, Kannada, Tamil, Malayalam and Tulu are the large literary languages with 2,000–2,800 recorded roots each, while Koṇḍa, Pengo and Manḍa have 400–900. A small recorded inventory depresses overlap automatically, regardless of relatedness. Second, contact: Telugu and Kannada have been neighbours in intense exchange for more than a millennium, and borrowed words look exactly like inherited ones in a raw count.

Shared vocabulary measures contact and documentation effort as much as it measures descent. The phylogenetic method uses cognate coding and historical calibration precisely to get past that — which is why its answer is the one to trust here, and why this cross-check is published as a disagreement rather than buried.

Iranian-related ancestry.

what is established, and the part that is not

The component usually called "Iranian farmer-related" is the largest single element in the ancestry of the Indus world and of everyone descended from it:

The part of the question the published data cannot answer

Which specific communities in each part of South India carry the highest Iranian-related proportions is not established at that resolution in the published literature. What exists is coarser: a study of West Maharashtra found the Brahmin caste there carries higher Ancient Iranian and Steppe contributions than the Kunbi Maratha caste, and the general pattern across the subcontinent is that these proportions correlate with traditional rank.

A community-by-community, region-by-region table for South India would be an invention, and this page will not produce one. The sampling that would support it has not been published. That is a gap in the record, and naming it is more useful than filling it.

Two different measurements, kept apart

Founder events and ancestry proportions are genome-wide. Haplogroup frequencies are not: a Y-haplogroup traces one unbroken father-to-son line, one ancestor out of more than a thousand in the same generation. A group can carry a lineage at high frequency and hold little of the ancestry that lineage arrived with, or the reverse. Nothing on this page converts one into the other, and where both appear they are labelled.

The honest limits.

Made visible: that a five-million-strong merchant caste descends from roughly a hundred people, that founder events run from Brahmin to Dalit rather than concentrating at the bottom, and that the Dravidian root date carries a credible interval spanning the pre-Indus, Indus and post-Indus periods alike. Kept dark: that the widely quoted 4,500-year figure is a midpoint, that the calibration is anchored on the first surviving inscription rather than on speech, and that the branching order flips under a different set of constraints. Put outside: everyone a community stopped marrying two thousand years ago, and the medical consequences their descendants still carry.

Exhibit

The claims, their sources, and what is still open

This page carries a reversed correction — a withdrawn interval that was later restored. The record of that belongs on the page.

Mela Keela is an independent digital museum. Sources, evidentiary limits and review status are identified on individual pages.