Skip to content
Mela Keela

Evidence page · lineage markers

Eighty-one per cent of Tamil male lineages are indigenous, and castes and tribes carry the same ones.

Y-chromosome haplogroups are the markers most often invoked when caste and ancestry are discussed together, and they are the most frequently misread. This page sets out which lineages the Dravidian-speaking populations actually carry, how they are distributed across caste and tribal groups, what structure exists in that distribution — and what the largest study of the question found the structure tracks, which is not rank.

The lineages.

what South Indian populations carry

HaplogroupOriginNote
H (H-M69, mostly H1)South AsianThe deepest indigenous paternal layer. Rarely found outside the subcontinent. The Romani carry H1a1a-M82 as their single South Asian paternal founder — the marker usually cited as establishing their South Asian origin. No source for this is given on the page; it is stated here as background, not as a finding of the study this page reports.
L (L1-M27)South AsianIndigenous; rarely found outside the subcontinent.
R2 (R2-M124)South AsianIndigenous; rarely found outside the subcontinent.
F* (F-M89)Ancient, South AsianA very deep lineage retained at appreciable frequency.
C5 (C5-M356)South AsianIndigenous.
J2West AsianAssociated with Neolithic-era connections to West Asia and the Iranian plateau.
R1a1 (R1a1-M17)Steppe-associatedThe lineage most often invoked in migration arguments — and, as below, the one whose treatment in the literature is most revealing.

Frequencies reported for Tamil castes: R1a1 27%, H 21% (predominantly H1), L 13% (predominantly L1), R2 11%, J2 11%, F* 10%. Each of F*, H1, J2, L1, R1a1 and R2 exceeds 5% in most castes — that is, the common lineages are shared across caste groups rather than sorted between them.

The finding that matters most.

1,680 Y chromosomes, 31 populations, Tamil Nadu

The largest study of the question this museum has located genotyped 1,680 Y chromosomes across 12 tribal and 19 caste endogamous populations from Tamil Nadu, taking advantage of the fact that both the Y chromosome and caste membership pass down the male line.

Two results, and the second is the one to hold onto

First: tribes and castes alike were characterised by an overwhelming proportion of indigenous South Asian lineages — H-M69, F-M89, R1a1-M17, L1-M27, R2-M124 and C5-M356 — 81% combined, with a shared heritage dating to the late Pleistocene, 10,000–30,000 years ago. Holocene migrations from western Eurasia contributed under 20% of male lineages.

Second: there is strong genetic structure in the data — and it is associated primarily with the current mode of subsistence. Not with caste rank. The paper's title states the conclusion: the differentiation of southern Indian male lineages correlates with agricultural expansions predating the caste system.

The structure is real. It tracks how people fed themselves, and it is older than the ranking.

The groups, and what each carries.

31 endogamous populations, sorted seven ways

The Tamil Nadu study did something unusual: rather than assume varṇa rank was the right way to group people, it tested several classifications against the data and let the genetics choose. The seven groups that emerged are defined by subsistence, funerary custom and mother tongue — not by rank.

GroupDefined byLanguage
HTF — Hill Tribe ForagersForagingTheir own Tamil/Malayalam dialects
HTC — Hill Tribes, CrematingCremation of the deadTamil
HTK — Hill Tribes, Kannada-speakingHunting and gatheringKannada — inside Tamil Nadu
SC — Scheduled CastesNon-landowning labourTamil
DLF — Dry Land FarmersMillets and grains, no irrigationTamil
AW — Artisans and WarriorsCraft and military occupationTamil
BRH — Brahmin-relatedVedic tradition; wetland irrigationTamil

Overall frequencies — seven haplogroups carry 82% of the variation

Inspect

sort by any column; filter to a lineage or a community

HaplogroupAll 31 populations
H1-M5217.4%
F*-M8916.3%
L1-M2714.0%
R1a1-M1712.7%
J2-M1729.4%
R2-M1248.2%
H-M694.7%

Eleven tables carry this page. The one that decides its argument is the group-by-group haplogroup table, and a reader who cannot sort it cannot test the claim that subsistence outperforms varna.

Where each lineage peaks — the actual per-group numbers

HaplogroupPeaks inFrequencypSignificantly absent from
F*-M89HTF — foragers53.3%<0.0001All BRH populations
H1-M52HTK — Kannada-speaking tribes42.5%<0.0001
R1a1-M17BRH — Brahmin-related42.2%<0.0001HTF (p=0.003)
L1-M27DLF — dry-land farmers24.1%<0.0001HTF (p=0.02)

But the study is explicit that these peaks conceal wide internal variation. F*-M89 in the forager group ranges from 75% down to 28.6% across its constituent populations. The same holds for H1-M52 within HTK and L1-M27 within DLF. Not all populations within a group share a genetic makeup — drift, fragmentation and isolation have acted differently on each.

Which classification the data actually prefers.

AMOVA, and the number that decides it

Grouping testedBetween-group variance
Seven groups by subsistence, custom and languageFCT = 0.065
Geographylower
Varṇa rank statuslower
Tribe / caste dichotomylower

The subsistence grouping maximised difference between groups and minimised variation within them — significantly better than grouping by geography or by varṇa rank. And breaking it down further is where it becomes interesting:

SubsetFCTReading
The three tribal groups alone0.095Strong differentiation
The four non-tribal groups alone0.015Six times less. Castes are far more similar to each other than tribes are.
Non-tribal groups with Brahmin-related removed0.004Structure almost vanishes. Nearly all the differentiation among Tamil castes is one group.
All groups with foragers removed0.027Between-group variance more than halves.
Read that third row again

Remove the Brahmin-related populations from the caste side and between-group genetic structure among Tamil castes drops from 1.5% to 0.4% — close to nothing. Scheduled Castes, dry-land farmers, artisans and warriors are barely distinguishable from one another.

The ranked order runs through all of them and this measurement does not track it.What the measurement supports, and what it does notIt would be broader than the result supports to conclude that whatever separates a Scheduled Caste community from a landholding one, it is not ancestry — in three ways at once. This is a Y-chromosome result, one ancestral line out of more than a thousand, as this page argues further down; it is a Tamil Nadu result, and the authors say so; and genome-wide work on endogamy does detect fine-grained differentiation between jātis. What the AMOVA shows is stated above and stands: between-group variance among the non-Brahmin-related Tamil caste groups is 0.4%, and the ranking does not explain what little structure there is. The museum’s own The genome of caste carries the matching guard — “NOT claimed: that one regional Y-chromosome study settles the all-India question”.

The same appears in the PCA and MDS plots: HTF, HTK and BRH form three distinct, distant clusters, and every other population is interspersed in the middle — the foragers at the F-M89 pole, the Kannada-speaking tribes at H1-M52, the Brahmin-related at R1a1-M17, and everyone else between.

The language connection.

who is closest to whom, and it is not who lives nearby

Tamil tribes resemble Dravidian tribes a thousand kilometres away more than their own neighbours

Compared against 97 populations from India and its neighbours, the Tamil Nadu hill tribes showed greater genetic similarity to Dravidian tribal groups in Andhra Pradesh and Odisha than to the caste populations living beside them. And the Tamil Nadu Brahmin-related populations clustered with Indo-European-speaking populations from multiple regions.

So genetic affinity here tracks linguistic and social community across long distances rather than geographic proximity. Two populations in the same district can be less alike than one of them is to a group in another state speaking a related language.

The Kannada-speaking hill tribes make the point sharply. They are hunter-gatherers inside Tamil Nadu speaking a different Dravidian language, they carry the highest H1-M52 frequency of any group, and they cluster with the foragers rather than with the Tamil-speaking populations around them. Language boundary, subsistence boundary and genetic boundary coincide — and the varṇa boundary runs across all three without matching any.

When the groups separated.

coalescence dates, and what they rule out

EventEstimate95% CI
All groups begin diverging7.1 kya5.5–9.2
Cremating tribes + all castes node6.2 kya4.7–8.0
Foragers and Kannada tribes split from the rest4.9 kya3.6–7.1
Youngest detectable admixture between groups3.0 kya2.3–4.3
Vanniyar expansion2.3 kyamatches historical record
Nadar expansion1.0 kyamatches historical record
The chronology that undoes the standard story

Endogamous social stratification in Tamil Nadu was in place between 6,000 and 4,000 years ago, with essentially no admixture between groups for the last 3,000. The varṇa system was implemented locally around 1,000 years ago, under the Pallavas and Cholas between the sixth and twelfth centuries CE, following Brahmin migration into the region.

The structure is three to five thousand years older than the system said to have created it. Sangam literature of 300 BCE–300 CE already names Valayar, Pulayar, Paliyan and Kadar, and the same names are carried by communities today. A historical claim without a citationThis is a literary-historical claim with no source on the page, and it is doing real work: it is the only non-genetic evidence offered that the groups the study measured are continuous with the groups named in the Sangam corpus. It needs a text-and-line citation, and until it has one it is not load-bearing here. Reading kuḍi — a broad term for a household, family or lineage — as an attested endogamous unit is an interpretation, and that reading is not made on this page. Name-continuity over two millennia is in any case weak evidence of population continuity. The study's conclusion is that varṇa was superimposed on a pre-existing structure without significant population transfer, and that its genetic impact was minimal.

And the R1a1 story does not hold up either

Brahmin-related populations do carry the highest R1a1-M17 frequency in the sample, at 42.2%. But three separate results argue against reading that as an incoming Indo-Aryan lineage:

Together the authors read these as arguing against introduction of these lineages through a single wave of Brahmin migration. And the Brahmin-related populations carry no significant frequency of any haplogroup with a likely origin outside India.

One more result that inverts an assumption

The non-tribal groups carry the older lineages

Tribes are routinely treated as the living remnant of India's earliest settlers. Measured by haplogroup age, the non-tribal populations showed older age estimates than the tribal ones for every haplogroup except R2-M124.

The explanation offered is the frontier model: expanding agricultural populations absorbed lineages from many pre-existing forager groups, accumulating diversity, while the foragers who were displaced into the Western Ghats went through founder effects and drift that stripped theirs. The tribal groups are not older; they are more isolated. Forager haplogroup diversity is 0.687 and 0.748 against 0.881 for the non-tribal populations — and F*-M89 is the only lineage showing clear population-specific clusters, in exactly those isolated groups.

The maternal line tells a different story.

and the asymmetry is the point

Where Y lineages show some structure, mitochondrial DNA shows very little. Tribal groups analysed for mtDNA coalesce at Indian-specific branches of haplogroups M and N that cover populations of different social rank across the whole subcontinent. Studies of caste groups in eastern India found Brahmins with high paternal affinity to eastern Europeans through R1a1 — while maternal polymorphisms revealed primarily Indian-specific lineages.

Men arrived; women largely did not

The asymmetry is consistent across studies: external input is visible in the paternal line and largely absent from the maternal one. Whatever movements occurred were strongly male-biased — incoming men, local women.

mtDNA also shows higher gene flow into caste populations than into tribal groups, and the Andhra pattern has been read as evidence of historical upward female mobility within the ranking — an interpretation offered in the secondary literature, not a result this page has verified against the study. Rank was, for a period, more permeable through women than through men — which is precisely the arrangement the law books spend their energy prohibiting.

How the data got handled when it was inconvenient.

the Chenchu case

A tribal group with high R1a, called "an aberrant phenomenon"

An influential proposal defined a package of haplogroups — J2, R1a, R2 and L — heuristically associated with Indo-European migration and the introduction of the caste system, on the grounds that these occurred at lower proportions in South Indian tribal groups.

But the Chenchus of Andhra Pradesh — a tribal group — carry R1a at high frequency. That observation was set aside as "an aberrant phenomenon." A marker defined as the signature of an incoming elite turned up in a forest-dwelling tribal population, and the response was to classify the data point as anomalous rather than to question the package. A later review criticised precisely this "one haplogroup equals one migration" reasoning.

What these markers actually measure.

the mechanics, because the findings are unreadable without them

A haplogroup traces one line out of thousands

Go back ten generations and you have 1,024 ancestors. The Y chromosome traces exactly one of them — father's father's father, unbroken. Mitochondrial DNA traces exactly one other — mother's mother's mother. Between them they capture two ancestors out of a thousand and tell you nothing about the other 1,022.

This is why haplogroup findings and genome-wide findings can look contradictory without conflicting. They are answering different questions. "81% of male lineages are indigenous" is a statement about one inherited line in each man. It is not a statement that these populations are 81% indigenous in overall ancestry — the genome-wide estimates, which integrate all those thousands of ancestors, give different proportions. Anyone quoting a haplogroup percentage as though it were an ancestry percentage has made a category error, and the error is extremely common.

Why the paternal and maternal records diverge

The consistent finding — external input visible in the Y, largely absent from mtDNA — follows from three mechanisms working together:

MechanismEffect
Male reproductive varianceA small number of men can father a very large number of children; women cannot. Male lineages can therefore sweep through a population far faster than female ones, and a militarily or politically dominant male group leaves a disproportionate Y signature relative to its actual numbers.
PatrilocalityWhere women move to their husband's community at marriage and men stay, mtDNA is redistributed locally between neighbouring groups while Y lineages stay put. This homogenises maternal variation regionally and sharpens paternal variation.
Directional hypergamyWhere marriage rules permit women to move upward in rank but not men, the maternal line carries ancestry up the hierarchy while the paternal line does not carry it down. The Andhra mtDNA pattern has been read this way.
What follows for the migration question

A strongly male-biased signal is compatible with a numerically small incoming group whose men married local women and whose descendants expanded. It does not require mass population replacement, and it does not support one either. The size of a genetic signature and the size of a migration are not the same quantity — which is exactly why the same data gets cited by people arguing for invasion and people arguing against it.

What endogamy does to a genome, and how fast

Sealing a community has predictable consequences that compound over generations:

The trap this sets

Measurable genetic differences between caste groups are exactly what endogamy alone predicts. They are evidence that groups stopped exchanging marriage partners. They are not evidence that the groups had different origins — and the difference between those two readings is the whole argument. A community that sealed in 500 BCE and drifted for a hundred generations will look distinct today whether its founders came from the next village or from Central Asia.

The unit problem: varṇa is not the endogamous unit

This is the flaw that runs through much of the literature, and it maps directly onto a distinction this platform enforces everywhere.

What it isDoes it seal?
VarṇaThe fourfold textual scheme — brāhmaṇa, kṣatriya, vaiśya, śūdra. A classification.No. Not a marriage circle.
JātiThe endogamous birth-group. Thousands of them, regionally specific.Yes. This is the unit that seals.

Endogamy operates at jāti level. A study that recruits samples labelled "Brahmin" or "Kshatriya" is aggregating dozens of distinct jātis that have not intermarried with each other for two thousand years — pooling populations that are, genetically, no more one group than they are one group with their neighbours.

Why this explains the confusing results

It accounts for the pattern that otherwise looks like noise. Rank correlations appear in one state and vanish in the next; a Brahmin sample turns out closest to a Relli sample; a tribal group carries the marker of the supposed incoming elite. These are what you would expect if the real structure is thousands of small sealed units, and the varṇa labels laid over them are a classification that does not track the sealing.

The finding that southern male-lineage structure correlates with subsistence mode rather than rank fits the same picture. Communities that farmed the same way in the same region were near each other and sometimes married each other. The ranking was a separate system, applied on top.

Two timescales that get collapsed

TimescaleWhat it holds
10,000–30,000 yearsThe shared late-Pleistocene South Asian inheritance. Almost everyone in the subcontinent, at every rank.
4,000–2,000 yearsAdmixture from the northwest, male-biased, differentially distributed. This is the layer that varies between communities.
3,000–2,000 years and afterSealing. Founder events, drift, and the differentiation visible today.

Arguments about caste and genetics almost always collapse these. A claim that caste groups are "genetically distinct" is usually true at the third timescale, marginally true at the second, and false at the first — and the three get stated as one. The deepest layer is shared, the middle layer varies, and the most recent layer is the one that produced today's boundaries. The boundaries are the youngest thing in the data.

What these markers cannot do.

stated because this is where genetics gets weaponised

The sources, audited.

run before relying on them

SourceFinding
ArunKumar et al. 2012 (PLOS One 7(11) e50269, with a published correction in 2013)
the 81% and subsistence result
Credentials sound; funding must be disclosed. Produced at The Genographic Laboratory, School of Biological Sciences, Madurai Kamaraj University, Madurai — an Indian institution working on its own region's populations, which is worth noting on a platform that tracks who gets to study whom. The corresponding author is named on the paper, with that affiliation. Authors declare no competing interests. A correction to this paper was published in 2013; readers should consult the corrected version, and the figures reproduced here are taken from it.

The disclosure: this is a Genographic Project output — a population-genetics programme funded by the National Geographic Society and IBM, which drew sustained objection from indigenous-rights organisations over consent and the commercial use of indigenous DNA. That criticism concerns the programme's collection practices, not the analysis reproduced here, and the paper's central result cuts against a migrationist reading rather than for one. Recorded so a reader can weigh it.

Post-publication: a claim that the first author now works in commercial genomics is not made here. No source is given for it, this museum could not verify it, and it is a claim about a named living person’s present employment. It is withdrawn under C-24. If it is documented it can be restored with a citation; as an unsourced aside about an individual it does not belong on the page, and nothing here depends on it.
Kivisild et al.
H, L, R2 shared across castes and tribes
Estonian Biocentre and Cambridge; a central and heavily cited figure in South Asian population genetics. A separate laboratory from both the Reich cluster and the Genographic programme — which makes the convergence between them on the shared-inheritance finding meaningful rather than circular.
Sahoo et al. 2006
the Chenchu "aberrant phenomenon"
PNAS. Used here for its critique of the "one haplogroup equals one migration" reasoning and for its record of how the Chenchu R1a observation was handled. Cited as documentation of the field's own argument, not as a settled finding.
Excluded, and why

Commercial ancestry-testing sites publish the tidiest per-caste haplogroup tables available — clean percentage ranges by community and by state, exactly the format this question invites. They carry no published sampling method, no sample sizes, no peer review, and a product to sell. None of their figures appear on this page. Where a number here looks less precise than one found elsewhere, that is the reason.

The honest limits.

Made visible: that the genetics of Tamil Nadu sorts by subsistence, custom and language rather than rank; that removing one group makes structure among the castes almost vanish; and that endogamous separation is three to five thousand years older than the varṇa system said to have produced it. Kept dark: that the highest R1a1 diversity sits in the Scheduled Castes and dry-land farmers rather than in the group with the highest frequency, and that non-tribal populations carry the older lineages. Put outside: the forager communities displaced into the Western Ghats, whose isolation is read as antiquity when it is loss.

Exhibit

Every claim on this page, and its source

The page reproduces seven tables from one paper. What the museum has read of that paper, and at what depth, belongs beside them.

Mela Keela is an independent digital museum. Sources, evidentiary limits and review status are identified on individual pages.