Evidence page · nine texts, counted

Caste was built, and the build is visible in the vocabulary.

Nine Sanskrit texts spanning roughly 1,700 years, counted with one script. The words that name the caste system are almost absent at the start and saturate the law code at the end. This is not an argument about the texts. It is a measurement of them.

315×increase in śūdra, Rigveda → Manu
194×increase in vaiśya
24×increase in dharma
0 → 5.0jāti, absent until the Upaniṣads
LAW CODES 0 27 54 81 108 135 per 10,000 words THE FOUR VARNAS brāhmaṇa 54.7 kṣatriya 14.2 vaiśya 19.4 śūdra 31.5 Rigveda Paippalāda AV Sāmaveda Pañcaviṃśa Br. Gopatha Br. Kauṣītaki Br. Śatapatha Br. Bṛhadāraṇyaka Chāndogya Āpastamba DhS Gautama DhS Baudhāyana DhS Vasiṣṭha DhS Manusmṛti

All four varna terms, fourteen texts, per 10,000 words. Note the scale: 0–135.

All four rise together, and brāhmaṇa rises most

Brāhmaṇa reaches 131.9 in the Vasiṣṭha Dharmasūtra and 54.7 in Manu. It is the most frequent of the four at every stage. Kṣatriya 21.6, vaiśya 43.5, śūdra 68.0 at their peaks.

One caveat: brāhmaṇa also names a genre — the Brāhmaṇa texts. Its high readings in the Kauṣītaki (94.9) and Gopatha (45.4) are inflated by self-reference. The Dharmasūtra and Manu figures are not — those texts are not Brāhmaṇas and have no reason to name the genre.

The word for the group at the top is the most frequent of the four in every law code.

LAW CODES 0 10 20 30 40 50 per 10,000 words THE CATEGORY WORDS varṇa 14.6 jāti 5.0 caṇḍāla 3.4 dvija 47.2 Rigveda Paippalāda AV Sāmaveda Pañcaviṃśa Br. Gopatha Br. Kauṣītaki Br. Śatapatha Br. Bṛhadāraṇyaka Chāndogya Āpastamba DhS Gautama DhS Baudhāyana DhS Vasiṣṭha DhS Manusmṛti

The category words. Scale 0–50.

These are the ones that arrive

Dvija, "twice-born", reaches 47.2 in Manu from 0.4 in the Rigveda and zero across every Brāhmaṇa. It is the largest single jump in the whole dataset and it is the term that divides those entitled to initiation from those who are not.

Jāti is zero until the Upaniṣads. Caṇḍāla is zero until the Chāndogya.

And varṇa is the flattest line on the chart — 1.3 in the Rigveda, 0.5 in the Śatapatha, 14.6 in Manu. It rises tenfold while śūdra rises three hundredfold. The category word grows slowly; the words for the people inside the categories grow explosively.

The table

TextDateWordsvarṇajātivaiśyaśūdracaṇḍālamlecchadharma
Rigvedac.1500–1000 BCE180,1961.30.00.10.10.00.03.4
Paippalāda Atharvavedac.1200–1000 BCE126,4591.20.00.21.10.10.01.4
Sāmavedac.1200–1000 BCE41,2851.20.00.00.00.00.03.6
Pañcaviṃśa Brāhmaṇac.900–700 BCE49,8621.80.01.00.60.00.00.8
Gopatha Brāhmaṇac.800–600 BCE36,0614.20.00.00.00.00.01.7
Kauṣītaki Brāhmaṇac.800–600 BCE19,0753.70.01.60.50.00.01.0
Bṛhadāraṇyaka Upaniṣadc.700–600 BCE85,4674.31.20.91.10.00.010.4
Chāndogya Upaniṣadc.700–600 BCE44,8802.70.90.41.60.90.08.5
Manusmṛtic.200 CE70,53814.65.019.431.53.40.782.5

What the curve shows

Four findings, each independently checkable

1 · jāti — caste as actually lived — does not exist until the Upaniṣads. Zero in the Rigveda, zero in the Atharvaveda, zero in the Sāmaveda, zero in all three Brāhmaṇas. It appears first in the Bṛhadāraṇyaka and Chāndogya, and reaches 5.0 per 10,000 in Manu. The unit of caste is younger than the philosophy that is supposed to justify it.

2 · śūdra rises 315-fold. One occurrence in the entire Rigveda. In Manu it is one of the most frequent nouns in the text.

3 · mleccha is entirely absent until Manu. Not Rigvedic, not Atharvavedic, not Brāhmaṇic in this sample, not Upaniṣadic. The word for the barbarian is a late arrival.

4 · caṇḍāla appears at 0.1 in the Atharvaveda, 0.9 in the Chāndogya, 3.4 in Manu. The category is invented, then filled in.

The vocabulary of caste is not inherited. It is assembled, and the assembly takes a thousand years.

And dharma is the runaway term

3.4 → 82.5 per 10,000 words

Dharma is roughly as frequent in the Rigveda as varṇa. In Manu it is more frequent than almost anything else in the text.

And the jump happens in two stages: a first rise in the Upaniṣads (10.4 and 8.5), where it means cosmic and personal order, and a second explosion in Manu, where it means the rules governing what your birth permits you to do.

The word did not become more important. It became a different word — and the count shows exactly where.

Sixteen texts — and it is a step, not a curve

Adding the four Dharmasūtras changes the shape of the finding

0 22 44 66 88 110 dharma 82.5 DHARMA — plotted alone because it exceeds every other scale Rigveda Paippalāda AV Sāmaveda Pañcaviṃśa Br. Gopatha Br. Kauṣītaki Br. Śatapatha Br. Bṛhadāraṇyaka Chāndogya Āpastamba DhS Gautama DhS Baudhāyana DhS Vasiṣṭha DhS Manusmṛti

Fourteen texts, per 10,000 words.

The discontinuity is at one boundary, and it is abrupt

śūdra: 1.6 in the Chāndogya → 37.2 in the Āpastamba Dharmasūtra. A twenty-three-fold jump in a single step. vaiśya: 0.4 → 28.6. varṇa: 2.7 → 22.9.

This is not a gradual accumulation across a thousand years. The Vedic and Upaniṣadic material sits flat and low. Then the Dharmasūtras arrive and the vocabulary saturates immediately — and it stays saturated through Manu and the later Smṛtis, seven hundred years on.

Caste vocabulary does not grow into the tradition. It arrives with a genre.

And the genre is law. The texts that carry it are not hymns, not ritual manuals, not philosophy — they are codes of conduct with penalties attached. Vasiṣṭha reaches 68.0 for śūdra, higher than Manu.

The date of that boundary

The Dharmasūtras are conventionally placed c. 500–300 BCE — which is the lifetime of the Buddha and Mahāvīra, and the period in which the śramaṇa movements were arguing that birth confers nothing.

The correlation is real and the causation is not established. Codification may be a response to challenge; it may be independent; the datings are loose enough that the ordering cannot be pressed. But the two things happen in the same window, and the older literature shows no comparable interest in ranking birth.

The second pass — twenty-one terms

Kinship, purity, mixture, and the vocabulary of ranked birth

The first table used ten words. Here are eleven more, including the ones that carry the actual machinery: mixture, direction of marriage, twice-born status, ritual purification, pollution.

TermRigvedaAVBrāhmaṇas
(max)
BṛhadĀr.Chānd.Manufactor
varṇasaṃkara
caste-mixing
000000.9
apasada
base-born
000000.9
antyaja
last-born, outcaste
0000.102.0
ambaṣṭha
a mixed category
000001.1
dvija
twice-born
0.40.106.96.049.3123×
kṣatriya0.51.23.71.80.214.729×
kula
family, lineage
0.10.20.30.40.46.565×
saṃskāra
purificatory rite
0011.52.81.14.8
aśuci / apavitra
impure
00.1000.22.6
patita
fallen
0.10.70.52.13.86.565×
pratiloma
against the grain
0000.70.40.6
varṇasaṃkara does not exist in Vedic literature

Zero occurrences. Not in the Rigveda, not in the Atharvaveda, not in the Sāmaveda, not in any of the three Brāhmaṇas sampled, not in the Bṛhadāraṇyaka, not in the Chāndogya. It appears in Manu.

And it is the entire basis of Arjuna's refusal to fight. Six verses at Bhagavad Gītā 1.39–44 — family destroyed, women corrupted, castes mixed, ancestors fallen — resting on a concept that has no attestation anywhere in the Vedic corpus.

The moral crisis at the opening of the Gītā is built on a word the Vedas never use.

The Gītā sits inside the Mahābhārata, which is not in this sample and should be. But the Vedic silence is complete, and it is the Vedas the Gītā claims continuity with.

And pratiloma enters as an epistemological term, not a sexual one

Its first appearances here are in the Upaniṣads — and the Bṛhadāraṇyaka use is the one already documented on this platform: Ajātaśatru telling the Brahmin Gārgya that a Brahmin becoming a Kṣatriya's pupil is pratiloma, against the grain.

The word appears first as against-the-grain teaching, and only later as against-the-grain marriage. The metaphor was applied to knowledge before it was applied to women.

The purity vocabulary arrives with the law code

apasada, antyaja, ambaṣṭha, aśuciall zero across every Vedic text sampled, all present in Manu.

dvija, "twice-born," rises 123-fold — and its first substantial use is in the Upaniṣads, not the hymns. The idea that ritual initiation constitutes a second birth, and that some people never get one, is not a Vedic idea. It is built on top of the Vedas.

Śatapatha Brāhmaṇa

166,253 words of Vedic ritual prose — the largest text in the corpus

TermRigvedaŚATAPATHAManu
varṇa1.30.514.6
jāti0.00.05.0
caṇḍāla0.00.03.4
dvija — twice-born0.40.047.2
varṇasaṃkara0.00.00.9
antyaja0.00.01.6
śūdra0.11.031.5
kṣatriya0.52.214.2
dharma3.42.082.5
The largest ritual text in Vedic literature has almost no caste vocabulary

166,000 words. Zero jāti. Zero caṇḍāla. Zero dvija. Zero varṇasaṃkara. Zero antyaja.

And varṇa at 0.5 — less than half the Rigveda's rate, in a text roughly the same length.

This is the text that contains the Bṛhadāraṇyaka Upaniṣad as its final section. It is the fullest surviving account of Vedic ritual practice — the sacrifices, the fire altars, the rājasūya, the aśvamedha, described procedure by procedure across fourteen books.

The most detailed record of what Vedic religion actually consisted of contains none of the vocabulary the religion is now identified with.

What it does have is kṣatriya at 2.2 — its highest rate anywhere before the law codes — which fits a text preoccupied with royal consecration. The word is doing work about kingship, not about a birth order.

And dharma at 2.0, lower than the Rigveda's 3.4. The term that reaches 82.5 in Manu is almost absent from the ritual literature it supposedly grows out of.

LAW CODES 0 14 28 42 56 70 per 10,000 words · SYNONYM SETS, not single words brahmin 21.5 kṣatriya + rājanya 17.0 vaiśya 19.4 śūdra 31.5 untouchable — 5 terms 9.8 dvija 47.2 Rigveda Paippalāda AV Sāmaveda Pañcaviṃśa Br. Gopatha Br. Kauṣītaki Br. Śatapatha Br. Bṛhadāraṇyaka Chāndogya Āpastamba DhS Gautama DhS Baudhāyana DhS Vasiṣṭha DhS Manusmṛti

Synonym sets, not single words. Fourteen texts, per 10,000 words.

Why single words undercount

Each category has more than one word, and using only one hides material.

CategoryTerms countedManu, single wordManu, full set
Brahminbrāhmaṇaḥ · vipra · bhūdeva54.721.5 — LOWER, see below
Kṣatriyakṣatriya · rājanya14.217.0
Untouchablecaṇḍāla · antyaja · antyāvasāyin · śvapāka · pukkasa · niṣāda3.49.8 — nearly three times higher

The untouchability vocabulary was undercounted by a factor of three, because Manu uses five different words: caṇḍāla 24, niṣāda 18, antyaja 14, pukkasa 12, śvapāka 6.

And brāhmaṇa is two different words

Masculine brāhmaṇa = a brahmin man. Neuter brāhmaṇa = a ritual explanation, and the name of the Brāhmaṇa texts. A search on the stem catches both.

The Kauṣītaki's apparent 94.9 is mostly brāhmaṇam at 71.8 — the neuter, meaning "the explanation." It is a text saying "the explanation is…", not a text about brahmins.

Using the unambiguous masculine nominative brāhmaṇaḥ instead: Rigveda 0.0, Śatapatha 1.0, Manu 5.4. That undercounts the other way, since the masculine appears in other cases too. The truth is between them, and the honest figure is a range, not a point.

The chart above uses brāhmaṇaḥ plus vipra — a conservative floor. The Rigveda's 5.9 is almost entirely vipra, the inspired singer, which is a functional term rather than a birth category.

The whole register — 17.9 million words

480 texts, every section of GRETIL over 3,000 words

JĀTI ACROSS 17.9 MILLION WORDS OF SANSKRIT raw occurrences per 10,000 words — before sense-disambiguation Vedic Saṃhitās — 732,756 w 0.0 Brāhmaṇas — 107,492 w 0.0 Rāmāyaṇa — 353,655 w 0.2 Upaniṣads — 189,487 w 2.1 Mahābhārata — 495,628 w 1.0 Purāṇas — 2,033,400 w 0.9 Buddhist corpus — 3,415,278 w 5.4 Philosophy — 4,503,469 w 4.6 Dharmaśāstra — 452,743 w 4.5 Arthaśāstra — 51,057 w 6.7

Raw string counts. See the disambiguation below before reading the upper bars.

Zero, in the two oldest sections, out of 840,248 words

Vedic Saṃhitās: 732,756 words. Brāhmaṇas: 107,492 words. Occurrences of jāti: none.

Everywhere else in Sanskrit literature the word is present. It is in the Upaniṣads, the epics, the Purāṇas, the philosophical corpus, the Buddhist corpus, the law books and the Arthaśāstra. The two sections where it does not occur are the two oldest.

Nearly eighteen million words of Sanskrit contain the word for birth-group. The eight hundred thousand oldest do not.

But jāti is three different words

Counting the string conflates senses that have nothing to do with each other
SenseWhere it livesWhat it means there
BIRTHBuddhist literature — 256 identified instancesjāti-jarā-maraṇa, birth-old-age-death. Jātismara — one who remembers former births. The rebirth cycle, with no social content at all
GENUSNyāya and the philosophical corpus — 157 identifiedA technical logic term. Nyāyasūtra 1.2.18 defines jāti as a specific fallacy of objection; elsewhere it is the universal against the particular, jāti against vyakti
CASTEDharmaśāstra — 35 identified in 452,743 words, the highest concentration anywherejāti-bhraṣṭa, fallen from one's birth-group · jāti-parivṛtti, changing it · jāti-ācāra, the conduct proper to it

So the Buddhist corpus's apparent 5.4 per ten thousand is largely not caste at all. It is the rebirth cycle. And the philosophical corpus's 4.6 is largely a term in formal logic.

Reading the raw counts without this would produce the exact wrong conclusion: that Buddhist literature is more preoccupied with caste than the law books are. The opposite is true, and the string count cannot see it.

And it makes the Vedic result unassailable

Disambiguation is only needed where the word occurs. In the Saṃhitās and the Brāhmaṇas the raw count is zero — so the word is absent in every sense: not as caste, not as birth, not as genus.

No amount of re-reading recovers it, because there is nothing to re-read.

Mahābhārata — 921,642 words, all eighteen parvans

The caste vocabulary is concentrated in the two latest books

CASTE VOCABULARY BY PARVAN — per 10,000 words ■ varṇa ■ śūdra ■ jāti gold rows = the two didactic parvans 0 3 6 9 12 Ādi dharma 61.8 · 91,152 words Sabhā dharma 76.1 · 29,814 words Vana dharma 62.2 · 129,277 words Virāṭa dharma 27.3 · 23,063 words Udyoga dharma 61.3 · 77,472 words Bhīṣma +Gītā dharma 16.3 · 64,238 words Droṇa dharma 22.7 · 96,093 words Karṇa dharma 28.1 · 49,148 words Śalya dharma 35.5 · 40,619 words Sauptika dharma 26.6 · 9,405 words Strī dharma 56.3 · 8,698 words ŚĀNTI dharma 101.1 · 164,223 words ANUŚĀSANA dharma 76.0 · 83,324 words Āśvamedhika dharma 71.5 · 34,829 words Āśramavāsika dharma 81.1 · 13,187 words Mausala dharma 11.9 · 3,350 words Mahāprasthānika dharma 123.4 · 1,378 words Svargārohaṇa dharma 118.0 · 2,372 words

Per 10,000 words. Shaded rows are Śānti and Anuśāsana, the didactic parvans.

The distribution matches the textual argument exactly

Śānti and Anuśāsana are the two parvans textual scholarship treats as the largest late didactic accretions — the deathbed instruction of Bhīṣma, a body of teaching attached to the narrative rather than arising from it.

They carry the caste vocabulary and the rest of the epic does not.

Parvanvarṇajātiśūdradharma
ŚĀNTI — 164,223 words8.61.74.6101.1
ANUŚĀSANA — 83,324 words8.01.911.276.0
Droṇa — the battle books1.60.00.122.7
Śalya0.50.00.235.5
Karṇa0.80.21.428.1

Anuśāsana has śūdra at 11.2 — a hundred times the Droṇa parvan's 0.1, and nine times the Bhīṣma parvan that contains the Gītā.

Śānti has dharma at 101.1 — higher than Manu's 82.5. A section of an epic is denser in the word than the law code itself.

The story does not carry the caste material. The instruction attached to the story does.

This is a textual argument made independently of the counts — Sukthankar's Critical Edition and a century of philology identify these parvans as accretions on stylistic and manuscript grounds. The vocabulary distribution agrees with that conclusion without being derived from it.

And two absences worth recording

Varṇasaṃkara — the mixing-of-classes anxiety — occurs 17 times in 921,642 words. That is 0.2 per ten thousand. It is present, it is concentrated in Ādi, Bhīṣma and the two didactic parvans, and it is nowhere near as prominent as its later reputation suggests.

Sapiṇḍa occurs 3 times and sagotra 4 — in an epic of nearly a million words. The marriage-prohibition machinery that defines north Indian kinship is effectively absent from the epic as it is absent from the Veda. It belongs to the Dharmaśāstra and to nothing earlier.

Taittirīya Saṃhitā

Added — the Black Yajurveda, 166,000 words

Termoccurrencesper 10k
brāhmaṇa925.5
varṇa523.1
dharma392.4
kṣatriya + rājanya372.2
śūdra120.7
vaiśya90.5
niṣāda30.2
jāti00.0
caṇḍāla00.0
dvija00.0
The count now stands at roughly 700,000 words

Rigveda · Paippalāda Atharvaveda · Sāmaveda · Taittirīya Saṃhitā · four Brāhmaṇas including the Śatapatha · the Bṛhadāraṇyaka.

Jāti: zero. Dvija: zero outside a handful in the Rigveda. Caṇḍāla: zero.

Seven hundred thousand words of Vedic text without a single occurrence of the word that names caste as it is actually lived.

The Taittirīya is the Yajurveda of the south — the recension of the Taittirīya śākhā, the dominant Vedic school of Tamil Nadu, Andhra, Karnataka and Kerala. It has all four varna words, at low rates, and none of the caste machinery.

Method note: this text is a padapāṭha — each word separated for recitation — which inflates the token count relative to continuous prose. The rates above are therefore conservative; the true per-word frequencies are higher, and the zeroes are unaffected.

A contamination, and the clean result

Two problems in the source files, both now fixed

The Upaniṣad files are mostly Śaṅkara

The GRETIL Bṛhadāraṇyaka file is 85,467 words. The Upaniṣad itself is about 17,000. The remainder is Śaṅkara's commentary, written around the eighth century CE — roughly 1,300 years after the text.

So every figure this page reported for "the Upaniṣads" was measuring Śaṅkara.

The fix is in the same source. The Bṛhadāraṇyaka is the fourteenth kāṇḍa of the Śatapatha Brāhmaṇa. Taking that book on its own gives the Upaniṣad with no commentary at all — and confirms it: 106 of the 145 occurrences of Yājñavalkya's name, and all 53 of Gārgī's, are in kāṇḍa 14.

TermBU + Śaṅkara
85,467 words
BU alone
17,091 words
Where the occurrences actually are
jāti1.20.0ALL in the commentary
untouchability terms
caṇḍāla · pukkasa · śvapāka · niṣāda
0.60.0ALL in the commentary
dvija0.40.0ALL in the commentary
varṇa4.30.6Mostly commentary
dharma10.48.2Genuinely in the Upaniṣad — and six times the rate of Śatapatha books 1–13
What this does to the finding

The claim that jāti first appears in the Upaniṣads is wrong. It does not appear in the Bṛhadāraṇyaka at all. It appears in an eighth-century commentary on the Bṛhadāraṇyaka.

The corrected result is stronger and simpler:

Jāti — the endogamous birth-group, caste as actually lived — is absent from the Rigveda, the Atharvaveda, the Sāmaveda, four Brāhmaṇas, and the Bṛhadāraṇyaka Upaniṣad. Roughly half a million words. Zero occurrences.

It enters with the law codes. Everything earlier is either silence or a commentator writing a thousand years later.

And the same correction applies to the untouchability vocabulary. Last week's count found pukkasa and śvapāka in the Bṛhadāraṇyaka. They are Śaṅkara's words, not the Upaniṣad's.

The Chāndogya file has the same problem and no clean equivalent is available here. Its caṇḍāla at 5.10.7 is genuinely in the text — that passage is well attested — but its jāti figure should be treated as unverified until the commentary can be stripped.

And the second contamination

My Śatapatha figures included kāṇḍa 14 — that is, they included the Bṛhadāraṇyaka. The two were not independent measurements. Books 1–13 alone, 149,152 words:

varṇa 0.5 · jāti 0.0 · untouchability 0.0 · dvija 0.0 · dharma 1.3

And kṣatriya plus rājanya at 5.7 — the highest rate anywhere before the law codes, which supports the reading that this text is preoccupied with kingship rather than with a birth order. That claim now rests on a clean measurement.

Meanwhile dharma jumps from 1.3 in books 1–13 to 8.2 in kāṇḍa 14 — a six-fold rise inside a single text, between its ritual books and its philosophical one. Same school, same recension, same manuscript tradition. That is the genre effect, isolated and measured.

The strongest objection to this whole chart

Genre, not chronology

A law code about social duties will mention social categories more than a ritual manual does — regardless of when it was written. The Dharmasūtras are about who may do what. The Brāhmaṇas are about how to build a fire altar. Some of the rise measures subject matter, not time.

Three things limit that objection without removing it:

1 · The Brāhmaṇas do discuss social matters — who may officiate, who may attend, who may receive. The Śatapatha spends 166,000 words on ritual and still records zero jāti, zero dvija, zero untouchability terms. A ritual text with strict eligibility rules and no vocabulary for birth-rank is informative.

2 · The Dharmasūtras are short. Āpastamba is 3,498 words; Gautama 4,636. Small denominators make rates volatile — a handful of occurrences produces a large figure. Manu, at 70,538 words, is the stable anchor and it still shows the pattern.

3 · And the genre effect can be isolated exactly. The Bṛhadāraṇyaka is the fourteenth book of the Śatapatha — same school, same recension, same manuscript. Between its ritual books and its philosophical one, dharma rises six-fold, 1.3 to 8.2 — while jāti, dvija and the untouchability vocabulary stay at zero in both.

Genre moves dharma. It does not move the caste vocabulary, because the caste vocabulary is not there to move.

Limits — read before citing

What this measurement can and cannot support

String matching, not lemmatisation. These are prefix matches on stems, run over GRETIL transliterated texts. They will catch compounds and inflected forms — which is intended — but they cannot distinguish senses. The Rigveda figures here differ slightly from this platform's lemmatised counts for that reason.

brāhmaṇa is excluded from the chart because in Brāhmaṇa-genre texts the word names the genre as well as the person, and the count is uninterpretable without disambiguation.

Sample, not census. 480 texts, 17.9 million words for the register-wide scan; sixteen texts run in detail. The Śaunaka Atharvaveda is not in this run — the last of these particularly, since the Gītā's varṇasaṃkara passage is the strongest counter-case to the zero above.

Single source. All texts are from one archive, GRETIL. No cross-validation against a second edition has been run, and the editions behind GRETIL files vary in quality and in editorial convention. A second corpus would test whether any of these counts are artefacts of a particular edition.

Frequency is not meaning. A word can be present and mean something else — which is precisely the finding on varṇa elsewhere on this platform. The curve shows when the vocabulary arrives. It does not by itself show when the system did.

Dating is conventional and contested at the edges, particularly for the Brāhmaṇas. The ordering, not the absolute dates, carries the argument.