Evidence page · nine texts, counted
Nine Sanskrit texts spanning roughly 1,700 years, counted with one script. The words that name the caste system are almost absent at the start and saturate the law code at the end. This is not an argument about the texts. It is a measurement of them.
All four varna terms, fourteen texts, per 10,000 words. Note the scale: 0–135.
Brāhmaṇa reaches 131.9 in the Vasiṣṭha Dharmasūtra and 54.7 in Manu. It is the most frequent of the four at every stage. Kṣatriya 21.6, vaiśya 43.5, śūdra 68.0 at their peaks.
One caveat: brāhmaṇa also names a genre — the Brāhmaṇa texts. Its high readings in the Kauṣītaki (94.9) and Gopatha (45.4) are inflated by self-reference. The Dharmasūtra and Manu figures are not — those texts are not Brāhmaṇas and have no reason to name the genre.
The word for the group at the top is the most frequent of the four in every law code.
The category words. Scale 0–50.
Dvija, "twice-born", reaches 47.2 in Manu from 0.4 in the Rigveda and zero across every Brāhmaṇa. It is the largest single jump in the whole dataset and it is the term that divides those entitled to initiation from those who are not.
Jāti is zero until the Upaniṣads. Caṇḍāla is zero until the Chāndogya.
And varṇa is the flattest line on the chart — 1.3 in the Rigveda, 0.5 in the Śatapatha, 14.6 in Manu. It rises tenfold while śūdra rises three hundredfold. The category word grows slowly; the words for the people inside the categories grow explosively.
| Text | Date | Words | varṇa | jāti | vaiśya | śūdra | caṇḍāla | mleccha | dharma |
|---|---|---|---|---|---|---|---|---|---|
| Rigveda | c.1500–1000 BCE | 180,196 | 1.3 | 0.0 | 0.1 | 0.1 | 0.0 | 0.0 | 3.4 |
| Paippalāda Atharvaveda | c.1200–1000 BCE | 126,459 | 1.2 | 0.0 | 0.2 | 1.1 | 0.1 | 0.0 | 1.4 |
| Sāmaveda | c.1200–1000 BCE | 41,285 | 1.2 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 3.6 |
| Pañcaviṃśa Brāhmaṇa | c.900–700 BCE | 49,862 | 1.8 | 0.0 | 1.0 | 0.6 | 0.0 | 0.0 | 0.8 |
| Gopatha Brāhmaṇa | c.800–600 BCE | 36,061 | 4.2 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 1.7 |
| Kauṣītaki Brāhmaṇa | c.800–600 BCE | 19,075 | 3.7 | 0.0 | 1.6 | 0.5 | 0.0 | 0.0 | 1.0 |
| Bṛhadāraṇyaka Upaniṣad | c.700–600 BCE | 85,467 | 4.3 | 1.2 | 0.9 | 1.1 | 0.0 | 0.0 | 10.4 |
| Chāndogya Upaniṣad | c.700–600 BCE | 44,880 | 2.7 | 0.9 | 0.4 | 1.6 | 0.9 | 0.0 | 8.5 |
| Manusmṛti | c.200 CE | 70,538 | 14.6 | 5.0 | 19.4 | 31.5 | 3.4 | 0.7 | 82.5 |
1 · jāti — caste as actually lived — does not exist until the Upaniṣads. Zero in the Rigveda, zero in the Atharvaveda, zero in the Sāmaveda, zero in all three Brāhmaṇas. It appears first in the Bṛhadāraṇyaka and Chāndogya, and reaches 5.0 per 10,000 in Manu. The unit of caste is younger than the philosophy that is supposed to justify it.
2 · śūdra rises 315-fold. One occurrence in the entire Rigveda. In Manu it is one of the most frequent nouns in the text.
3 · mleccha is entirely absent until Manu. Not Rigvedic, not Atharvavedic, not Brāhmaṇic in this sample, not Upaniṣadic. The word for the barbarian is a late arrival.
4 · caṇḍāla appears at 0.1 in the Atharvaveda, 0.9 in the Chāndogya, 3.4 in Manu. The category is invented, then filled in.
The vocabulary of caste is not inherited. It is assembled, and the assembly takes a thousand years.
Dharma is roughly as frequent in the Rigveda as varṇa. In Manu it is more frequent than almost anything else in the text.
And the jump happens in two stages: a first rise in the Upaniṣads (10.4 and 8.5), where it means cosmic and personal order, and a second explosion in Manu, where it means the rules governing what your birth permits you to do.
The word did not become more important. It became a different word — and the count shows exactly where.
Adding the four Dharmasūtras changes the shape of the finding
Fourteen texts, per 10,000 words.
śūdra: 1.6 in the Chāndogya → 37.2 in the Āpastamba Dharmasūtra. A twenty-three-fold jump in a single step. vaiśya: 0.4 → 28.6. varṇa: 2.7 → 22.9.
This is not a gradual accumulation across a thousand years. The Vedic and Upaniṣadic material sits flat and low. Then the Dharmasūtras arrive and the vocabulary saturates immediately — and it stays saturated through Manu and the later Smṛtis, seven hundred years on.
Caste vocabulary does not grow into the tradition. It arrives with a genre.
And the genre is law. The texts that carry it are not hymns, not ritual manuals, not philosophy — they are codes of conduct with penalties attached. Vasiṣṭha reaches 68.0 for śūdra, higher than Manu.
The Dharmasūtras are conventionally placed c. 500–300 BCE — which is the lifetime of the Buddha and Mahāvīra, and the period in which the śramaṇa movements were arguing that birth confers nothing.
The correlation is real and the causation is not established. Codification may be a response to challenge; it may be independent; the datings are loose enough that the ordering cannot be pressed. But the two things happen in the same window, and the older literature shows no comparable interest in ranking birth.
Kinship, purity, mixture, and the vocabulary of ranked birth
The first table used ten words. Here are eleven more, including the ones that carry the actual machinery: mixture, direction of marriage, twice-born status, ritual purification, pollution.
| Term | Rigveda | AV | Brāhmaṇas (max) | BṛhadĀr. | Chānd. | Manu | factor |
|---|---|---|---|---|---|---|---|
| varṇasaṃkara caste-mixing | 0 | 0 | 0 | 0 | 0 | 0.9 | ∞ |
| apasada base-born | 0 | 0 | 0 | 0 | 0 | 0.9 | ∞ |
| antyaja last-born, outcaste | 0 | 0 | 0 | 0.1 | 0 | 2.0 | ∞ |
| ambaṣṭha a mixed category | 0 | 0 | 0 | 0 | 0 | 1.1 | ∞ |
| dvija twice-born | 0.4 | 0.1 | 0 | 6.9 | 6.0 | 49.3 | 123× |
| kṣatriya | 0.5 | 1.2 | 3.7 | 1.8 | 0.2 | 14.7 | 29× |
| kula family, lineage | 0.1 | 0.2 | 0.3 | 0.4 | 0.4 | 6.5 | 65× |
| saṃskāra purificatory rite | 0 | 0 | 11.5 | 2.8 | 1.1 | 4.8 | ∞ |
| aśuci / apavitra impure | 0 | 0.1 | 0 | 0 | 0.2 | 2.6 | ∞ |
| patita fallen | 0.1 | 0.7 | 0.5 | 2.1 | 3.8 | 6.5 | 65× |
| pratiloma against the grain | 0 | 0 | 0 | 0.7 | 0.4 | 0.6 | — |
Zero occurrences. Not in the Rigveda, not in the Atharvaveda, not in the Sāmaveda, not in any of the three Brāhmaṇas sampled, not in the Bṛhadāraṇyaka, not in the Chāndogya. It appears in Manu.
And it is the entire basis of Arjuna's refusal to fight. Six verses at Bhagavad Gītā 1.39–44 — family destroyed, women corrupted, castes mixed, ancestors fallen — resting on a concept that has no attestation anywhere in the Vedic corpus.
The moral crisis at the opening of the Gītā is built on a word the Vedas never use.
The Gītā sits inside the Mahābhārata, which is not in this sample and should be. But the Vedic silence is complete, and it is the Vedas the Gītā claims continuity with.
Its first appearances here are in the Upaniṣads — and the Bṛhadāraṇyaka use is the one already documented on this platform: Ajātaśatru telling the Brahmin Gārgya that a Brahmin becoming a Kṣatriya's pupil is pratiloma, against the grain.
The word appears first as against-the-grain teaching, and only later as against-the-grain marriage. The metaphor was applied to knowledge before it was applied to women.
apasada, antyaja, ambaṣṭha, aśuci — all zero across every Vedic text sampled, all present in Manu.
dvija, "twice-born," rises 123-fold — and its first substantial use is in the Upaniṣads, not the hymns. The idea that ritual initiation constitutes a second birth, and that some people never get one, is not a Vedic idea. It is built on top of the Vedas.
166,253 words of Vedic ritual prose — the largest text in the corpus
| Term | Rigveda | ŚATAPATHA | Manu |
|---|---|---|---|
| varṇa | 1.3 | 0.5 | 14.6 |
| jāti | 0.0 | 0.0 | 5.0 |
| caṇḍāla | 0.0 | 0.0 | 3.4 |
| dvija — twice-born | 0.4 | 0.0 | 47.2 |
| varṇasaṃkara | 0.0 | 0.0 | 0.9 |
| antyaja | 0.0 | 0.0 | 1.6 |
| śūdra | 0.1 | 1.0 | 31.5 |
| kṣatriya | 0.5 | 2.2 | 14.2 |
| dharma | 3.4 | 2.0 | 82.5 |
166,000 words. Zero jāti. Zero caṇḍāla. Zero dvija. Zero varṇasaṃkara. Zero antyaja.
And varṇa at 0.5 — less than half the Rigveda's rate, in a text roughly the same length.
This is the text that contains the Bṛhadāraṇyaka Upaniṣad as its final section. It is the fullest surviving account of Vedic ritual practice — the sacrifices, the fire altars, the rājasūya, the aśvamedha, described procedure by procedure across fourteen books.
The most detailed record of what Vedic religion actually consisted of contains none of the vocabulary the religion is now identified with.
What it does have is kṣatriya at 2.2 — its highest rate anywhere before the law codes — which fits a text preoccupied with royal consecration. The word is doing work about kingship, not about a birth order.
And dharma at 2.0, lower than the Rigveda's 3.4. The term that reaches 82.5 in Manu is almost absent from the ritual literature it supposedly grows out of.
Synonym sets, not single words. Fourteen texts, per 10,000 words.
Each category has more than one word, and using only one hides material.
| Category | Terms counted | Manu, single word | Manu, full set |
|---|---|---|---|
| Brahmin | brāhmaṇaḥ · vipra · bhūdeva | 54.7 | 21.5 — LOWER, see below |
| Kṣatriya | kṣatriya · rājanya | 14.2 | 17.0 |
| Untouchable | caṇḍāla · antyaja · antyāvasāyin · śvapāka · pukkasa · niṣāda | 3.4 | 9.8 — nearly three times higher |
The untouchability vocabulary was undercounted by a factor of three, because Manu uses five different words: caṇḍāla 24, niṣāda 18, antyaja 14, pukkasa 12, śvapāka 6.
Masculine brāhmaṇa = a brahmin man. Neuter brāhmaṇa = a ritual explanation, and the name of the Brāhmaṇa texts. A search on the stem catches both.
The Kauṣītaki's apparent 94.9 is mostly brāhmaṇam at 71.8 — the neuter, meaning "the explanation." It is a text saying "the explanation is…", not a text about brahmins.
Using the unambiguous masculine nominative brāhmaṇaḥ instead: Rigveda 0.0, Śatapatha 1.0, Manu 5.4. That undercounts the other way, since the masculine appears in other cases too. The truth is between them, and the honest figure is a range, not a point.
The chart above uses brāhmaṇaḥ plus vipra — a conservative floor. The Rigveda's 5.9 is almost entirely vipra, the inspired singer, which is a functional term rather than a birth category.
480 texts, every section of GRETIL over 3,000 words
Raw string counts. See the disambiguation below before reading the upper bars.
Vedic Saṃhitās: 732,756 words. Brāhmaṇas: 107,492 words. Occurrences of jāti: none.
Everywhere else in Sanskrit literature the word is present. It is in the Upaniṣads, the epics, the Purāṇas, the philosophical corpus, the Buddhist corpus, the law books and the Arthaśāstra. The two sections where it does not occur are the two oldest.
Nearly eighteen million words of Sanskrit contain the word for birth-group. The eight hundred thousand oldest do not.
| Sense | Where it lives | What it means there |
|---|---|---|
| BIRTH | Buddhist literature — 256 identified instances | jāti-jarā-maraṇa, birth-old-age-death. Jātismara — one who remembers former births. The rebirth cycle, with no social content at all |
| GENUS | Nyāya and the philosophical corpus — 157 identified | A technical logic term. Nyāyasūtra 1.2.18 defines jāti as a specific fallacy of objection; elsewhere it is the universal against the particular, jāti against vyakti |
| CASTE | Dharmaśāstra — 35 identified in 452,743 words, the highest concentration anywhere | jāti-bhraṣṭa, fallen from one's birth-group · jāti-parivṛtti, changing it · jāti-ācāra, the conduct proper to it |
So the Buddhist corpus's apparent 5.4 per ten thousand is largely not caste at all. It is the rebirth cycle. And the philosophical corpus's 4.6 is largely a term in formal logic.
Reading the raw counts without this would produce the exact wrong conclusion: that Buddhist literature is more preoccupied with caste than the law books are. The opposite is true, and the string count cannot see it.
Disambiguation is only needed where the word occurs. In the Saṃhitās and the Brāhmaṇas the raw count is zero — so the word is absent in every sense: not as caste, not as birth, not as genus.
No amount of re-reading recovers it, because there is nothing to re-read.
The caste vocabulary is concentrated in the two latest books
Per 10,000 words. Shaded rows are Śānti and Anuśāsana, the didactic parvans.
Śānti and Anuśāsana are the two parvans textual scholarship treats as the largest late didactic accretions — the deathbed instruction of Bhīṣma, a body of teaching attached to the narrative rather than arising from it.
They carry the caste vocabulary and the rest of the epic does not.
| Parvan | varṇa | jāti | śūdra | dharma |
|---|---|---|---|---|
| ŚĀNTI — 164,223 words | 8.6 | 1.7 | 4.6 | 101.1 |
| ANUŚĀSANA — 83,324 words | 8.0 | 1.9 | 11.2 | 76.0 |
| Droṇa — the battle books | 1.6 | 0.0 | 0.1 | 22.7 |
| Śalya | 0.5 | 0.0 | 0.2 | 35.5 |
| Karṇa | 0.8 | 0.2 | 1.4 | 28.1 |
Anuśāsana has śūdra at 11.2 — a hundred times the Droṇa parvan's 0.1, and nine times the Bhīṣma parvan that contains the Gītā.
Śānti has dharma at 101.1 — higher than Manu's 82.5. A section of an epic is denser in the word than the law code itself.
The story does not carry the caste material. The instruction attached to the story does.
This is a textual argument made independently of the counts — Sukthankar's Critical Edition and a century of philology identify these parvans as accretions on stylistic and manuscript grounds. The vocabulary distribution agrees with that conclusion without being derived from it.
Varṇasaṃkara — the mixing-of-classes anxiety — occurs 17 times in 921,642 words. That is 0.2 per ten thousand. It is present, it is concentrated in Ādi, Bhīṣma and the two didactic parvans, and it is nowhere near as prominent as its later reputation suggests.
Sapiṇḍa occurs 3 times and sagotra 4 — in an epic of nearly a million words. The marriage-prohibition machinery that defines north Indian kinship is effectively absent from the epic as it is absent from the Veda. It belongs to the Dharmaśāstra and to nothing earlier.
Added — the Black Yajurveda, 166,000 words
| Term | occurrences | per 10k |
|---|---|---|
| brāhmaṇa | 92 | 5.5 |
| varṇa | 52 | 3.1 |
| dharma | 39 | 2.4 |
| kṣatriya + rājanya | 37 | 2.2 |
| śūdra | 12 | 0.7 |
| vaiśya | 9 | 0.5 |
| niṣāda | 3 | 0.2 |
| jāti | 0 | 0.0 |
| caṇḍāla | 0 | 0.0 |
| dvija | 0 | 0.0 |
Rigveda · Paippalāda Atharvaveda · Sāmaveda · Taittirīya Saṃhitā · four Brāhmaṇas including the Śatapatha · the Bṛhadāraṇyaka.
Jāti: zero. Dvija: zero outside a handful in the Rigveda. Caṇḍāla: zero.
Seven hundred thousand words of Vedic text without a single occurrence of the word that names caste as it is actually lived.
The Taittirīya is the Yajurveda of the south — the recension of the Taittirīya śākhā, the dominant Vedic school of Tamil Nadu, Andhra, Karnataka and Kerala. It has all four varna words, at low rates, and none of the caste machinery.
Method note: this text is a padapāṭha — each word separated for recitation — which inflates the token count relative to continuous prose. The rates above are therefore conservative; the true per-word frequencies are higher, and the zeroes are unaffected.
Two problems in the source files, both now fixed
The GRETIL Bṛhadāraṇyaka file is 85,467 words. The Upaniṣad itself is about 17,000. The remainder is Śaṅkara's commentary, written around the eighth century CE — roughly 1,300 years after the text.
So every figure this page reported for "the Upaniṣads" was measuring Śaṅkara.
The fix is in the same source. The Bṛhadāraṇyaka is the fourteenth kāṇḍa of the Śatapatha Brāhmaṇa. Taking that book on its own gives the Upaniṣad with no commentary at all — and confirms it: 106 of the 145 occurrences of Yājñavalkya's name, and all 53 of Gārgī's, are in kāṇḍa 14.
| Term | BU + Śaṅkara 85,467 words | BU alone 17,091 words | Where the occurrences actually are |
|---|---|---|---|
| jāti | 1.2 | 0.0 | ALL in the commentary |
| untouchability terms caṇḍāla · pukkasa · śvapāka · niṣāda | 0.6 | 0.0 | ALL in the commentary |
| dvija | 0.4 | 0.0 | ALL in the commentary |
| varṇa | 4.3 | 0.6 | Mostly commentary |
| dharma | 10.4 | 8.2 | Genuinely in the Upaniṣad — and six times the rate of Śatapatha books 1–13 |
The claim that jāti first appears in the Upaniṣads is wrong. It does not appear in the Bṛhadāraṇyaka at all. It appears in an eighth-century commentary on the Bṛhadāraṇyaka.
The corrected result is stronger and simpler:
Jāti — the endogamous birth-group, caste as actually lived — is absent from the Rigveda, the Atharvaveda, the Sāmaveda, four Brāhmaṇas, and the Bṛhadāraṇyaka Upaniṣad. Roughly half a million words. Zero occurrences.
It enters with the law codes. Everything earlier is either silence or a commentator writing a thousand years later.
And the same correction applies to the untouchability vocabulary. Last week's count found pukkasa and śvapāka in the Bṛhadāraṇyaka. They are Śaṅkara's words, not the Upaniṣad's.
The Chāndogya file has the same problem and no clean equivalent is available here. Its caṇḍāla at 5.10.7 is genuinely in the text — that passage is well attested — but its jāti figure should be treated as unverified until the commentary can be stripped.
My Śatapatha figures included kāṇḍa 14 — that is, they included the Bṛhadāraṇyaka. The two were not independent measurements. Books 1–13 alone, 149,152 words:
varṇa 0.5 · jāti 0.0 · untouchability 0.0 · dvija 0.0 · dharma 1.3
And kṣatriya plus rājanya at 5.7 — the highest rate anywhere before the law codes, which supports the reading that this text is preoccupied with kingship rather than with a birth order. That claim now rests on a clean measurement.
Meanwhile dharma jumps from 1.3 in books 1–13 to 8.2 in kāṇḍa 14 — a six-fold rise inside a single text, between its ritual books and its philosophical one. Same school, same recension, same manuscript tradition. That is the genre effect, isolated and measured.
A law code about social duties will mention social categories more than a ritual manual does — regardless of when it was written. The Dharmasūtras are about who may do what. The Brāhmaṇas are about how to build a fire altar. Some of the rise measures subject matter, not time.
Three things limit that objection without removing it:
1 · The Brāhmaṇas do discuss social matters — who may officiate, who may attend, who may receive. The Śatapatha spends 166,000 words on ritual and still records zero jāti, zero dvija, zero untouchability terms. A ritual text with strict eligibility rules and no vocabulary for birth-rank is informative.
2 · The Dharmasūtras are short. Āpastamba is 3,498 words; Gautama 4,636. Small denominators make rates volatile — a handful of occurrences produces a large figure. Manu, at 70,538 words, is the stable anchor and it still shows the pattern.
3 · And the genre effect can be isolated exactly. The Bṛhadāraṇyaka is the fourteenth book of the Śatapatha — same school, same recension, same manuscript. Between its ritual books and its philosophical one, dharma rises six-fold, 1.3 to 8.2 — while jāti, dvija and the untouchability vocabulary stay at zero in both.
Genre moves dharma. It does not move the caste vocabulary, because the caste vocabulary is not there to move.
String matching, not lemmatisation. These are prefix matches on stems, run over GRETIL transliterated texts. They will catch compounds and inflected forms — which is intended — but they cannot distinguish senses. The Rigveda figures here differ slightly from this platform's lemmatised counts for that reason.
brāhmaṇa is excluded from the chart because in Brāhmaṇa-genre texts the word names the genre as well as the person, and the count is uninterpretable without disambiguation.
Sample, not census. 480 texts, 17.9 million words for the register-wide scan; sixteen texts run in detail. The Śaunaka Atharvaveda is not in this run — the last of these particularly, since the Gītā's varṇasaṃkara passage is the strongest counter-case to the zero above.
Single source. All texts are from one archive, GRETIL. No cross-validation against a second edition has been run, and the editions behind GRETIL files vary in quality and in editorial convention. A second corpus would test whether any of these counts are artefacts of a particular edition.
Frequency is not meaning. A word can be present and mean something else — which is precisely the finding on varṇa elsewhere on this platform. The curve shows when the vocabulary arrives. It does not by itself show when the system did.
Dating is conventional and contested at the edges, particularly for the Brāhmaṇas. The ordering, not the absolute dates, carries the argument.