Evidence page · machine-counted
If caste were the ancient, eternal order of Indian society, its vocabulary would be distributed evenly through the literature. It is not. Counted across 480 Sanskrit texts, the words of exclusion are nearly absent from the oldest layer and become dense in the legal one — and the untouchable category enters later still. This page publishes the counts, the layers, and the limits of the method, so the result can be checked rather than believed.
occurrences per 100,000 words, by textual layer
Counted across 758 texts, 18.8 million words, searching each term together with its vṛddhi and orthographic variants. Raw counts are given beneath.
| Layer | Files | Words | caṇḍāla | aspṛśya | antyaja | śūdra | śūdra rate | |
|---|---|---|---|---|---|---|---|---|
| Saṃhitā | 5 | 735,043 | 0.1 | 0.0 | 0.1 | 1.8 | 1.8 | |
| Brāhmaṇa | 3 | 107,492 | 0.0 | 0.0 | 0.0 | 2.8 | 2.8 | |
| Āraṇyaka | 1 | 12,993 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | |
| Upaniṣad | 18 | 202,464 | 2.0 | 0.5 | 0.5 | 5.4 | 5.4 | |
| Vedāṅga / Sūtra | 18 | 324,857 | 2.2 | 13.5 | 0.3 | 19.7 | 19.7 | |
| Epic | 7 | 849,283 | 1.4 | 2.5 | 0.0 | 7.5 | 7.5 | |
| Buddhist | 213 | 3,526,352 | 3.1 | 0.6 | 0.1 | 3.9 | 3.9 | |
| Purāṇa | 19 | 2,033,400 | 6.2 | 2.6 | 0.6 | 28.3 | 28.3 | |
| Arthaśāstra | 1 | 51,065 | 17.6 | 0.0 | 0.0 | 43.1 | 43.1 | |
| Dharmaśāstra (law books) | 19 | 460,496 | 35.4 | 10.2 | 4.6 | 221.5 | 221.5 | |
gret_scan_v2.json ships with the site.Aspṛśya — literally "not-to-be-touched" — is the term that names the practice rather than the people. It occurs zero times in the Saṃhitās and zero times in the Brāhmaṇas. It appears in the Upaniṣads twice, and then jumps to 13.5 per 100,000 words in the Vedāṅga and Sūtra literature — 44 occurrences — before settling at 10.2 in the law books. The concept of untouchability, stated in its own word, is not a Vedic inheritance. It is a rule-book development.
The śūdra rate rises more than a hundredfold from the oldest layer to the legal one. Not because society was newly stratified in the first millennium CE, but because a body of literature appeared whose subject is rank — and which had to name the ranked in order to regulate them.
caṇḍāla across the whole corpus
The strongest single result concerns caṇḍāla, the term for the outcaste. Of 480 texts scanned, 132 contain it. The Vedic section contains almost none.
| Text | caṇḍāla | Note |
|---|---|---|
| Rigveda (both editions) | 0 | Absent. |
| Sāmaveda Saṃhitā | 0 | Absent. |
| Atharvaveda, Paippalāda Saṃhitā | 1 | The only occurrence in any Vedic Saṃhitā in this corpus. |
| Pañcaviṃśa, Gopatha, Kauṣītaki Brāhmaṇas | 0 | Absent across the Brāhmaṇa layer sampled. |
| Taittirīya Saṃhitā | 0 | Direct file check, 282,992 tokens. Confirmed absent. |
| Śatapatha Brāhmaṇa, books 1–13 | 0 | The Brāhmaṇa proper, 149,152 tokens. Absent — as are paulkasa and niṣāda. |
| Śatapatha book 14 (= Bṛhadāraṇyaka Upaniṣad) | 1 | ŚB 14.7.1.22. The latest stratum — and the category is negated there, not applied. See below. |
| Śāṅkhāyana Gṛhyasūtra | 2 | Sūtra layer. |
| Chāndogya Upaniṣad with commentary | 4 | Commentary is far later than the Upaniṣad; see limits. |
| Manusmṛti | 24 | 34.0 per 100k. |
| Parāśarasmṛti (ācāra/prāyaścitta) | 50 | 314.8 per 100k — the densest text in the entire corpus. |
| Kauṭilya, Arthaśāstra | 15 | 29.4 per 100k — the category is administratively operational. |
A category that will later carry the weight of untouchability — with its own rules of distance, dwelling, food, and punishment — appears exactly once in the Vedic hymn corpus, and then becomes one of the densest words in the legal literature. Its home is the law book and the treatise on statecraft, not the hymn.
This is what a constructed category looks like in a corpus: absent where the tradition claims its origins lie, ubiquitous where the rules are written. The tradition's own texts date its arrival.
ŚB 14.7.1.22 — where the category is dissolved
A count is not a reading. The single occurrence of the outcaste term anywhere in the Śatapatha was located and checked directly in the file. It falls in book 14 — which is the Bṛhadāraṇyaka Upaniṣad, the latest layer of the text — and it does not do what a reader might expect.
…atra steno 'steno bhavati bhrūṇahā 'bhrūṇahā paulkaso 'paulkasaś cāṇḍālo 'cāṇḍālaḥ śramaṇo 'śramaṇas tāpaso 'tāpaso…
There a thief is not a thief; a killer of an embryo, not a killer of an embryo; a paulkasa not a paulkasa; a caṇḍāla not a caṇḍāla; a śramaṇa not a śramaṇa; an ascetic not an ascetic.
The context is the description of the liberated or deep-sleeping self, in which social and moral categories cease to apply. The word's only appearance in the entire Śatapatha is in a sentence stating that the category does not hold.
The result is often quoted as caṇḍāla = 0 in the Śatapatha Brāhmaṇa. That figure comes from searching the form caṇḍāla and missing the vṛddhi form cāṇḍāla. Checked directly against the text, the picture is sharper than the flat zero:
Zero in books 1–13 — the Brāhmaṇa proper, 149,152 tokens. Also zero for paulkasa and niṣāda there. One occurrence in book 14, the Upaniṣadic appendix, alongside paulkasa and śramaṇa, each appearing exactly once, in that same sentence.
So the vocabulary of exclusion does not merely arrive late — in this text it arrives in a passage that suspends it. The ranked categories enter the Śatapatha at the moment the text describes a state in which ranking stops.
the first shift
The Atharvaveda — here the Paippalāda recension, 126,459 words — is where the vocabulary of social ranking first becomes noticeable. Set against the Rigveda:
| Term | Rigveda /100k | Atharvaveda /100k | Change |
|---|---|---|---|
| śūdra | 0.6 | 11.1 | ~18× denser |
| caṇḍāla | 0.0 | 0.8 | First appearance |
| strī (woman) | 2.2 | 15.8 | ~7× denser |
| varṇa | 12.8 | 11.9 | Flat |
| jāti (birth-group) | 0.0 | 0.0 | Still absent from both |
Two things are worth noticing together. Varṇa is flat — it is present from the start, because it is an ordinary word (and, as this platform documents, it means colour before it means rank). But śūdra, the name of a ranked class, and strī, the marked category of woman, both rise sharply in the same text. The Atharvaveda is where the literature begins to sort people. And jāti — the birth-group that eventually becomes the operative unit of caste — is absent from every Vedic Saṃhitā in this corpus, arriving only later. See jātization.
a figure that was unusable, now measured
The raw count for jāti — "birth" — cannot be read as caste vocabulary, because in Buddhist Sanskrit it is also a core doctrinal term meaning birth as rebirth, a link in dependent origination. The raw Buddhist rate of 58 per 100,000 words says nothing about social rank.
Each of the 3,273 occurrences across three layers was therefore classified by the words around it: doctrinal where the context carries the vocabulary of dependent origination (jarāmaraṇa, bhava, upādāna, saṃsāra, cyuti), social where it carries the vocabulary of rank (brāhmaṇa, śūdra, varṇa, kula, gotra, saṃkara), and unclear where neither dominates.
| Layer | Total | Doctrinal | Social | Unclear |
|---|---|---|---|---|
| Buddhist | 2,428 | 1,339 | 235 | 854 |
| Purāṇa | 439 | 90 | 72 | 277 |
| Dharmaśāstra | 406 | 49 | 184 | 173 |
Note the Dharmaśāstra row first. Only 12% of its jāti occurrences are doctrinal. The law books use the word almost entirely for social rank — while in the Buddhist corpus 55% of occurrences are the rebirth doctrine.
| Layer | Social jāti per 100,000 words | |
|---|---|---|
| Purāṇa | 3.5 – 17.2 | |
| Buddhist | 6.7 – 30.9 | |
| Dharmaśāstra | 40.0 – 77.5 | |
The Dharmaśāstra's lowest estimate (40.0) is higher than the Buddhist corpus's highest (30.9). The two ranges do not overlap. Whichever end of the uncertainty is taken, birth-group vocabulary is concentrated in the legal literature.
And the raw figure inverted the picture entirely. Uncorrected, the Buddhist corpus appeared to be the most jāti-preoccupied body of literature in the sample. Corrected, it is the least concerned with birth as a social category and the most concerned with birth as a cosmological one.
how this could be wrong
caṇḍāla without the vṛddhi form cāṇḍāla produces a false zero for the Śatapatha — a difference of one occurrence that changes the reading entirely. Any term here may carry variants not searched: the counts are a floor, not a ceiling, and a zero is the most fragile figure a scan can produce.corpus, query, gap
Corpus: GRETIL (Göttingen Register of Electronic Texts in Indian Languages), via the tokushige-koyasan/gretil-corpus plain-text mirror — 784 files, of which 758 exceeding 200 tokens were scanned, totalling 18,777,671 tokens, in IAST transliteration.
Query: Unicode-NFC normalised, case-folded substring counts for each term together with its vṛddhi and orthographic variants (e.g. caṇḍāla + cāṇḍāla + caṇḍala + cāṇḍala), across every file; densities as occurrences × 100,000 ÷ tokens; files grouped by GRETIL directory section. For jāti, each occurrence was additionally classified on a ±220-character window, scoring doctrinal markers (jarāmaraṇa, bhava, upādāna, saṃsāra, cyuti) against rank markers (brāhmaṇa, śūdra, varṇa, kula, gotra, saṃkara). Output: gret_scan_v2.json, shipped with the site.
Gap: no sense disambiguation; the Śaunaka Atharvaveda is absent from GRETIL itself, so the Paippalāda recension stands in for it; commentaries are not separated from base texts; and varṇa is excluded from the argument as contaminated (it also means colour, letter and praise). Jāti was disambiguated by context window — doctrinal markers against rank markers, 3,273 occurrences classified — and is reported as a range rather than a point.