Veḷi

Evidence page · machine-counted

Śūdra appears once in the Rigveda. It appears 1,053 times in the law books.

If caste were the ancient, eternal order of Indian society, its vocabulary would be distributed evenly through the literature. It is not. Counted across 480 Sanskrit texts, the words of exclusion are nearly absent from the oldest layer and become dense in the legal one — and the untouchable category enters later still. This page publishes the counts, the layers, and the limits of the method, so the result can be checked rather than believed.

The curve.

occurrences per 100,000 words, by textual layer

Counted across 758 texts, 18.8 million words, searching each term together with its vṛddhi and orthographic variants. Raw counts are given beneath.

LayerFilesWordscaṇḍālaaspṛśyaantyajaśūdraśūdra rate
Saṃhitā5735,0430.10.00.11.81.8
Brāhmaṇa3107,4920.00.00.02.82.8
Āraṇyaka112,9930.00.00.00.00.0
Upaniṣad18202,4642.00.50.55.45.4
Vedāṅga / Sūtra18324,8572.213.50.319.719.7
Epic7849,2831.42.50.07.57.5
Buddhist2133,526,3523.10.60.13.93.9
Purāṇa192,033,4006.22.60.628.328.3
Arthaśāstra151,06517.60.00.043.143.1
Dharmaśāstra (law books)19460,49635.410.24.6221.5221.5
Raw counts — caṇḍāla: Saṃhitā 1 · Brāhmaṇa 0 · Upaniṣad 4 · Sūtra 7 · Epic 12 · Purāṇa 126 · Dharmaśāstra 163. śūdra: Saṃhitā 13 · Brāhmaṇa 3 · Upaniṣad 11 · Sūtra 64 · Purāṇa 576 · Dharmaśāstra 1,020. aspṛśya: zero in Saṃhitā and Brāhmaṇa · 44 in the Sūtra layer. Scan file gret_scan_v2.json ships with the site.
The word "untouchable" itself

Aspṛśya — literally "not-to-be-touched" — is the term that names the practice rather than the people. It occurs zero times in the Saṃhitās and zero times in the Brāhmaṇas. It appears in the Upaniṣads twice, and then jumps to 13.5 per 100,000 words in the Vedāṅga and Sūtra literature — 44 occurrences — before settling at 10.2 in the law books. The concept of untouchability, stated in its own word, is not a Vedic inheritance. It is a rule-book development.

The śūdra rate rises more than a hundredfold from the oldest layer to the legal one. Not because society was newly stratified in the first millennium CE, but because a body of literature appeared whose subject is rank — and which had to name the ranked in order to regulate them.

The vocabulary of exclusion is thin where the hymns are, and dense where the law is.

The first appearance of the outcaste.

caṇḍāla across the whole corpus

The strongest single result concerns caṇḍāla, the term for the outcaste. Of 480 texts scanned, 132 contain it. The Vedic section contains almost none.

TextcaṇḍālaNote
Rigveda (both editions)0Absent.
Sāmaveda Saṃhitā0Absent.
Atharvaveda, Paippalāda Saṃhitā1The only occurrence in any Vedic Saṃhitā in this corpus.
Pañcaviṃśa, Gopatha, Kauṣītaki Brāhmaṇas0Absent across the Brāhmaṇa layer sampled.
Taittirīya Saṃhitā0Direct file check, 282,992 tokens. Confirmed absent.
Śatapatha Brāhmaṇa, books 1–130The Brāhmaṇa proper, 149,152 tokens. Absent — as are paulkasa and niṣāda.
Śatapatha book 14 (= Bṛhadāraṇyaka Upaniṣad)1ŚB 14.7.1.22. The latest stratum — and the category is negated there, not applied. See below.
Śāṅkhāyana Gṛhyasūtra2Sūtra layer.
Chāndogya Upaniṣad with commentary4Commentary is far later than the Upaniṣad; see limits.
Manusmṛti2434.0 per 100k.
Parāśarasmṛti (ācāra/prāyaścitta)50314.8 per 100k — the densest text in the entire corpus.
Kauṭilya, Arthaśāstra1529.4 per 100k — the category is administratively operational.
What the distribution shows

A category that will later carry the weight of untouchability — with its own rules of distance, dwelling, food, and punishment — appears exactly once in the Vedic hymn corpus, and then becomes one of the densest words in the legal literature. Its home is the law book and the treatise on statecraft, not the hymn.

This is what a constructed category looks like in a corpus: absent where the tradition claims its origins lie, ubiquitous where the rules are written. The tradition's own texts date its arrival.

The one Vedic occurrence, read.

ŚB 14.7.1.22 — where the category is dissolved

A count is not a reading. The single occurrence of the outcaste term anywhere in the Śatapatha was located and checked directly in the file. It falls in book 14 — which is the Bṛhadāraṇyaka Upaniṣad, the latest layer of the text — and it does not do what a reader might expect.

The passage, in transliteration

…atra steno 'steno bhavati bhrūṇahā 'bhrūṇahā paulkaso 'paulkasaś cāṇḍālo 'cāṇḍālaḥ śramaṇo 'śramaṇas tāpaso 'tāpaso…

There a thief is not a thief; a killer of an embryo, not a killer of an embryo; a paulkasa not a paulkasa; a caṇḍāla not a caṇḍāla; a śramaṇa not a śramaṇa; an ascetic not an ascetic.

The context is the description of the liberated or deep-sleeping self, in which social and moral categories cease to apply. The word's only appearance in the entire Śatapatha is in a sentence stating that the category does not hold.

The distribution, stated precisely

The result is often quoted as caṇḍāla = 0 in the Śatapatha Brāhmaṇa. That figure comes from searching the form caṇḍāla and missing the vṛddhi form cāṇḍāla. Checked directly against the text, the picture is sharper than the flat zero:

Zero in books 1–13 — the Brāhmaṇa proper, 149,152 tokens. Also zero for paulkasa and niṣāda there. One occurrence in book 14, the Upaniṣadic appendix, alongside paulkasa and śramaṇa, each appearing exactly once, in that same sentence.

So the vocabulary of exclusion does not merely arrive late — in this text it arrives in a passage that suspends it. The ranked categories enter the Śatapatha at the moment the text describes a state in which ranking stops.

What the Atharvaveda adds.

the first shift

The Atharvaveda — here the Paippalāda recension, 126,459 words — is where the vocabulary of social ranking first becomes noticeable. Set against the Rigveda:

TermRigveda /100kAtharvaveda /100kChange
śūdra0.611.1~18× denser
caṇḍāla0.00.8First appearance
strī (woman)2.215.8~7× denser
varṇa12.811.9Flat
jāti (birth-group)0.00.0Still absent from both

Two things are worth noticing together. Varṇa is flat — it is present from the start, because it is an ordinary word (and, as this platform documents, it means colour before it means rank). But śūdra, the name of a ranked class, and strī, the marked category of woman, both rise sharply in the same text. The Atharvaveda is where the literature begins to sort people. And jāti — the birth-group that eventually becomes the operative unit of caste — is absent from every Vedic Saṃhitā in this corpus, arriving only later. See jātization.

Jāti, with the rebirth sense removed.

a figure that was unusable, now measured

The raw count for jāti — "birth" — cannot be read as caste vocabulary, because in Buddhist Sanskrit it is also a core doctrinal term meaning birth as rebirth, a link in dependent origination. The raw Buddhist rate of 58 per 100,000 words says nothing about social rank.

Each of the 3,273 occurrences across three layers was therefore classified by the words around it: doctrinal where the context carries the vocabulary of dependent origination (jarāmaraṇa, bhava, upādāna, saṃsāra, cyuti), social where it carries the vocabulary of rank (brāhmaṇa, śūdra, varṇa, kula, gotra, saṃkara), and unclear where neither dominates.

LayerTotalDoctrinalSocialUnclear
Buddhist2,4281,339235854
Purāṇa4399072277
Dharmaśāstra40649184173

Note the Dharmaśāstra row first. Only 12% of its jāti occurrences are doctrinal. The law books use the word almost entirely for social rank — while in the Buddhist corpus 55% of occurrences are the rebirth doctrine.

LayerSocial jāti per 100,000 words
Purāṇa3.5 – 17.2
Buddhist6.7 – 30.9
Dharmaśāstra40.0 – 77.5
Each range runs from the strict count (social only) to the loose count (social plus unclear). Ranges rather than points, because the classification is contextual.
The result holds however the line is drawn

The Dharmaśāstra's lowest estimate (40.0) is higher than the Buddhist corpus's highest (30.9). The two ranges do not overlap. Whichever end of the uncertainty is taken, birth-group vocabulary is concentrated in the legal literature.

And the raw figure inverted the picture entirely. Uncorrected, the Buddhist corpus appeared to be the most jāti-preoccupied body of literature in the sample. Corrected, it is the least concerned with birth as a social category and the most concerned with birth as a cosmological one.

Limits.

how this could be wrong

Reproduce this.

corpus, query, gap

The method in full

Corpus: GRETIL (Göttingen Register of Electronic Texts in Indian Languages), via the tokushige-koyasan/gretil-corpus plain-text mirror — 784 files, of which 758 exceeding 200 tokens were scanned, totalling 18,777,671 tokens, in IAST transliteration.

Query: Unicode-NFC normalised, case-folded substring counts for each term together with its vṛddhi and orthographic variants (e.g. caṇḍāla + cāṇḍāla + caṇḍala + cāṇḍala), across every file; densities as occurrences × 100,000 ÷ tokens; files grouped by GRETIL directory section. For jāti, each occurrence was additionally classified on a ±220-character window, scoring doctrinal markers (jarāmaraṇa, bhava, upādāna, saṃsāra, cyuti) against rank markers (brāhmaṇa, śūdra, varṇa, kula, gotra, saṃkara). Output: gret_scan_v2.json, shipped with the site.

Gap: no sense disambiguation; the Śaunaka Atharvaveda is absent from GRETIL itself, so the Paippalāda recension stands in for it; commentaries are not separated from base texts; and varṇa is excluded from the argument as contaminated (it also means colour, letter and praise). Jāti was disambiguated by context window — doctrinal markers against rank markers, 3,273 occurrences classified — and is reported as a range rather than a point.