Evidence page · the open question
Four thousand inscriptions. Average length, five signs. Not one sign has been deciphered.
A century after the seals were first excavated, the Indus script remains unread — and the reason is not a shortage of effort. Over a hundred decipherment claims have been published. The obstacle is structural: the texts are too short to test anything against. That single fact is why every competing theory survives, why the argument is fought over politics rather than evidence, and why a state government has now attached a million dollars to it.
What actually exists.
the corpus
| The material | What it is |
|---|---|
| Around 4,000 inscriptions | Almost exclusively on seals, terracotta tablets and pottery. No archive, no monument text, no bilingual. |
| Average length ~5 signs | Most inscriptions would fit on a luggage tag. |
| The longest single text | The Dholavira signboard — around ten feet wide, with ten characters some 37 cm tall, set over the city's northern gateway. The longest surviving Indus text is ten signs on a sign. |
| 67 signs carry 80% | Of a total sign inventory usually estimated in the hundreds, 67 signs account for roughly 80% of all usage — a heavily skewed distribution. |
| Signs deciphered | None. "Not a single sign is deciphered yet." |
Why brevity defeats decipherment.
the structural problem
With five-sign texts, almost any proposed reading can be made to fit somewhere and none can be broken anywhere. A decipherment is normally tested by applying it to a text it was not built on and seeing whether the result is coherent. There is no such text. That is why over a hundred claims exist, why they contradict each other, and why the field has no mechanism to reject any of them. The obstacle is not that nobody is clever enough. It is that the evidence cannot adjudicate.
Three positions.
and what each would need to be right
| Position | Case | Difficulty |
|---|---|---|
| It encodes a language Parpola, Mahadevan, Rao | Sign order is regular and positionally structured. A 2009 study in Science found the script's conditional entropy closer to natural languages than to various non-linguistic sign systems. | Entropy establishes structured sequence, not language. Comparable regularity appears in non-linguistic systems too, and the point was contested immediately. |
| It is not writing Farmer, Sproat & Witzel, 2004 | "The Collapse of the Indus-Script Thesis: The Myth of a Literate Harappan Civilization." Texts are far too short; sign repetition is unlike known scripts; a literate state should have left longer documents somewhere. | Cannot explain the positional regularity or the standardisation of an inventory across a million square kilometres. |
| Administrative notation | The seals record ownership, commodity, quantity or licence — a system that carries information without encoding speech. | Would explain both the brevity and the structure. Testing it requires exactly the long texts that do not exist. |
The three are not equally distant from each other. "Structured notation that is not full writing" is compatible with much of the evidence both other camps produce — and is the least satisfying answer, because it means the seals were never going to tell us what the Indus people thought.
Why the fight is political.
what a decipherment would settle
One position holds that the Indus language was Indo-Aryan — that Sanskrit and its relatives originated in the Indus Valley and spread outward toward Europe. As one specialist summarised the claim: "Everything was within India to begin with. Nothing came from outside." The other holds it was ancestral Dravidian — that Dravidian languages were spoken widely across the region before being displaced in the north.
Each would settle the origin question this platform documents at length. Which is why the script is not treated as an open technical problem but as a contested title deed — and why researchers in the field report receiving death threats.
In January 2025 the Chief Minister of Tamil Nadu, M. K. Stalin, announced a prize of one million dollars for a decipherment satisfying archaeological experts, saying the riddle had gone unanswered for a hundred years. Stalin is among those who hold the Indus language to be a Dravidian ancestor. The prize is real and the problem is real; the announcement also came from a party to the dispute.
The study that prompted it
The prize followed a Tamil Nadu government-commissioned study by K. Rajan and R. Sivananthan, comparing graffiti marks on over 14,000 ceramic sherds from Tamil Nadu against Indus signs and reporting that around 60% showed morphological parallels.
What it establishes: that potters' marks in Iron Age Tamil Nadu resemble Indus signs in shape. What it does not establish: that they mean the same thing, belong to one system, or descend from one another. Simple marks — crosses, arrows, ladders, chevrons — resemble each other across unrelated traditions worldwide because the hand and the tool constrain them. A morphological match rate is a starting point for investigation, not a demonstration of continuity. And the study was commissioned by the government that then offered a prize for the conclusion it points toward, which is worth a reader knowing.
The claims that arrive weekly.
and how to read them
A researcher who has worked on the script for over a decade reports receiving emails every week from people who say they have solved it and that the case is closed. Recent claims include cryptographic approaches asserting that every inscription can now be read.
A genuine decipherment does three things: it is derived from one set of texts and then successfully applied to texts it was not built on; it produces readings that are contextually appropriate — a seal reading as an owner's name rather than a philosophical maxim; and it predicts, so that a newly excavated inscription can be read before anyone knows what it says.
No claim so far has passed the third test, and with five-sign texts the first two are weak filters. Any decipherment announced without a prediction should be treated as unproven — including one that wins a prize.
The honest limits.
- Sources: corpus size, sign distribution and the "not a single sign" assessment from Nisha Yadav (Tata Institute of Fundamental Research) via BBC News reporting; the entropy finding from Rao et al., Science, 2009; the non-linguistic thesis from Farmer, Sproat & Witzel, Electronic Journal of Vedic Studies 11 (2004); the concordance from Mahadevan (ASI Memoirs, 1977); the prize announcement and the two-camp characterisation from CNN and Smithsonian reporting, January–February 2025; the Rajan and Sivananthan graffiti study as reported at the time of the announcement.
- Not used: self-published and AI-assisted blog treatments of the script, of which several circulate with confident conclusions and no peer review. Pseudonymous decipherment claims posted to preprint and paper-sharing sites are noted as existing and are not treated as findings.
- The sign inventory is itself disputed — estimates range widely depending on how variants are counted, which is part of why the corpus is hard to analyse.
- NOT claimed: that the script is or is not writing. All three positions are set out; none is endorsed.
- NOT claimed: anything about what language the Indus people spoke. No genome indicates language, and no undeciphered script does either.
What is missing, and why it defeats decipherment
four thousand inscriptions, average length five signs
This page is about an absence of evidence. Drawing the absence is more honest than describing it.