Santhal DNA
Genetic origins and closest populations · India · South Asia
The Santhal are one of the largest Adivasi peoples of India, living mainly in Jharkhand, West Bengal, Odisha and Bihar, with communities in Bangladesh and Nepal. Their language, Santali, belongs to the Munda branch of the Austroasiatic family and has its own script, Ol Chiki, created in the 20th century. The Santhal rebellion of 1855 is remembered as one of the great uprisings against colonial rule and local landlords. In our model, the Santhal average is 82.0% South Asian hunter-gatherer related (Irula proxy) and 18.0% Southeast Asian Neolithic farmer. Genetic studies link that Southeast Asian share to the arrival of Austroasiatic speech from Southeast Asia, which then mixed with the much larger local population. The closest modern averages are other Santhal samples from Jharkhand and Bangladesh, the Bhumij of Odisha (0.0096) and Ho speakers, all Munda-speaking. The fit is loose (0.0479), and no ancient genome is close: the Great Andamanese, about 1850 AD, are at 0.1138.
- Ancient components in our model: South Asian hunter-gatherer related (Irula proxy) 82.0%, Southeast Asian Neolithic farmers 18.0%
- Closest modern population in our data: Santhal (Jharkhand), distance 0.0086
- Closest ancient group in our data: Andaman Islands 100BP (Great Andamanese), distance 0.1138
Where does Santhal DNA sit on the genetic map of the world? To answer, we take the averaged Global25 (G25) coordinates of the Santhal individuals sampled in India and compare them with deep ancestral reference populations, with other modern groups and with ancient genomes. It is one of 335 populations in our atlas.
Ancient make-up of the Santhal average
The ancient make-up of the Santhal average is led by South Asian hunter-gatherer related (Irula proxy) (82.0%), followed by Southeast Asian Neolithic farmers (18.0%). The South Asian hunter-gatherer related (Irula proxy) component reflects the deep hunter-gatherer ancestry of South Asia (Ancient Ancestral South Indians). The Southeast Asian Neolithic farmers share (18.0%) reflects the Neolithic farmers of mainland Southeast Asia. With well over half of the total, this single source dominates the profile.
The fit is loose (fit distance 0.048), which usually means that none of the available reference populations is a close stand-in for part of this ancestry, so read the percentages as rough. This population was modelled with a global set of reference populations (ancient genomes, plus modern stand-ins where no suitable ancient genome exists), because West Eurasian sources alone cannot describe it.
Closest modern populations to Santhal
Among the modern populations in the G25 data, the closest to the Santhal average is Santhal (Jharkhand), at a distance of 0.009 (very close). Next come Santhal (Bangladesh) at 0.009 and Bhumij (Odisha) at 0.010. Leaving aside the other regional samples listed under the same name in the G25 sheet, the nearest other populations are Bhumij (Odisha) (0.010), Bhumij (Jharkhand) (0.014) and Bhunjia (Chhattisgarh) (0.015). Even the tenth closest, Bhumij, is only 0.018 away: Santhal belongs to a dense cluster of related populations, so small differences inside that cluster should not be over-interpreted.
| # | Population | Distance | Closeness |
|---|---|---|---|
| 1 | Santhal (Jharkhand) | 0.0086 | Very close |
| 2 | Santhal (Bangladesh) | 0.0089 | Very close |
| 3 | Bhumij (Odisha) | 0.0096 | Very close |
| 4 | Bhumij (Jharkhand) | 0.0141 | Very close |
| 5 | Bhunjia (Chhattisgarh) | 0.0150 | Very close |
| 6 | Ho (Odisha) | 0.0158 | Very close |
| 7 | Korwa | 0.0161 | Very close |
| 8 | Bharia | 0.0168 | Very close |
| 9 | Birhor | 0.0174 | Very close |
| 10 | Bhumij | 0.0182 | Very close |
Distances are Euclidean distances between averaged G25 coordinates. On our scale, below 0.025 is very close, below 0.050 close, below 0.080 moderate, and beyond that distant.
Closest ancient populations to Santhal
If we search ancient DNA for the best match to Santhal, Andaman Islands 100BP (Great Andamanese) (c. 1850 AD) comes first, at 0.114, with Laos Hoabinhian next. Even the best ancient match is distant, so the modern population is better described as a blend of several ancestral sources than as the continuation of one. A close ancient match is not proof of direct descent: it means those individuals carried a similar overall mix of ancestry.
| # | Ancient sample or group | Period | Distance |
|---|---|---|---|
| 1 | Andaman Islands 100BP (Great Andamanese) | c. 1850 AD | 0.1138 |
| 2 | Laos Hoabinhian | Hoabinhian foragers (pre-Neolithic) | 0.1202 |
| 3 | Malaysia Hoabinhian | Hoabinhian foragers (pre-Neolithic) | 0.1370 |
| 4 | China Yunnan EN Xingyi | Early Neolithic | 0.1491 |
| 5 | China Amur River 33000BP | c. 31050 BC | 0.1675 |
| 6 | China Tianyuan | Early Upper Palaeolithic | 0.1764 |
| 7 | India Uttarakhand Early Medieval Roopkund c.800CE (South Asian Profile) | c. 800 AD | 0.1863 |
| 8 | Mongolia EUP Salkhit | Early Upper Palaeolithic | 0.2045 |
| 9 | Tibet Early Modern Guge Kingdom Guge Ganshidong | Early Modern, c. 1500-1800 AD | 0.2080 |
| 10 | China Longlin 10500BP | c. 8550 BC | 0.2169 |
Ancient DNA from India
Our ancient DNA database holds 59 individuals excavated in present-day India, dated from about 2,350 BC to 1860 AD. The most frequent Y-DNA haplogroups among them are R2a (4), H1a (3) and J2b (3), and the most frequent mtDNA haplogroups are M3 (6), H1 (4) and U2 (4). These are people who lived on the same land in the past, not necessarily ancestors of today's Santhal population.
Most frequent Y-DNA haplogroups
Ancient DNA studies on India
- Blending borders: reconstructing the genetic history of the Sindhi population (2026)
- Blending borders: reconstructing the genetic history of the Sindhi population (2026)
- The Old Lady spider cave skeletons in Ladakh have diverse maternal genetic origin (2026)
- Layers in the sand: The genetic imprint of migration, culture, and Indus craft in the Thar desert (2026)
- Admixture and Genetic Connectivity: Autosomal Insights Into Indo-Aryan Speakers at the Eastern Edge of the Indian Subcontinent (2026)
- Novel 4400-year-old ancestral component in a tribe speaking a Dravidian language (2025)
- Ancient mitogenomes from Neolithic, megalithic and medieval burials suggest complex genetic history of Kashmir valley, India (2025)
- 50,000 years of evolutionary history of India: Impact on health and disease variation (2025)
Browse them in our ancient DNA database: 59 from India.
Compare yourself with Santhal
Paste your G25 coordinates (scaled, one line, with or without a name in front) and we compute your genetic distance to the Santhal average and to its closest neighbours, right in your browser. Nothing is uploaded or stored.
| # | Population | Your distance | Closeness |
|---|
This list only covers Santhal and its neighbours. To find out which of our reports actually fits your DNA, run the free Report Finder: it runs the same fit test that every report uses before an order.
No G25 coordinates yet? Get a free simulated G25 from your raw DNA file, or order G25 coordinates.
About these numbers
Keep in mind that a population average is a statistical summary of a limited number of sampled individuals. Real people inside any group vary, some carry more of one ancestry and some less, and no genetic profile decides who is or is not Santhal. Use these numbers as a map of deep ancestry, not as a label.
Method: averaged G25 coordinates, Euclidean distances to other modern and ancient averages, and a non-negative least-squares model against deep ancestral reference populations (data generated 2026-10-01). Read the full method.
Go deeper than the average
This page describes the Santhal average. For your own DNA, Advanced Ancestry Report (16€) gives a full ancestry breakdown with no regional focus, in a personal PDF report.
Comprehensive DNA analysis comparing your genetics to over 2,000 modern reference populations.
More populations from South Asia
Frequently asked questions
What is the ancient genetic make-up of Santhal?
Modelled with deep ancestral reference populations, the Santhal average is about 82.0% South Asian hunter-gatherer related (Irula proxy) and 18.0% Southeast Asian Neolithic farmers. These proportions are model estimates for a group average, not exact values for any one person.
Which populations are genetically closest to Santhal?
Leaving aside other regional samples listed under the same name in the G25 sheet (such as Santhal (Jharkhand)), the closest modern populations to the Santhal average are Bhumij (Odisha), Bhumij (Jharkhand) and Bhunjia (Chhattisgarh). The closest ancient matches are Andaman Islands 100BP (Great Andamanese), Laos Hoabinhian and Malaysia Hoabinhian.
Can a DNA test tell me if I am Santhal?
No DNA test can confirm an ethnicity or a nationality. What DNA can show is how similar your genome is to the sampled Santhal average and to its neighbours. If you have G25 coordinates, paste them in the comparison box on this page to check that for free.
Which ExploreYourDNA report suits Santhal ancestry?
Our regional report for this part of the world does not fit the Santhal average closely, so the full Advanced Ancestry Report (no regional focus) is the better choice. The DNA Mega-Analysis adds Neanderthal %, traits and closest ancient populations. The free Report Finder checks every report against your own DNA first.
Where does the data on this page come from?
From the averaged Global25 (G25) coordinates of the sampled individuals: genetic distances to other modern and ancient population averages, and a non-negative least-squares model against deep ancestral reference populations. The full method is described at https://www.exploreyourdna.com/populations#method.