The Malay homeland lies on both sides of the Strait of Malacca, on the Malay Peninsula and the east coast of Sumatra. The Sultanate of Malacca, a great trading port of the 15th century, helped spread Islam and the Malay language, which belongs to the Austronesian family and became a lingua franca of the archipelago. The Malay average comes out at 56.4% Southeast Asian Neolithic farmer, 37.3% Southern Chinese Neolithic farmer and 6.3% South Asian hunter-gatherer related in our model. The Southern Chinese Neolithic share, which in this region tracks the Austronesian expansion from Taiwan, is much higher than in the Javanese (12.5%), and the small South Asian-related share fits centuries of trade with India across the Bay of Bengal. Beyond other Malay samples, the closest modern averages are the Batak Toba of Sumatra (0.0266) and the Seletar, sea people of the Johor coast. The nearest ancient group is Early Neolithic Dushan in Guangxi, at 0.0432.

  • Ancient components in our model: Southeast Asian Neolithic farmers 56.4%, Southern Chinese Neolithic farmers 37.3%, South Asian hunter-gatherer related (Irula proxy) 6.3%
  • Closest modern population in our data: Malay (Singapore), distance 0.0125
  • Closest ancient group in our data: China Guangxi EN Dushan, distance 0.0432

Malay is one of the 335 modern populations in our population genetics atlas. The numbers on this page describe the average Global25 (G25) genome of the Malay individuals sampled in Malaysia. Like every population average, it smooths over a lot of variation between individuals.

Ancient make-up of the Malay average

Southeast Asian Neolithic farmers
56.4%
Southern Chinese Neolithic farmers
37.3%
South Asian hunter-gatherer related (Irula proxy)
6.3%

The ancient make-up of the Malay average is led by Southeast Asian Neolithic farmers (56.4%), followed by Southern Chinese Neolithic farmers (37.3%) and South Asian hunter-gatherer related (Irula proxy) (6.3%). The Southeast Asian Neolithic farmers component reflects the Neolithic farmers of mainland Southeast Asia. The Southern Chinese Neolithic farmers share (37.3%) reflects the Neolithic rice and millet farmers of southern China, whose descendants spread into Southeast Asia.

The model fits well (fit distance 0.024), so these proportions are a reasonable summary. This population was modelled with a global set of reference populations (ancient genomes, plus modern stand-ins where no suitable ancient genome exists), because West Eurasian sources alone cannot describe it.

Best-matching report
Lapita – Austronesian & Pacific Report

Lapita – Austronesian & Pacific Report

Fit test on the Malay average: Excellent fit · 97% regional core
See the report · 15€ Check your own DNA first, free

Closest modern populations to Malay

Genetically, the Malay average sits nearest to Malay (Singapore) at 0.013, Malay (Malaysia) at 0.018 and Batak Toba at 0.027. On our scale, a distance of 0.013 counts as very close. Leaving aside the other regional samples listed under the same name in the G25 sheet, the nearest other populations are Batak Toba (0.027), Seletar (0.034) and Balinese (0.037). By the tenth closest (Indonesian (Jakarta, Java/Bali Profile)) the distance grows to 0.046, a common pattern for a population that shares ancestry with its neighbours while keeping a profile of its own.

#PopulationDistanceCloseness
1 Malay (Singapore) 0.0125 Very close
2 Malay (Malaysia) 0.0183 Very close
3 Batak Toba 0.0266 Close
4 Seletar 0.0342 Close
5 Balinese 0.0366 Close
6 Sundanese 0.0414 Close
7 Nyah (Kur) 0.0433 Close
8 Cambodian 0.0439 Close
9 Khmer (Thailand) 0.0458 Close
10 Indonesian (Jakarta, Java/Bali Profile) 0.0461 Close

Distances are Euclidean distances between averaged G25 coordinates. On our scale, below 0.025 is very close, below 0.050 close, below 0.080 moderate, and beyond that distant.

Closest ancient populations to Malay

If we search ancient DNA for the best match to Malay, China Guangxi EN Dushan (Early Neolithic) comes first, at 0.043, with Vietnam LN Mai Da Dieu next. That is a close match: much of the modern profile was already present then, while later movements of people added further layers. A close ancient match is not proof of direct descent: it means those individuals carried a similar overall mix of ancestry.

#Ancient sample or groupPeriodDistance
1 China Guangxi EN Dushan Early Neolithic 0.0432
2 Vietnam LN Mai Da Dieu Late Neolithic 0.0477
3 Vietnam N Man Bac Neolithic 0.0486
4 Vietnam LN Ha Long Hhon Hai Co Tien Late Neolithic 0.0623
5 Liangdao N Neolithic 0.0642
6 China Liangdao2 N Neolithic 0.0644
7 Laos BA Tam Hang Bronze Age 0.0659
8 Vietnam LN Nam Tun Late Neolithic 0.0691
9 China Liangdao1 N Neolithic 0.0695
10 Indonesia Flores N-Proto-Metallic Liang Bua Neolithic 0.0722

Ancient DNA from Malaysia

Our ancient DNA database holds 7 individuals excavated in present-day Malaysia, dated from about 2,353 BC to 1950 AD. The most frequent Y-DNA haplogroups among them are D1a (1), O1a (1) and O1b (1), and the most frequent mtDNA haplogroups are M21 (2), B4 (1) and B5 (1). These are people who lived on the same land in the past, not necessarily ancestors of today's Malay population.

Most frequent Y-DNA haplogroups

HaplogroupIndividualsExamples
D1a 1 Ma911
O1a 1 Ma554
O1b 1 Ma912
O2a 1 Ma555

Most frequent mtDNA haplogroups

HaplogroupIndividualsExamples
M21 2 Ma911 JHF05
B4 1 Ma555
B5 1 Ma548
F3 1 Ma554
M13 1 Ma912
N9 1 Ma525

Ancient DNA studies on Malaysia

Browse them in our ancient DNA database: 7 from Malaysia.

Compare yourself with Malay

Paste your G25 coordinates (scaled, one line, with or without a name in front) and we compute your genetic distance to the Malay average and to its closest neighbours, right in your browser. Nothing is uploaded or stored.

No G25 coordinates yet? Get a free simulated G25 from your raw DNA file, or order G25 coordinates.

About these numbers

A word of caution: these figures describe the average of the Malay individuals who were sampled, not every person who identifies as Malay. Individuals vary around that average, and identity is about history, language and family, not about a genetic score. Read this page as a map of deep ancestry, not as a definition.

Method: averaged G25 coordinates, Euclidean distances to other modern and ancient averages, and a non-negative least-squares model against deep ancestral reference populations (data generated 2026-10-01). Read the full method.

Go deeper than the average

This page describes the Malay average. Your own DNA has its own story: Lapita – Austronesian & Pacific Report (15€) models your genome against the ancient sources and modern communities of the region, era by era, in a personal PDF report.

A report built for people of Austronesian / Indo-Pacific ancestry: ancient regional core (the Taiwan Neolithic homeland, the Lapita culture, Guam Unai and Latte periods, the Medieval voyage to Madagascar) versus non-regional outside...

Frequently asked questions

What is the ancient genetic make-up of Malay?

Modelled with deep ancestral reference populations, the Malay average is about 56.4% Southeast Asian Neolithic farmers, 37.3% Southern Chinese Neolithic farmers and 6.3% South Asian hunter-gatherer related (Irula proxy). These proportions are model estimates for a group average, not exact values for any one person.

Which populations are genetically closest to Malay?

Leaving aside other regional samples listed under the same name in the G25 sheet (such as Malay (Singapore)), the closest modern populations to the Malay average are Batak Toba, Seletar and Balinese. The closest ancient matches are China Guangxi EN Dushan, Vietnam LN Mai Da Dieu and Vietnam N Man Bac.

Can a DNA test tell me if I am Malay?

No DNA test can confirm an ethnicity or a nationality. What DNA can show is how similar your genome is to the sampled Malay average and to its neighbours. If you have G25 coordinates, paste them in the comparison box on this page to check that for free.

Which ExploreYourDNA report suits Malay ancestry?

For the Malay average, the regional report to choose is Lapita – Austronesian & Pacific Report, which our fit test rates as an excellent match. Your own DNA may differ from the average, so the free Report Finder checks every report against your file before you buy.

Where does the data on this page come from?

From the averaged Global25 (G25) coordinates of the sampled individuals: genetic distances to other modern and ancient population averages, and a non-negative least-squares model against deep ancestral reference populations. The full method is described at https://www.exploreyourdna.com/populations#method.

We use cookies to enhance your experience. By continuing to visit this site you agree to our use of cookies. Learn more

🧭 Find your report
📬 Stay in the loop

New tools and new reports, straight to your inbox. No spam, unsubscribe anytime.