The Emerging Limbs and Twigs of the East Asian mtDNA Tree
Toomas Kivisild, Helle-Viivi Tolk, Jüri Parik, Yiming Wang, Surinder S Papiha, Hans-Jürgen Bandelt, Richard Villems*
*Department of Evolutionary Biology, Tartu University and Estonian Biocentre, Estonia;
Department of Medical Genetics, Sun Yat-Sen University of Medical Sciences, People's Republic of China;
Department of Human Genetics, University of Newcastle-upon-Tyne;
Department of Mathematics, University of Hamburg, Germany
We determine the phylogenetic backbone of the East Asian mtDNA tree by using published complete mtDNA sequences and assessing both coding and control region variation in 69 Han individuals from southern China. This approach assists in interpreting published mtDNA data on East Asians based on either control region sequencing or restriction fragment length polymorphism (RFLP) typing. Our results confirm that the East Asian mtDNA pool is locally region-specific and completely covered by the two subhaplogroups, M and N. The phylogenetic partitioning based on complete mtDNA sequences corroborates existing RFLP-based classification of Asian mtDNA types and supports the distinction between northern and southern populations. We describe new haplogroups M7, M8, M9, N9, and R9 and demonstrate by way of example that hierarchically subdividing the major branches of the mtDNA tree aids in recognizing the settlement processes of any particular region in an appropriate time scale. This is illustrated by the characteristically southern distribution of haplogroup M7 in East Asia. In contrast, its daughter-groups, M7a and M7b2, specific for Japanese and Korean populations, testify to a presumably (pre-)Jomon contribution to the modern mtDNA pool of Japan.
Fig. 3.—Phylogenetic reconstruction and geographic distribution of haplogroup M7. a, A network of HVS-I haplotypes comprises the superposition of the most parsimonious trees for the three postulated sets of M7a, M7b, and M7c sequences. The mutations along the bold links were only analyzed for a few Japanese sequences (Ozawa et al. 1991 ; Ozawa 1995 ; Nishino et al. 1996 ) and—toward the root of M—for some Chinese sequences (this study): the corresponding individuals with (partial) coding region information are boxed. Numbers along links indicate transitions; recurrent HVS-I mutations are underlined. The age of mtDNA clades is calculated (along the tree indicated by unbroken lines) according to Forster et al. (1996) , with standard errors estimated as in Saillard et al. (2000) . Sample codes (and sources): AI—Ainu (Horai et al. 1996 ); CH—Chinese (Betty et al. 1996 ; Nishimaki et al. 1999 ; Qian et al. 2001 ; Yao et al. 2002 ; this study); IN—Indonesian (Redd and Stoneking 1999 ); JP—Japanese (Ozawa et al. 1991 ; Ozawa 1995 ; Horai et al. 1996 ; Nishino et al. 1996 ; Seo et al. 1998 ; Nishimaki et al. 1999 ); KN—Koreans (Horai et al. 1996 ; Lee et al. 1997 ; Pfeiffer et al. 1998 ); MA—Mansi (Derbeneva et al. 2002 ); MJ—Majuro (Sykes et al. 1995 ); MO—Mongolians (Kolman, Sambuughin, and Bermingham 1996 ); PH—Philippines (Sykes et al. 1995 ; Maca-Meyer 2001 ); RY—Ryukyuans (Horai et al. 1996 ); SB—Sabah (Sykes et al. 1995 ); TW—Taiwanese Han (Horai et al. 1996 ) and aboriginals (Melton et al. 1998 ); UI—Uighur (Comas et al. 1998 ; Yao et al. 2000 ); YA—Yakuts (Derenko and Shields 1997 ). b, Frequencies of the subgroups of M7 in Asian populations are inferred from the preceding HVS-I as well as partial HVS-I and RFLP data (VN—Vietnamese: Ballinger et al. 1992 ; Lum et al. 1998 ). Mainland Han Chinese are denoted as follows: GD—Guangdong, LN—Liaoning, QD—Qingdao, WH—Wuhan, XJ—Xinjiang, YU—Yunnan (Yao et al. 2002 ), SH—Shanghai (Nishimaki et al. 1999 ). The number of M7 sequences in relation to the sample size is indicated under each pie slice proportional to the M7 frequency
Fig. 2. Frequency distributions of the eight Y-chromosome haplotypes for the 14 global populations, with their approximate geographic locations. The frequencies of the eight haplotypes are shown as coloured pie charts (for colour codes, see upper left insert). JP Japanese
Only four Japanese populations exhibited ht1 (defined only by YAP+) at various frequencies (also see Table 1). The highest frequency (87.5%) was found in JP-Ainu, followed by JP-
Okinawa (55.6%) living in the southwestern islands of Japan, JP-Honshu (36.6%), and JP-
Kyushu (27.9%). The ht2 haplotype (defined by YAP+/M15+) was found in only two males, one each from Thais and Thai-Khmers; ht3 (defined by YAP+/SRY4064-A) was completely absent in the Asian populations examined, whereas Jewish in the Uzbekistan and African populations had this haplotype with a frequency of 28.3% and 100%, respectively. Thus, the YAP+ lineage was found in restricted populations among Asian populations, consistent with previous reports (Hammer and Horai 1995; Hammer et al. 1997; Shinka et al. 1999).
The ht4 haplotype (defined only by M9-G) was widely distributed among north, east, and southeast Asian populations, except for the Ainu. This haplotype was frequent (60.5%) in overall Asian populations (Table 1). Among them, the Han Chinese and southeast Asian populations were characterized by high frequencies ranging from 81.0% to 96.0%. In contrast to ht4, ht5 (defined by M9-G/DYS257108-A) and ht6 (defined by M9-G/DYS257108-A/SRY10831-A) were small contributors to Asian populations. The highest frequency of ht5 was observed in Nivkhi (19.0%) and that of the ht6 in Thai-Khmers (10.8%). The ht5 haplotype is widely distributed among European, Asian, and Native American populations and is proposed to be one of the candidates for founder haplotypes in the Americas (Karafet et al. 1999). Furthermore, high frequencies of ht6 were observed in north Europe, central Asia, and India (Karafet et al. 1999). Thus, the presence of ht5 in Nivkhi may account for the founder effect of peopling of the Americas.
The ht7 haplotype (defined by RPS4Y-T) was also widely distributed throughout Asia with the exceptions of Malaysia and the Philippines, whereas this was absent in two non-Asian populations. The highest frequency of ht7 was found in Buryats (83.6%), followed by Nivkhi (38.1%). Thus, the geographic distribution of ht7 in Asia appears to contrast with that of ht4.
Only eight individuals (1.4%) in Asia belonged to ht8, which was the major haplotype in Jewish population (Table 1). The ht8 haplotype may not be useful for inferring population relatedness among Asian populations because it is defined by no mutations. Additional Y-polymorphic markers such as M89 and M168 (Underhill et al. 2000; Ke et al. 2001) will be needed to investigate details of the formation of modern Asian populations.