Title

P255 – Development of a Genome-wide Screening Pipeline for Ancestry-informative MHs and Rare MHs in East Asian Populations

11:13
Thursday August 20th
Station 19
Duration: 12 minutes 
08. Phenotyping
Meili Lv

As a crucial method for individual identification, forensic ancestry inference aims to determine the biogeographical origin of unknown individuals by analyzing ancestry informative markers (AIMs) in comparison with reference population data. Microhaplotypes (MHs), as a new type of compound genetic markers, possess several advantages, including multiallelism, balanced amplification, and high levels of polymorphism, making them increasingly valuable in forensic ancestry inference. The current studies have predominantly focused on intercontinental populations, and there remains a lack of systematic genome-wide investigations of MHs specifically in East Asian subpopulations. In this study, we established a complete framework suitable for ancestry inference of East Asian subpopulations based on genome-wide MHs. Traditional parameter thresholds (In and FST) were applied to screen candidate loci, and loci with undesirable characteristics were removed. A feature extraction strategy was applied to identify 250-bp ancestry-informative microhaplotypes (AIM-MHs) using the AIM-SNPtag approach proposed by Zhao et al, with Jensen–Shannon (JS) divergence as the classification ability (CA) and normalized mutual information (NMI) as the penalty term. MHs were ultimately obtained. To further enhance the ancestry inference performance of the panel, we innovatively introduced rare MHs for the first time. These were defined as MHs harboring population-unique haplotypes. Based on our previously developed MH screening algorithm and 1000 Genomes Project Phase 3 dataset, the top 10 most frequent rare MHs within 250bp were selected for each of Han Chinese in Beijing (CHB), Han Chinese South (CHS), and Japanese in Tokyo (JPT), enabling improved discrimination among these closely related populations. STRUCTURE analysis and four machine learning models, including eXtreme Gradient Boosting (XGBoost), Random Forest (RF), Partial Least Squares Discriminant Analysis (PLS-DA), and Support Vector Machine (SVM), were performed based on the combined set of 451 AIM-MHs and 30 rare MHs. The results indicated that the overall model accuracy exceeded 90%, and the inclusion of rare MHs further improve classification performance to a certain extent. Overall, this study demonstrates that the constructed AIM-MHs panel incorporating rare MHs represents a robust and effective tool for high-precision ancestry differentiation among East Asian subpopulations.

Authors

  • Meili Lv (West China School of Basic Medical Sciences and Forensic Medicine, Sichuan University, China)
  • Shanshan Wei (West China School of Basic Medical Sciences and Forensic Medicine, Sichuan University, China)
  • Jiaming Xue (West China School of Basic Medical Sciences and Forensic Medicine, Sichuan University, China)
  • Mengyu Tan (West China School of Basic Medical Sciences and Forensic Medicine, Sichuan University, China)
  • Qiushuo Wu (West China School of Basic Medical Sciences and Forensic Medicine, Sichuan University, China)
  • Shengqiu Qu (West China School of Basic Medical Sciences and Forensic Medicine, Sichuan University, China)
  • Weibo Liang (West China School of Basic Medical Sciences and Forensic Medicine, Sichuan University, China)

On the same topic