Title

P169 – Machine Learning–Based Estimation of the Number of Contributors in Y‑STR Mixtures

11:01
Thursday August 20th
Station 02
Duration: 12 minutes 
02. PGS
Shota Inokuchi

Estimating the number of contributors (NoC) in forensic DNA mixtures is a fundamental step in mixture interpretation and evaluation of evidential weight. While probabilistic genotyping has significantly advanced autosomal STR analysis, robust approaches for NoC estimation in Y-chromosomal STR (Y-STR) mixtures remain limited due to their haploid inheritance and lack of heterozygosity. Conventional threshold-based methods, such as maximum allele count (MAC) and total allele count (TAC), are simple and interpretable but often show limited accuracy for complex mixtures. Although machine learning (ML) has demonstrated strong potential for autosomal STR mixtures, its application to Y-STR NoC estimation has been largely unexplored. In this study, we evaluated ten widely used ML algorithms for estimating NoC in Y-STR profiles and directly compared their performance with conventional approaches using identical datasets. A total of 120,000 simulated Y-STR profiles per population (NIST and Henan Han) were generated from haplotype frequency data, covering 1–6 contributor mixtures across the 27 loci included in the Yfiler™ Plus kit. Each dataset was randomly divided into training and test sets with balanced class representation. Fifty-two allele count–based features were used while excluding peak height and frequency-based variables to ensure general applicability. All ML models consistently outperformed conventional MAC and TAC methods. The highest overall accuracy reached 0.8665 (SVC) in the NIST population and 0.8996 (XGB) in the Henan Han population, compared with 0.7920 and 0.8385 obtained using conventional approaches. Tree-based models demonstrated particularly strong and stable performance across both populations. Given the importance of interpretability in forensic applications, decision tree models were further examined. The final models achieved accuracies of 0.8572 (NIST) and 0.8882 (Henan Han) and produced transparent decision structures. Notably, accurate predictions were achieved using only two features—MAC and TAC. Learning curve analysis indicated rapid convergence and stable generalization. These results demonstrate that ML substantially improves the accuracy and robustness of NoC estimation for Y-STR mixtures while preserving interpretability, supporting the integration of interpretable ML approaches into forensic Y-STR mixture interpretation.

Authors

  • Shota Inokuchi (Department of Forensic Medicine, Graduate School of Medicine, Juntendo University, Japan)
  • Tetsuya Sato (Forensic Science Laboratory, Kumamoto Prefectural Police Headquarters, Japan)
  • Hiroaki Nakanishi (Department of Forensic Medicine, Graduate School of Medicine, Juntendo University, Japan)
  • Aya Takada (Department of Forensic Medicine, Saitama Medical University, Japan)
  • Kazuyuki Saito (Department of Forensic Medicine, Graduate School of Medicine, Juntendo University, Japan)

On the same topic