Title

P219 – Comparison of Error Correction Methods for UMI-based Microhaplotype Analysis in DNA Mixtures

11:01
Thursday August 20th
Station 12
Duration: 12 minutes 
10. NGS & SNPs
Ye-Lim Kwon

Microhaplotypes, defined as multi-SNP markers within short genomic segments, have surfaced as a promising alternative for forensic DNA mixture analysis using massively parallel sequencing (MPS). However, the interpretation of mixed DNA remains challenging due to background noise from PCR and sequencing errors. While unique molecular identifiers (UMIs) provide a robust solution for molecule-level error correction, residual noise often persists, necessitating additional computational error correction strategies. Recently, UMI-based error correction tools (e.g., UMIc and UMI-nea) have been developed to correct UMI sequence errors using edit distance-based algorithms without reference alignment; however, their efficacy in forensic contexts requires further validation. In this study, we established UMI error correction pipelines based on UMIc and UMI-nea, respectively, and evaluated their performance against the conventional FGBIO toolkit. Thirty-two microhaplotypes were analyzed across highly imbalanced mixtures (two to four contributors; 5 ng and 1 ng DNA input), where the minor contributor accounted for 0.5% to 10% of the total DNA. The MPS library was generated by attaching UMI barcodes via PCR and subsequently sequenced on the Illumina NextSeq platform. The performance of error correction was evaluated based on the number of UMI families, observed noise levels, and likelihood ratios (LR) calculated through EuroForMix. As a result, the number of UMI families per marker remained consistent across all methods, averaging 1,050 for 5 ng and 249 for 1 ng, indicates stable molecular resolution. No significant differences in noise levels were observed between the methods; however, at 5 ng input, a 2.1% noise level in FGBIO was effectively eliminated in UMIc and UMI-nea. Consequently, UMI error correction tools yielded higher log10LR values than the conventional FGBIO toolkit in most 5 ng DNA mixtures. In conclusion, the alignment-free UMI error correction pipelines established in this study provide a streamlined, highly precise alternative to conventional tool for microhaplotype analysis of DNA mixtures.

Authors

  • Ye-Lim Kwon (Yonsei University College of Medicine, South Korea)
  • Kyoung-Jin Shin (Yonsei University College of Medicine, South Korea)

On the same topic