Forensic genetics is increasingly challenged by trace evidence yielding only fragmented, low-template DNA. To utilize such compromised samples, forensic profiling has transitioned from traditional Short Tandem Repeat (STR) CE-analysis toward Massively Parallel Sequencing (MPS) for human identification (HID). Single Nucleotide Polymorphisms (SNPs) have emerged as superior markers for analysing highly degraded and low-coverage traces, offering a viable alternative when STR profiling fails. However, the shift toward genome-wide approaches, e.g. whole genome sequencing (WGS) or targeted capture data, necessitates robust analytical frameworks capable of processing sparse sequencing data to produce reliable weight of evidence (WoE) calculations.
To address these challenges, the SNP-Based Forensic Identification Tools (SNP-FIT) consortium was established to describe methods and to evaluate SNP analyses for HID from low-coverage WGS data.
In this study, the SNP-FIT consortium systematically investigated six computational tools: VerifyBamID, IBDGem, FamLink2 (lcNGS), wgsLR, EuroForMix, and DBLR™. Using data from the 1000 Genomes Project, downsampled via the snp2id pipeline, we evaluated HID performance across ten sequencing depths ranging from 0.001x to 10x, with 30x coverage as the ground truth (reference sample). Because low-coverage data results in an incomplete and stochastic breadth and depth of coverage, the available data may vary from sample to sample, and we propose a framework for dynamic SNP selection ensuring that only statistically independent markers contribute to the final likelihood ratios.
Preliminary results from both donor and non-donor comparisons demonstrated that all trace samples with ≥1x coverage were assigned an evidential weight as expected by all analysed tools, suggesting a universal minimum coverage threshold for reliable HID analysis. To confirm these findings on forensic relevant material, we analysed rootless hair samples as an example of a challenging, low-quality trace sample. These hair profiles were compared against high-quality reference buccal swabs from both donors and non-donors. We further discuss the comparative performance of hard-called genotyping versus probabilistic genotyping.
In conclusion, this research provides an evaluation of different tools and a methodology for implementing SNP-based HID in low-coverage WGS data in forensic casework. By exploring low depth sequencing data in combination with a dynamic SNP selection procedure, the SNP-FIT consortium offers a path forward for interpreting challenging, low-coverage genomic data in a forensic context.