The use of genome-wide SNP data, i.e., whole-genome sequencing (WGS) data, introduces new opportunities and challenges for forensic DNA mixture interpretation. While high-density SNP data provides substantial information content, existing probabilistic genotyping approaches are not optimized for efficient inference at this scale, nor for robust calibration under complex mixture scenarios.
In this study, we investigate deconvolution of WGS-based SNP mixtures (~500,000 markers) using an extension of the EuroForMix framework for massively parallel sequencing (MPS) data (EFMmps), specifically optimized for fast preprocessing and inference in high-dimensional settings. The framework further enables likelihood ratio (LR)-based inference for relatedness between contributors and persons of interest. To achieve computationally efficient and well-calibrated inference, we explore shrinkage-based extensions of marker-specific amplification efficiency (MAE) estimation, as shrinkage can stabilize parameter estimates, reduce overconfidence, and improve genotype posterior distributions for downstream inference.
A large in silico dataset is generated by combining real single-source WGS profiles to create two- and three-person mixtures across a range of mixture ratios and relatedness scenarios. This enables systematic evaluation under known ground truth. Performance is assessed using deconvolution accuracy, coverage, calibration, and the stability of LR-based inference across degrees of relatedness.
This work aims to provide a scalable and robust framework for WGS-based mixture interpretation, and to establish practical strategies for MAE estimation in ultra-high-dimensional SNP data, with particular emphasis on reliable relatedness inference in complex mixture settings.