Recovering individual genetic information from low-coverage or mixed sequencing data remains a central challenge in forensic genomics, where DNA is often degraded, present in trace amounts, and derived from multiple contributors. Under these conditions, standard genotype-based approaches can be difficult to apply, particularly in complex mixtures such as touch DNA or highly degraded material, where allelic dropout and stochastic effects limit interpretability. In this context, we explored whether heterozygosity levels could provide an indication of mixtures in degraded DNA, based on the expectation that composite samples may exhibit an excess of heterozygous sites relative to a single individual, even at low coverage.
Building on this, we are developing an unsupervised read-clustering framework aimed at disentangling DNA contributed by multiple individuals in mixed samples. The method operates directly on sequencing reads and leverages linkage information within reads and overlapping fragments to jointly support haplotype phasing and mixture separation. Through probabilistic pattern resolution, reads are grouped by their haplotypic origin, enabling the integration of weak signals across loci under challenging conditions.
In forensic contexts, this read-level strategy offers a complementary route from mixture detection toward contributor separation, moving beyond genotype-level summaries to make use of haplotypic structure. The same framework is also applicable to archaeogenomic datasets, including sedimentary ancient DNA and commingled remains, where DNA from multiple individuals is highly fragmented, and may enable recovery of individual-level genetic signals in contexts where this has previously been difficult.