At present, forensic kinship analysis mainly relies on fixed marker panels. However, Severe degradation often leads to limitations to current strategy. DNA fragments may be randomly lost, and sequencing depth was usually too low to genotype loci with confidence. All of these lead to a reduction in the number of available loci in this fixed panel, and the remaining markers may not provide sufficient power for kinship inference. Therefore, we developed a whole-genome, sample-specific, dynamic iterative marker selection strategy aiming to identify a relatively global optimal marker set for each sample. Genotype likelihoods were used to weight the likelihood ratio (LR) so that all possible genotypes were included in the probability calculation. In addition, to make fuller use of the genetic information in degraded samples, this study included both SNPs and MHs.
In this study, whole-genome sequencing (WGS) data from the target sample were used for genotyping. We analyzed all variant loci with minor allele frequency (MAF) > 0 in the Southern Han Chinese (CHS) population. We also analyzed all MH in a previously established whole-genome candidate MH dataset developed by our group. Genotype likelihoods were obtained for all genetic markers with GATK and an in-house MH genotyping software. A greedy algorithm was then applied for random iterative screening of all genotyped markers. The cumulative power of exclusion (CPE) was calculated in each iteration round. The selected markers had to meet two conditions. First, the genetic distance between each two markers had to be greater than a defined threshold. Second, the pairwise linkage disequilibrium had to satisfy r² < 0.2. Two stopping rules were tested, including a CPE threshold and a limit on iteration number or iteration time. Under each stopping rule, the marker set with the highest CPE was selected as the relatively global optimal marker set for that sample for kinship analysis. Based on the computational framework of Familias, the genotype likelihood of each genetic marker was incorporated into LR calculation as a weight.
In summary, this study is the first to propose a sample-specific dynamic iterative marker selection strategy for samples with different levels of degradation. Within a controllable time cost, this method can identify the most informative genome-wide markers for kinship analysis. By incorporating all possible genotypes into LR calculation, the method reduces the negative effects of sequencing errors and low sequencing depth on kinship inference.