Using genetic information to investigate the way in which two individuals may be related is an important application of forensic DNA typing. Pairwise kinship assessments will often compare the probability of observing the individuals’ genotype data under two competing relationship hypotheses, for example they are related as full-siblings or they are unrelated (i.e. any alleles that may be shared, are shared by chance). The result of such a kinship evaluation is presented as a likelihood ratio (LR), whereby the greater the LR value, the stronger the statistical support for a chosen relationship hypothesis.
It can be helpful to understand the expected range and distribution of LR values based on the specifics of a given case. In many forensic disciplines, such evaluations can be made based on “ground truth” data, however, this is often hard to come by in a kinship context. Therefore, a commonly used method for estimating the expected likelihood ratios for kinship scenarios involves simulating genotype data. These simulations can be conditioned on the allele frequencies of typed loci, the pedigree hypotheses compared and other relevant parameters such as recombination rates between linked markers. The distribution of the expected LR values can demonstrate the general capability of targeted genetic loci and/or analysis methodologies for resolving such relationship cases. Overall, this can provide important information, either to aid in deciding what markers to use in a given case or giving context to an LR value when reporting results. A simulation-based approach, however, relies upon the assumption that the simulated case genotypes and subsequently calculated LRs are representative of real cases.
This work aimed to explore the factors that influence the accuracy of simulated likelihood ratios, by comparing the distribution of theoretical LRs with the actual distribution of LRs generated from a large number of known sibling pairs. The impact of multiple calculation parameters was assessed. In particular, the results highlighted that the concordance between the simulated and real-life data sets was greatly influenced by the population-of-origin of the sibling pairs, and the corresponding population substructure. Additionally, accounting for genetic linkage when using expanded marker sets was found to be notably important.
Overall, understanding how representative simulated LR distributions are of genuine case LRs is an important and under-studied aspect of kinship analysis. Given that these simulated datasets are frequently used for evaluations in multiple contexts, it is crucial to understand the reliability of using this approach.