Accurately inferring genetic relatedness from genomic data underpins a wide range of applications in human, plant, and conservation genetics, and it has become increasingly important in investigative genetic genealogy (IGG). Although high quality DNA sources routinely yield reliable identity by descent (IBD) segment detection, many real world forensic and population genetic samples originate from low-quality/quantity material. Telogen hairs, in particular, present substantial challenges due to degraded DNA, resulting in sparse genomic coverage and elevated genotyping error rates. These limitations complicate kinship inference and highlight the need for a systematic framework capable of quantifying how reduced DNA quality, quantity, and coverage influence relationship classification.
In this study, we address this gap through two complementary advances. First, we introduce a conditional simulation framework that generates relatives of multiple degrees using high quality donor genotypes as a reference. This approach enables controlled, repeatable evaluation of kinship inference methods under varying error profiles and data completeness. Second, we empirically assess whether short DNA fragments obtained from single telogen hairs can provide sufficient SNP information for IGG analyses. Using paired buccal swabs and rootless hairs collected from 50 donors available through dbGap (accession phs002979.v2.p1), we generated whole genome SNP profiles, augmented through genotype imputation, and compared these empirical profiles with simulated relatives across a range of relationship categories up to seventh degree.
Despite extremely low DNA yields from rootless hairs, imputation successfully recovered more than 60-80% of genotypes with concordance exceeding 99% relative to high quality reference profiles. Most genotyping errors reflected allelic dropout, and these errors were broadly and evenly distributed across the genome rather than concentrated in specific regions. Relationship inference analyses demonstrated that hair derived profiles generally enabled accurate kinship classification, although samples with higher error rates showed reduced accuracy, particularly for close relatives where IBD detection is most sensitive to incorrect genotypes. Importantly, tuning IBD detection parameters—such as increasing allowable error rates or adjusting minimum marker thresholds—improved classification performance without introducing false relationships.
Together, these findings show that even highly compromised DNA sources, such as single rootless hairs, can yield informative SNP profiles for IGG when supported by imputation and carefully optimized analytical parameters. Our conditional simulation framework further provides a robust foundation for evaluating kinship inference across diverse data quality scenarios.