Investigative genetic genealogy (IGG) approaches for lead generation typically proceed from single nucleotide polymorphism (SNP) profiles that can be uploaded to a suitable database for matching. The development of SNP profiles for IGG can be approached in numerous ways. Wet lab processing options include whole genome sequencing (WGS) as well as approaches such as microarrays, PCR-based assays, and hybridization capture that target a defined set of markers. For next generation sequencing (NGS) based methodologies, numerous library preparation procedures are available and different chemistries and platforms may be used for sequencing. Few of these options have associated software packages that can be used to produce the IGG SNP genotypes for database queries. As a result, custom analysis pipelines are often developed, which involves both coding and selecting software packages from a variety of publicly available tools.
Forensic laboratories deliberating these options for SNP profile generation may gravitate toward a particular workflow based on their prior experience (e.g., with specific library preparation kits or assay types), instrument availability (e.g., for microarray typing, or sequencing at low versus high throughput), and bioinformatics resources, among other factors. Further, given the level of effort required to operationalize SNP typing for IGG, individual laboratories will likely focus on just a single workflow at the outset. However, considering the diversity of sample types and conditions that may be encountered in forensic casework, it is probable that no one workflow will be ideal for all circumstances.
We have undertaken a collaborative study to evaluate the performance of different workflows that aim to genotype the same approximately 1.3 million SNPs for IGG. We will present results comparing WGS to hybridization capture using a custom probe panel, different library preparation methods, custom analysis pipelines independently developed by different laboratories, and direct genotype calling versus imputation. Forensically relevant sample types tested included saliva, semen, and bone, and extracts with both damaged and degraded DNA of varying quantities. A variety of metrics were used to assess outcomes, including (where possible) concordance with known reference profiles. The results of the study will assist labs in method selection depending on their resources and expected casework.
This study was supported by a Peter M. Schneider ISFG Fellowship. This work was funded under Contract No. HSHQDC-15-C-00064, awarded by the DHS S&T to NBACC, a DHS federal laboratory operated by BNBI. Views and conclusions contained herein are those of the authors.