Title

P020 – A Unique Sequence Identifier Detected in STR and SNP Raw Reads Generated Using the ForenSeq™ DNA Signature Prep Kit

11:13
Wednesday August 19th
Station 04
Duration: 12 minutes 
05. STR typing
Chiara Saccardo

In 1997, the ISFH DNA Commission complained that the enormous increase in sequence and substructure data for many loci repeatedly raised nomenclature issues, underscoring the need to establish a standardized, unambiguous nomenclature to ensure reproducibility, comparability, and the interlaboratory exchange of genetic data.

Over the years, STR analysis by capillary electrophoresis (CE) became the analytical gold standard in forensic genetics, supported by consolidated knowledge of the features and number of STRs analyzed, and the advantages, limitations, and pitfalls of CE technology.

To expand information on commonly used markers not deducible from CE analysis, massively parallel sequencing (MPS) has been introduced in forensic genetics over the past decade. Detection of sequence variations within the marker amplicon is the defining feature of MPS, enabling the identification of isoalleles and microhaplotypes, thereby increasing discriminatory power and strengthening genetic evidence. 

However, as occurred for STR-CE analysis, sequences generated by MPS are heterogeneous, mainly due to different MPS kits and data analysis software that generate sequences of variable nucleotide lengths, potentially limiting the detection of sequence variations.

To harmonize MPS sequence data, in 2024, the ISFG DNA Commission considered it appropriate to redefine the minimum nucleotide sequence range to facilitate the detection of variations also present in flanking regions adjacent to the target locus.

In an effort to follow the DNA Commission recommendations and to bypass limitations imposed by MPS commercial kits and software, each 351 bp raw read within the R1 FASTQ files generated by Universal Analysis Software for the 232 DNA samples sequenced on the MiSeq FGx™ system using the ForenSeq™ DNA Signature Prep kit was manually inspected.

Regardless of marker type, a unique 32 bp sequence was identified in every read, which, as it does not align with the human reference genome reported in Genome Browser, suggests a synthetic origin.

This sequence, interposed between the sample index adapters and the locus-specific sequence, allowed determination within each 351 bp raw read of the actual amplicon size for each marker, including, beyond the target region, the flanking regions and primer sequences.

This approach extended the analyzable sequence range beyond that achievable through UAS and STRait Razor v3 analysis, enabling more comprehensive comparison with Forensic Sequence STRucture Guide v6.1 beta and Genome Browser reference sequences and revealing additional sequence variations that would otherwise remain undetected.

Authors

  • Chiara Saccardo (Department of Diagnostics and Public Health, Section of Forensic Medicine, University of Verona, Italy)
  • Stefania Turrina (Department of Diagnostics and Public Health, Section of Forensic Medicine, University of Verona, Italy)

On the same topic