Human identification in forensic genetics relies predominantly on short tandem repeats (STR) profiling. However, this approach captures only a tiny fraction of inter-individual genetic variation, leaving much of the non-coding and repetitive genome unexplored. Repetitive DNA, including tandem repeats (TRs) and transposable elements (TEs), comprises the majority of the human genome and represents a rich but largely untapped source of genetic and epigenetic variation with potential forensic relevance. Recent advances in long-read sequencing technologies, such as Oxford Nanopore Technologies (ONT), enable direct and simultaneous interrogation of sequence variation and DNA methylation at single-molecule resolution. These developments offer new opportunities to characterize (long and complex) repetitive regions that have historically remained inaccessible using short-read sequencing and array-based approaches. However, an integrated computational framework for joint (epi)genome-wide analysis of repeat variation from ONT data is currently lacking.
Here we present ECHO, a comprehensive Snakemake-based pipeline for the (Epi)genomic Characterisation of Human Repetitive Elements using ONT Sequencing. The pipeline enables joint analysis of major TR (STRs/VNTRs) and TE (LINEs/SINEs) classes, generating genome-wide, locus-specific, haplotype-resolved repeat profiles that integrate sequence variation with DNA methylation information across previously inaccessible regions. Our approach leverages the latest genome-wide repeat annotation catalogues and publicly available Genome-In-A-Bottle ONT datasets, and integrates specialized repeat analysis tools with custom methods to incorporate DNA methylation profiles. Using the HG002 diploid genome benchmark and independent whole-genome bisulfite sequencing data, we showed that ECHO accurately and robustly profiles both repeat sequence and CpG methylation levels across diverse repetitive loci.
These combined repetitive (epi)genomic signatures provide a new foundation for forensic applications, including for improved individual identification, tissue/cell type differentiation, and investigation of relevant personal phenotypes (e.g. ageing). Future work will focus on applying ECHO to larger, population-scale datasets spanning multiple cell types to also characterize potential somatic variation. Overall, our work contributes to the emerging field of human repeatomics and highlights the potential of long-read (epi)genomics to expand the forensic genetics toolkit.