Computational methods to discover, annotate, and genotype complex variation using long-read and biobank data - Project Summary The past five years have seen a dramatic increase in genomes sequenced, both in large-scale biobanks using short-read sequencing to map variation associated with traits and disease, and pangenome studies using long-read sequencing to assemble near complete reference genomes. During this period we contributed multiple novel approaches to discover and genotype variation in both long- and short-read data, including an algorithm to map variation in long-reads, lra; a method, danbing-tk, to map variation in tandem repeat sequences from short- read data and pangenome graphs, an algorithm; vamos, to detect tandem repeat variation in long-read sequences; and ctyper, an algorithm to genotype sequence-resolved copy-number variation using short-read sequencing and pangenomes. Each of these methods was applied collaboratively. Two notable features of variation are identified from long-read studies: the majority of structural variants: insertions, deletions, and rearrangements at least fifty bases, appear in short-tandem repeat, and variable number tandem repeat sequences, and the plurality of bases affected by structural variation are in large copy-number variants. The vamos method was created to solve the problem of aggregating tandem repeat variation across populations sequenced by long reads, and the ctyper algorithm was created to genotype previously inaccessible variation in copy-number variable genes. Because there are new disease cohorts sequenced using long-read sequencing, and because the majority of novel variation found by long read sequencing is in tandem repeat sequences, we will develop new approaches to curate tandem repeat variation in large long- read datasets, including discovering rare variants in tandem repeat loci, and new methods to compare tandem repeat variation across cohorts. These methods will be applied to test for associations between tandem repeat variation and dementia. We will additionally develop new methods to extend the ctyper algorithm to genotype variation from pangenomes genome-wide, in addition to sequence-resolved copy-number variation. This method is alignment-free, and extremely efficient, so that it may be applied to biobanks with over 400,000 genomes in a cost-effective manner. We will apply our approach to genotype variation in the UK Biobank, and replicate genotyping in the All of Us, each with nearly half a million samples. Finally, we will develop novel approaches to associate complex, phased variation with traits and disease.