Unlocking Pediatric Genomic Variation: A Proposal to Index and Genotype the Kids First Cohort with The Great Genotyper - Project Summary Structural variants (SVs) are a major yet underexplored source of genetic diversity and disease in children, contributing to congenital heart defects, neurodevelopmental syndromes, and pediatric cancers. However, short-read sequencing data, commonly used in large-scale studies, have limited sensitivity for SV discovery, and current genotyping tools are too slow to scale to tens of thousands of genomes. The Gabriella Miller Kids First (KF) program includes over 30,000 whole-genome sequencing (WGS) datasets, representing an unparalleled opportunity to study pediatric disease genetics, yet the true SV landscape of this cohort remains largely uncharacterized. We recently developed The Great Genotyper (TGG), a graph-based population genotyping framework that enables efficient representation, compression, and querying of genetic variants at unprecedented scale. TGG uses a counting colored de Bruijn graph (CCDG) to encode thousands of genomes with >200× compression and near real-time querying, allowing rapid allele-frequency and genotype analyses directly from indexed data. In this project, we will extend TGG to build the first queryable pediatric genomic graph by indexing the entire Kids First WGS cohort. In Aim 1, we will perform alignment-free quality control of all samples using a novel alignment-free tool, then construct a highly compressed, graph-based index of the KF repository. In Aim 2, we will genotype millions of curated SVs from dbVar and human pangenomes across all samples, validate results using long-read data, and produce a fully annotated, cohort-scale SV genotype matrix. In Aim 3, we will develop a secure, cloud-based web portal that enables researchers to query SV frequencies and genotype distributions in real time, integrated with the Kids First Data Resource Center (KFDRC). This project will transform static pediatric genomic data into a dynamic, interactive, and reusable resource for the research community. By combining scalable graph genomics, efficient data compression, and open-access infrastructure, our work will (1) enables fast, scalable genotyping of known SVs in existing short-read data, (2) supports incremental updates when new WGS datasets are added to the repository or SV catalogs emerge, and (3) provides real-time allele frequency lookup for novel variants. The resulting resources will accelerate the discovery of pathogenic SVs, improve variant interpretation, advance our understanding of the genetic architecture of childhood diseases, and maximize the translational impact of the Kids First program.