Machine learning and statistical tools for subcellular spatial biology - PROJECT SUMMARY Spatial omics is the new frontier in biotechnology – a series of innovations over the last ten years that give us exquisitely detailed views into molecular events and interactions inside cells, across all cells in a tissue sample. Some of these technologies can reveal a complete map of gene transcripts inside each cell and such “subcellular spatial transcriptomics” (SST) technology has immense and widely recognized potential for biomedical applications. Yet, current uses of this technology typically aggregate the available information at the level of an entire cell, rarely exploring the richness of subcellular information available from the assay. This project's goal is to develop a comprehensive toolkit for analyzing subcellular spatial transcriptomics (SST) data, extracting interpretable biological patterns and testable mechanistic insights into tissue function and pathology. The proposed approach will employ innovative spatial analysis techniques, leveraging state-of-the-art machine learning methods and robust statistical procedures. A major thrust will be on identifying subcellular spatial patterns involving individual genes, gene pairs and modules of genes, while being aware of biological variations from cell to cell. A new functionality in the toolkit will be to quantify changes in genes' subcellular distribution patterns between conditions, paving the way to a novel class of biomarkers. Planned approaches will build on recent publications from the PI's laboratory, improving the statistical power and scalability of state-of-the-art tools and exploring complementary modeling techniques. Another major goal will be to describe the subcellular space in useful ways, such as partitioning a cell's landscape into functionally distinct components, annotating axons and dendrites in brain data, and representing each cell's spatial transcriptome in a format that lends itself to machine learning algorithms. Tools developed for this goal will facilitate more accurate discovery of interpretable spatial patterns, charting of intercellular communication in brain SST data, and machine learning-based characterization of cells, ultimately leading to new ways of describing disease and biological conditions. The third plank of the proposed project is to discover how functional patterns at the subcellular level are encoded in gene sequences. For this task, machine learning tools will be implemented that relate gene sequence patterns to gene transcript distribution inside cells, and the discovered sequence patterns will then point to key regulators of those genes, thus providing potential targets for intervention. All functionalities of the proposed toolkit will be subjected to rigorous testing for robustness and reproducibility, and then applied to SST data sets from diverse biological systems, demonstrating their real-world utility. Furthermore, special attention will be given to software and data sharing, through adherence to “FAIR” (findable, accessible, interoperable, reusable) principles popularized by the NIH. This project will not only establish SST analytics on a firm footing, it will also generalize to other “omics” assays of subcellular resolution, that are under development today.