Skip to main content
An official website of the United States government
Grant Details

Grant Number: 1DP2CA325380-01 Interpret this number
Primary Investigator: Sakaue, Saori
Organization: University Of Washington
Project Title: Illuminating the Hidden Genome: Population-Scale Inference of Human Centromeric Variation and Its Clinical Sequelae
Fiscal Year: 2026


Abstract

Project Summary Human centromeres represent the last frontier of genomic dark matter. In each cell division, centromeres recruit the kinetochore to the chromosome, which attaches to the mitotic spindles to ensure accurate separation of sister chromatids. Defects in centromere function can lead to chromosome missegregation and genome instability. Paradoxically, centromeres are also among the fastest-evolving loci in the genome. This unique combination of functional indispensability and rapid evolution makes centromeres strong candidates for influencing human traits and disease susceptibility. Indeed, inter-individual variation in centromeric satellite DNA has been implicated in key traits including aneuploidy, infertility, and cancer since 1990s using small cohorts. Despite their importance, centromeres remain among the least characterized genomic regions. Conventional reference genomes (GRCh38) leave centromeres structurally unresolved with abundant near-identical yet highly complex repeat sequences. Due to the lack of representation in the reference genome, modern genome-wide association studies (GWAS) have entirely excluded centromeric variations. Consequently, we lack the landscape of centromeric variation to human health and diseases. To address this gap, we will significantly scale-up the catalog of human centromeric variations to millions of individuals by developing statistical inference algorithms that enable imputation of centromeric variations and robust association studies with human traits and diseases. We will construct a haplotype-resolved reference panel leveraging recently developed telomere-to-telomere long-read centromere assemblies and use it as a basis for computational inference of centromeric variation from whole-genome sequence data. We will apply these methods to existing large-scale biobanks with millions of genomes and deep phenotype data and perform phenome-wide association studies to comprehensively characterize clinical sequalae of centromeric variations. In addition, we will develop computational algorithms to detect somatic copy number alterations involving centromeres in tumor tissues and profile their non-coding RNA expression. We will thereby investigate the links between centromeric somatic mosaicism, genome instability, transcriptional desilencing, aging, and cancer. This project can transform the conventional paradigm of centromere research by bringing centromere back into the state-of-the-art population-scale genomics—recasting these regions from unsequenceable gaps into disease-relevant loci with novel insights into clinical practice and biology.



Publications


None

Back to Top