Grant Details
| Grant Number: |
1DP2CA325380-01 Interpret this number |
| Primary Investigator: |
Sakaue, Saori |
| Organization: |
University Of Washington |
| Project Title: |
Illuminating the Hidden Genome: Population-Scale Inference of Human Centromeric Variation and Its Clinical Sequelae |
| Fiscal Year: |
2026 |
Abstract
Project Summary
Human centromeres represent the last frontier of genomic dark matter. In each cell division, centromeres recruit
the kinetochore to the chromosome, which attaches to the mitotic spindles to ensure accurate separation of
sister chromatids. Defects in centromere function can lead to chromosome missegregation and genome
instability. Paradoxically, centromeres are also among the fastest-evolving loci in the genome. This unique
combination of functional indispensability and rapid evolution makes centromeres strong candidates for
influencing human traits and disease susceptibility. Indeed, inter-individual variation in centromeric satellite DNA
has been implicated in key traits including aneuploidy, infertility, and cancer since 1990s using small cohorts.
Despite their importance, centromeres remain among the least characterized genomic regions. Conventional
reference genomes (GRCh38) leave centromeres structurally unresolved with abundant near-identical yet highly
complex repeat sequences. Due to the lack of representation in the reference genome, modern genome-wide
association studies (GWAS) have entirely excluded centromeric variations. Consequently, we lack the landscape
of centromeric variation to human health and diseases.
To address this gap, we will significantly scale-up the catalog of human centromeric variations to millions of
individuals by developing statistical inference algorithms that enable imputation of centromeric variations and
robust association studies with human traits and diseases. We will construct a haplotype-resolved reference
panel leveraging recently developed telomere-to-telomere long-read centromere assemblies and use it as a
basis for computational inference of centromeric variation from whole-genome sequence data. We will apply
these methods to existing large-scale biobanks with millions of genomes and deep phenotype data and perform
phenome-wide association studies to comprehensively characterize clinical sequalae of centromeric variations.
In addition, we will develop computational algorithms to detect somatic copy number alterations involving
centromeres in tumor tissues and profile their non-coding RNA expression. We will thereby investigate the links
between centromeric somatic mosaicism, genome instability, transcriptional desilencing, aging, and cancer.
This project can transform the conventional paradigm of centromere research by bringing centromere back into
the state-of-the-art population-scale genomics—recasting these regions from unsequenceable gaps into
disease-relevant loci with novel insights into clinical practice and biology.
Publications
None