In a recent paper published in Nature, researchers analyzed 150,119 genome sequences from the United Kingdom Biobank (UKB).
Study: The sequences of 150,119 genomes in the UK Biobank. Image credit: Yurchanka Siarhei/Shutterstock
background
A thorough and accurate characterization of sequences and phenotypic variation is necessary for a detailed understanding of how variations in the human genome sequence influence phenotypic diversity. Over the last ten years, insights into this association have been uncovered by whole genome sequencing (WGS) or whole exome sequencing (WES) of large cohorts with rich phenotypic information.
With a healthy participant bias, the UKB records the phenotypic diversity of 500,000 people across the UK. The UKB WGS Consortium sequences the entire genome of each participant to an average depth of at least 23.5 base pairs.
About the study
In the present study, researchers report WGS analysis of 150,119 UKB participants. From the UKB pool of volunteers, pseudo-random individuals were chosen and divided into the two sequencing sites. The authors stated that using a depletion score (DR) of genome-spanning windows, this extensive database of variants allows the assessment of selection based on sequence diversity within a population.
Overall, the study report on the initial data release contains a large collection of WGS-focused sequence variants from 150,119 individuals, including short insertions or deletions (indels), single-nucleotide polymorphisms ( SNPs), structural variants (SVs) and microsatellites. .
Each variant call was conducted jointly across all participants to provide accurate data comparison. The resulting dataset provided a rare opportunity to investigate human sequence diversity and how it affects phenotypic variation.
In addition, the team describes some of the discoveries made possible by this huge new WGS data resource that would be difficult or impossible to make with WES and SNP array datasets.
results
The researchers noted that the dataset generated by sequencing the whole genomes of more than 150,000 UKB participants was unparalleled in scale, providing the most comprehensive analysis of sequence heterogeneity in the genomes to date. the germline of a population.
The team provided two pairs of variant classes often not examined in genome-wide association studies (GWAS), namely 1) indel and SNP data and 2) SV and microsatellite data, identifying a large number of sequence variants among WGS participants. This group comprises a series of high-quality variants consisting of 58,707,036 indels and 585,040,410 SNPs, representing 7% of all potential human SNPs.
DR examination reveals that coding exons constitute only a small fraction of the areas of the genome prone to significant sequence conservation. The authors identified three cohorts under the UKB: a smaller African cohort, a South Asian cohort and a sizeable British Irish cohort.
The study provides a haplotype reference panel, which facilitates accurate imputation of the majority of variants harbored by three or more sequenced subjects. The team discovered two types of variants that have typically been left out of extensive WGS analyses, namely 2,536,688 microsatellites and 895,055 SVs.
Compared to the WES of the same individuals, the number of indels and SNPs was 40 times greater. Even within the identified coding exons, WES missed 10.7% of the variants discovered by WGS. The majority of the remaining genome was not covered by WES, including untranslated regions (UTRs), functionally significant promoter regions, and unannotated exons. The identification of rare non-coding sequence variants with drastic impacts on menarche and height versus any variant revealed in GWAS to date serves as an illustration of the importance of these regions.
Conclusions
The current research provides numerous examples of trait relationships for uncommon variants with deep impacts using this powerful new WGS resource not previously discovered through WES- or imputation-based research.
Collectively, the scientists anticipate that the DR score discussed in the paper will be a valuable tool for recognizing genomic areas of functional importance. However, further studies are warranted to fully understand its characteristics, implications, and how it compares with other metrics of sequence conservation and constraint.
Although the coding exons were subjected to strong purifying selection, as shown by a low DR value, they constitute only an insignificant part of the areas with a low DR value. The authors mentioned that the description of the extensive sequencing of the present research and the ongoing efforts to sequence the entire UKB are expected to significantly advance the knowledge of the role and relevance of the non-coding genome.
The current findings should greatly improve the understanding of the connection between phenotypic variety and human genome variability when combined with in-depth analysis of phenotypic variation across the UKB.