The sequences of 150,119 genomes in the UK Biobank

Datasets

KB data

UKB phenotype and genotype data were collected after informed consent obtained from all participants. The Northwest Research Ethics Committee reviewed and approved the scientific protocol and operational procedures (REC reference number: 06 / MRE08 / 65) of the UKB. Data were obtained for this study and investigations were conducted under application license numbers UKB 24898, 52293, 68574 and 69804. The sequence data were processed as described in Additional Notes 1-4, Supplementary figs. 5–8 and Supplementary Tables 16 and 17.

Phenotypes were downloaded from the UKB. A total of 8,180, 1,291, and 459 phenotypes were constructed for the XBI, XAF, and XSA cohorts, respectively. The examples presented here were selected as notable representative examples of association. The processing of the phenotypes presented here, with reference to the field identity in the UKB data showcase, is provided in Supplementary Table 15.

Icelandic data

The set of gout samples60, a total of 1,740 Icelandic individuals, was recruited through multiple sources. A subset of these individuals were regular users of gout medications corresponding to class M04 of the therapeutic anatomical chemical classification system (ATC-M04). People using ATC-M04 were identified using questionnaires at the time of entry into genetics projects at deCODE and provided by the Directorate of Health from entry in the Register of Prescribed Drugs (2005-2020) or the Register of RAI Assessments and Minimum Data Set (MDS). ) for residents and applicants for nursing homes (1993–2018). In addition, approximately half had received a clinical diagnosis of gout (International Classification of Diseases: ICD-9 code 274 or ICD-10 M10 code) between 1984 and 2019 at Landspitali, the National University Hospital of Iceland, or in two rheumatology clinics. , or this diagnosis was determined by examining the medical records of RAI and MDS.

Serum uric acid levels in blood samples from 95,086 Icelandic individuals were obtained from Landspitali, the National University Hospital of Iceland and the Laboratory of the Icelandic Medical Center (Laeknasetrid) in Mjodd (RAM) between 1990 and 2020. Serum uric acid levels were normalized to a standard normal distribution by quantile-quantile normalization and then adjusted for sex, year of birth, and age to measure. For people for whom more than one measure was available, we used the mean normalized value. Serum uric acid levels were determined from an enzymatic reaction in which uricase oxidizes urate to allantoin and hydrogen peroxide, which, with the help of peroxidase and a dye, forms a complex. of colors that can be measured in a photometer at a wavelength of 670 nm. .

All participants who donated blood signed the informed consent. Participants’ identities were encrypted using a third-party system approved and supervised by the Data Protection Authority of Iceland. The study was approved by the National Bioethics Committee of Iceland (approval No. VSN-15-023) following the evaluation by the Icelandic Data Protection Authority. All data processing complies with the instructions of the Data Protection Authority (PV_2017060950ÞS).

RNA sequence data analysis was approved by the Icelandic Data Protection Authority and the National Bioethics Committee of Iceland (No. VSNb2015030021).

Danish data

Data were provided from the Danish Blood Donor Study (DBDS) 61. The DBDS genetic study has been approved by the Danish National Committee for Ethics in Health Research (NVK-1700407) and by the Data Protection Office of the Danish Capital Region (P-2019-99).

Call SNP and indel with GraphTyper

Before running GraphTyper, we preprocessed all indexes (CRAI) from the compressed reference-oriented alignment map (CRAM) by extracting a single large file containing all the CRAI index entries with the sample identifier for to a 50 kb window (with 1 kb fill in each). side of the region) for all samples. Then, for each region, we created a chopped CRAI for each sample by processing the large file of the corresponding region, substantially reducing the number of CRAI index entries read.

In addition, we created a script cache of the reference FASTA file using the script ‘seq_cache_populate.pl’ distributed with samtools 1.9. In each region, we copied the sequence cache to the local disk and used it to read the CRAM files by configuring the “REF_CACHE” environment variable.

We ran GraphTyper (v2.7.1) using the ‘genotype’ suborder. The complete order we executed was in the format:

graphtyper genotype $ {UKBIO_REFERENCE} –sams = $ {SAMS} –sams_index = $ {CRAI_TMP} /crai_filelist.txt –avg_cov_by_readlen = $ {COVERAGES} –region = $ {REGION} –threads = $ {THREADS } –verbous

Where UKBIO_REFERENCE is the GRCh38_full_analysis_set_plus_decoy_hla sequence file FASTA, SAMS is a list of all input BAM / CRAM files, CRAI_TMP is a path to CRAI files cut to local disk, COVERAGES is the coverage divided by each read length input, REGION is the genotyping region and THREADS is the number of threads to use.

SNP and indel calls with GATK are shown in Additional Note 5. Detailed comparisons of the GraphTyper and GATK call sets are provided in Additional Notes 6 and 7, Additional Figs. 9–12 and supplementary tables 18–21.

Execution time

All work was run on 12 cores with 60 GB of reserved RAM. Approximately 1% of the work was run again with 24 cores with 120 GB of reserved RAM. A few jobs that require more cores and memory, with a single job ending with 48 cores and 1,000 GB of RAM. The total CPU time reserved for the cluster was 5.8 million CPU hours and the total effective computation time was 5.0 million CPU hours. The difference in these numbers is explained by the fact that not all kernels reserved for the program may not use them all at once.

SV calls with Manta and GraphTyper

We ran an SV genotyping pipeline similar to the one we had previously applied to 49,962 Icelandic individuals50. In summary, we ran Manta48 v1.6 to discover SV in the 150,119 individuals in the genotyping set. We also created a set of highly reliable common SVs (imputation information greater than 0.95, with a frequency greater than 0.1%) from our previous studies using both short Illumina50 readings and long reading data from Oxford Nanopore49. Finally, we inferred a set of SVs from six publicly available assembly datasets using dipcall62, as described above50. We used svimmer50 to combine these different SV datasets and named the resulting SVs using GraphTyper50 version 2.7.1. By incorporating long-read data and high-quality assemblies, we were able to call more true SVs than using only short reads, especially for regular SVs.

A total of 895,054 variants were called, of which 637,321 variants were noted as “Passes”. Variant counts are presented for variants annotated by GraphTyper as “Pass”, unless otherwise noted.

Most SVs are deletions (81.3%); however, we observed only slightly more deletions than average insertions and duplications per individual (Fig. 4a). This is because the source of many inserts is long reads and assembly data and therefore many rare inserts are missing. Deletions are usually easier to detect in short read data. Individuals belonging to the XAF cohort carry more SV than the other cohorts (Fig. 4b).

Imputation and classification in phases

UKB samples were genotyped with an SNP chip with a bespoke Affymetrix chip, UK BiLEVE Axiom, in the first 50,000 individuals63, and the Affymetrix matrix UKB Axiom64 in the remaining participants. We used the existing long-range phase of genotyped samples with SNP5 chip.

We excluded SNP and indel sequence variants in which at least 50% of the samples had no coverage (GQ score = 0), if Hardy-Weinberg P the value was less than 10-30 or if the excess heterozygote was less than 0.5 or more than 1.5.

We used the remaining sequence variants and long-range chip data to create a haplotype reference panel using internal tools1,65. We then imputed the haplotype reference panel variants to the chip genotyped samples using internal tools and methods described above1,65.

Imputation consists of estimating, for each haplotype, the sharing of haplotypes with haplotypes in the haplotype reference panel, giving the weights of the haplotypes for each haplotype. These weights together with the allele probabilities for each haplotype of the haplotype reference panel allow imputation with a Li and Stephens66 model similar to that used in IMPUTE2 (ref. 67). Estimation of haplotype weights was based on long-range chip haplotypes.

The sequence variant phase consists of iteratively imputing the phase in each sequenced sample based on the other sequenced samples and the estimated phase of the last iteration. The imputed genotypes, along with the original genotypes, are weighted together to estimate new allele probabilities for haplotypes. The imputation is done as described above.

We calculated an omission rscore 2 (L1oR2) as a correlation squared (r2) of the original genotype calls, with the genotypes imputed for each sample by excluding the original genotype from the sample from the imputation entry.

Batch effects from the sequencing center were discovered in both the raw genotype (Supplementary Table 21) and the imputed data (Supplementary Table 22).

Identification of functionally important regions

To identify functionally important regions, we began by estimating whether reliable base calls can be expected to be made at each site in the genome. Sequence coverage at each base pair in GRCh38 was calculated for each of the 1,000 randomly selected individuals. At each base pair, we then calculated the mean and SD of coverage among the 1,000 individuals. Base pairs with an average coverage of at least 20 and an SD coverage of up to 12 were considered reliable base pairs. Only variants in GraphTyperHQ …

Leave a Comment

Your email address will not be published. Required fields are marked *