Publication Date
12-1-2020
Document Type
Article
Publication Title
Scientific Data
Volume
7
Issue
1
DOI
10.1038/s41597-020-00664-2
Abstract
Here we report whole genome sequencing of four individuals (H3, H4, H5, and H6) from a family of Pakistani descent. Whole genome sequencing yielded 1084.92, 894.73, 1068.62, and 1005.77 million mapped reads corresponding to 162.73, 134.21, 160.29, and 150.86 Gb sequence data and 52.49x, 43.29x, 51.70x, and 48.66x average coverage for H3, H4, H5, and H6, respectively. We identified 3,529,659, 3,478,495, 3,407,895, and 3,426,862 variants in the genomes of H3, H4, H5, and H6, respectively, including 1,668,024 variants common in the four genomes. Further, we identified 42,422, 39,824, 28,599, and 35,206 novel variants in the genomes of H3, H4, H5, and H6, respectively. A major fraction of the variants identified in the four genomes reside within the intergenic regions of the genome. Single nucleotide polymorphism (SNP) genotype based comparative analysis with ethnic populations of 1000 Genomes database linked the ancestry of all four genomes with the South Asian populations, which was further supported by mitochondria based haplogroup analysis. In conclusion, we report whole genome sequencing of four individuals of Pakistani descent.
Funding Number
R01EY022714
Funding Sponsor
National Eye Institute
Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.
Department
Computer Science
Recommended Citation
Shahid Y. Khan, Muhammad Ali, Mei Chong W. Lee, Zhiwei Ma, Pooja Biswas, Asma A. Khan, Muhammad Asif Naeem, Saima Riazuddin, Sheikh Riazuddin, Radha Ayyagari, J. Fielding Hejtmancik, and S. Amer Riazuddin. "Whole genome sequencing data of multiple individuals of Pakistani descent" Scientific Data (2020). https://doi.org/10.1038/s41597-020-00664-2
Comments
This is the Version of Record and can also be read online here.