In August 2025, the UK Biobank released the largest single-study whole-genome sequencing dataset assembled to date. It included 9,674 participants of South Asian ancestry among 490,640 total. That is twice the size of the next-largest South Asian cohort available anywhere: gnomAD v3's 2,419 South Asian genomes. Read as a headline, that sounds like real progress: the world's most-used genomic reference resource just doubled its South Asian sample overnight.

It is worth slowing down on that. Doubling a number only matters relative to what it started from, and what it started from was close to nothing. South Asians make up about 25% of the world's population but less than 2% of participants in genome-wide association studies (GWAS), the large gene-hunting studies that most disease-risk tools are built from. The UK Biobank's new release does not move that ratio in any meaningful way: 9,674 of 490,640 participants is about 2% of the release, not 25%.

A quarter of humanity, under one percent of the studies.

World Bank data put South Asia's population at about 1.69 billion in 2025, roughly one in four people alive today. The genomic infrastructure that increasingly decides who gets an accurate disease-risk score and who gets the right drug dose was built almost entirely on someone else's genome.

Bar chart comparing South Asians' 25 percent share of the world's population against their 0.8 percent share of genome-wide association study participants.

Source: World Bank; Nature Medicine, 2022, Fatumo et al.. Chart: The Signal.

Grading medicine on a biased curve

The imbalance runs through the whole field, well beyond the UK Biobank. A 2022 Nature Medicine roadmap paper on genomic diversity reports that genome-wide association studies have been conducted in individuals of European descent 86.3% of the time, versus 0.8% in South Asian populations. A polygenic risk score sums the small effects of thousands of genetic variants, each one calibrated on whichever population the underlying GWAS was run in. When almost all of that calibration data comes from European-ancestry cohorts, the resulting score is only as portable as it is accurate elsewhere, and it does not port well. A 2023 Brain Communications study reports that a European-derived polygenic risk score explains 4.8% of disease-risk variance in European-ancestry UK Biobank participants but only 1.1% in a South Asian cohort: more than four times the predictive power on one ancestry group than the other, from the identical score.

Bar chart showing a European-derived polygenic risk score explains 4.8 percent of disease-risk variance in a European ancestry cohort versus 1.1 percent in a South Asian ancestry cohort.

Source: Brain Communications, 2023. Chart: The Signal.

A pill that works differently

The gap is not confined to risk forecasting; it reaches the prescription pad. A 2023 study in JACC: Advances, drawing on the Genes & Health cohort of 44,396 British South Asians, reports that 57% are CYP2C19 intermediate or poor metabolizers and 13% are poor metabolizers, versus 2.4% poor metabolizers among Europeans. CYP2C19 is the enzyme that activates clopidogrel, a blood thinner routinely prescribed after a heart attack to prevent a second one. The same study found that poor metabolizers taking clopidogrel had 3.1 times higher odds of a recurrent heart attack than patients whose bodies process the drug normally. This is not a hypothetical calibration problem. It is a drug that already, measurably, protects South Asian patients less well on average, prescribed at standard dose regardless of ancestry in most routine care.

The fix for that exists as a clinical service, just not here. A 2025 scoping review in The Pharmacogenomics Journal identified 37 implemented CYP2C19 genotype-guided clopidogrel testing services worldwide, 76% of them in the United States, spread across a handful of other high-income health systems. None were in India or anywhere else in South Asia, the region the Genes & Health data says needs it most.

Bar chart showing 13 percent of British South Asians are poor CYP2C19 metabolizers versus 2.4 percent of Europeans, affecting how well clopidogrel works after a heart attack.

Source: JACC: Advances, 2023. Chart: The Signal.

The reference problem underneath the studies

Who gets studied is only part of the shortfall. Another part is what the reference databases themselves contain. A 2020 Nucleic Acids Research paper on India's IndiGenomes database reports that of 55.9 million genetic variants catalogued from 1,029 whole-sequenced Indian genomes, 32.23% were entirely absent from the gnomAD and 1000 Genomes reference databases. Those are the same reference databases that most clinical genome interpretation software checks a patient's variants against. A variant that is common in an Indian genome but missing from the reference panel cannot be correctly flagged as ordinary background variation; it risks being misread as rare, or simply not evaluated at all. Nearly a third of the variant catalogue from a single Indian genome study had nowhere to be checked against.

India's own answer to that gap is still catching up to the scale of the problem. A 2025 Nature Genetics paper on the GenomeIndia project reports that its national whole-genome sequencing effort has sequenced and analyzed genomes from 9,772 individuals, the country's largest genomic dataset to date. That is a real base, built specifically on Indian diversity rather than borrowed from elsewhere, but it is still smaller than the 9,674 South Asian genomes inside the UK Biobank's 2025 release alone, a foreign biobank's ancestry subgroup outsizing an entire national genome project.

Even the newest release is still small in absolute count.

Reference panelSouth Asian genomes included
UK Biobank whole-genome sequencing (2025)9,674
Trans-Omics for Precision Medicine (TOPMed)4,599
gnomAD v32,419
1000 Genomes Project601
Human Genome Diversity Project181

Source: Nature, 2025, UK Biobank whole-genome sequencing consortium. Table: The Signal.

What fixing it looks like

None of this is intractable, and there is already a working demonstration of the fix. A 2023 study in Therapeutic Advances in Endocrinology and Metabolism reports that a type 2 diabetes polygenic risk score built from Asian Indian GWAS data outperformed the European-derived score by about 12%, an odds ratio of 1.38 versus 1.23, when both were validated in the same South Asian UK Biobank cohort. The improvement came from something simpler than a conceptual breakthrough: training data drawn from the population the score was meant to predict for.

The honest objection

The strongest case against calling this a standing failure is that the trend line is real. The UK Biobank's 2025 whole-genome release already put its South Asian cohort at twice the size of the next-largest one available anywhere, and that release is little more than a year old. Momentum built from a near-zero base can compound, and a consortium that has just proven it can recruit and sequence tens of thousands of South Asian genomes in one push has a plausible path to doing it again, at larger scale, sooner than the historical pace of the field would suggest.

That case is real, but it describes a rate of change, not a current state. GWAS participation overall is still under 1% South Asian, and even doubling 2,419 to 9,674 only takes the UK Biobank's own release to about 2% South Asian, against a 25% population share. A single release, however large by absolute count, does not close a proportional gap of that size, and every polygenic risk score and prescribing algorithm trained on the data collected before 2025 was calibrated on the smaller base.

The Signal

The databases are getting bigger. The proportion is barely moving. Every hospital, insurer, and drugmaker that adopts a genomic risk tool trained mostly on European-ancestry data is asking a quarter of humanity to trust a score calibrated on someone else's genome, and the clopidogrel odds ratio and the polygenic-risk-score variance gap show that is not a theoretical concern: it already changes which patients a standard prescription protects and how reliable their disease forecast is. The fix that already works, ancestry-matched training data for the type 2 diabetes score, is not a mystery to the field. It is a resourcing choice. Watch whether the next major genomic release reports its South Asian sample as a share of the cohort, not just as an absolute headcount. That is the number that will say whether the gap is actually closing, or just growing more slowly than everything else around it.

Reporting basis: the South Asian share of world population and of GWAS participation is per a 2023 review in the journal Behavior Genetics. South Asia's 2025 population total is World Bank data. The UK Biobank's 2025 whole-genome sequencing cohort, and its comparison to gnomAD v3, TOPMed, the 1000 Genomes Project and the Human Genome Diversity Project, are from the Nature paper describing that release. The 86.3% versus 0.8% GWAS ancestry breakdown is from a 2022 Nature Medicine roadmap paper by Fatumo and colleagues. The polygenic-risk-score variance figures are from a 2023 Brain Communications study of the Genes & Health and UK Biobank cohorts. The CYP2C19 metabolizer rates and the clopidogrel odds ratio are from a 2023 JACC: Advances study of the Genes & Health cohort of 44,396 British South Asians. The IndiGenomes variant figures are from a 2020 Nucleic Acids Research paper by the CSIR-Institute of Genomics and Integrative Biology. The type 2 diabetes polygenic-risk-score comparison is from a 2023 study in Therapeutic Advances in Endocrinology and Metabolism. The GenomeIndia project's genome count is from a 2025 Nature Genetics paper by the GenomeIndia Consortium. The count and geographic spread of implemented CYP2C19-guided clopidogrel testing services is from a 2025 scoping review in The Pharmacogenomics Journal. The percentage-point comparisons and ratios stated in the text are The Signal's calculations from those figures.