On 20 July 2026, The New York Times reported that the genetic risk tools entering clinical use do not perform equally well across populations. Polygenic risk scores, which aggregate hundreds or thousands of small-effect variants into a single estimate of disease risk, were trained overwhelmingly on European-ancestry DNA. For some conditions, the article reports, predictions for people of color are "almost no better than flipping a coin."
The finding lands weeks after the NIH announced the world's largest integrated database of human genomes, the All of Us program. Many of the most promising clinical risk tools are being built on it. A tool that misreads risk by ancestry will not stay a research artifact. It will shape who gets a statin early, who is screened for cancer, and who is wrongly told they are low risk.
The skew is well documented. An analysis of polygenic score studies published between 2008 and 2017 found that 67% included only European-ancestry participants. The pattern persists in the cohorts that anchor the field today. UK Biobank is roughly 94% white. The Million Veteran Program is close to three-quarters white. The All of Us cohort, built explicitly to be more representative, remains about 50% European.
The imbalance is also getting harder to correct. Funding for All of Us has been cut 72% since 2023. A major funding stream tied to the 21st Century Cures Act is set to expire at the end of the fiscal year.
Present-day European populations descend from a small group that left Africa roughly 50,000 years ago. That makes them among the least genetically diverse populations a researcher could study. The clearest illustration is PCSK9. A variant found mostly in a small share of Black participants pointed to a mechanism that lowers coronary heart disease risk by roughly 88%. It surfaced only because the Dallas Heart Study deliberately built a cohort that was more than half Black. It later led to a class of cholesterol drugs now used across all populations.
The Times frames the problem as one of data and modeling, and better statistics will help. Multi-ancestry methods that draw weighted signal from many cohorts already narrow the gap. Work by Cavazos and Witte shows that including variants discovered in African-ancestry populations improves prediction across groups. There is a limit to what modeling can recover. As Ding and colleagues have shown, score accuracy declines continuously with genetic distance from the training data. You cannot debias a tool using data you were never able to collect.
That places the bottleneck upstream. The rate-limiting step is diverse, consented recruitment, and recruitment depends on trust. The communities most underrepresented in genomic databases often have the most reason for caution. That caution is grounded in a documented history that includes Tuskegee, discriminatory sickle-cell screening, and the case of Henrietta Lacks. Better algorithms do nothing to address that history. Representative data has to be collected before it can be modeled, and it will only be collected where trust is rebuilt.
Sano's approach to precision patient finding combines electronic health records, existing or new genetic data, and lifestyle data. The goal is to identify eligible, representative patients earlier. The same infrastructure that makes recruitment efficient can also make it more equitable, because it reaches patients that conventional site-based enrollment tends to miss. Reaching people is only part of the work. Sustained participation depends on reciprocity: genetic counselling, clear consent, and returning value to participants so engagement continues before, during, and between studies.
A concrete example is Sano's Innovate UK-funded program with Predictive Health Intelligence and Somerset NHS Foundation Trust. It uses risk-based identification to find people at risk of non-alcoholic fatty liver disease within existing health databases. It then educates those patients and offers early intervention or trial opportunities. It is a working demonstration of representative, risk-based patient finding operating inside a national health system, rather than a proposal on a slide.
The mechanism generalizes. Risk-based identification tells you who to reach; the engagement layer determines whether they participate and stay. Both are needed to build the diverse, longitudinal datasets that portable risk tools depend on.
The useful shift is to treat representative recruitment as core scientific infrastructure, not a downstream data-quality fix. For teams building or buying polygenic tools, a few questions follow:
A better algorithm answers none of these on its own. They are answered by the operational work of building trust, reaching underserved populations, and feeding that representative data back into the models. The equity of the next generation of genetic risk tools will be decided as much in recruitment as in code.
To discuss representative recruitment and precision patient finding, get in touch.