In the latest episode of The Genetics Podcast, Patrick is joined by Dr. Jeffrey Barrett, Professor of Computational Predictive Medicine at the University of Helsinki and the Ellis Institute Finland. Jeff is a statistical geneticist who founded Open Targets and served as CSO of Genomics plc and of Nightingale. He also returned to the Wellcome Sanger Institute to help lead the UK's COVID-19 genome sequencing effort.
His argument is direct. Common chronic diseases likely contain many molecular subtypes that current care does not distinguish. The main constraint on finding them is access to data: linking biobanks so models can learn from millions of people, and using large language models (LLMs) to read the clinical notes that record treatment response.
Jeff starts with a contrast. Most cancers now come with companion diagnostics, usually based on tumor sequence or methylation profile, that point patients toward targeted drugs. Common chronic disease has almost none of that, even in the GLP-1 era. The gap persists even though a recent GLP-1 trial proteomics paper showed a dramatic shift in the plasma proteome versus placebo.
His clearest example comes from inflammatory bowel disease (IBD). An Oxford paper published this year by Holm Uhlig and colleagues found that about 3–4% of IBD patients carry anti-IL-10 autoantibodies. Across IBD, the HLA association has an odds ratio of about 2. In this subgroup, it rises to roughly 20–50.
These patients respond poorly to traditional therapy and undergo more surgery. As Jeff puts it: "But it is a molecular subtype of IBD defined almost entirely by a genotype. It has been kind of hiding in plain sight for forty years." For trial sponsors, the consequence is concrete. A subgroup with a distinct genetic profile and a distinct treatment response can sit inside a broad IBD population for decades without being recognized.
Jeff expects most subtypes to be harder to find than the IL-10 group, which is why he puts data ahead of algorithms. In his words, "my hunch is that the much bigger part of the current problem is in being able to access all of the data that are already out there, and this is really very concrete when you look at biobank data sets."
The obstacles are practical. Trusted research environments (TREs) require a human to review every data export. Federated learning, where a model trains across separate datasets without pooling them, needs thousands of model exchanges. A proof of principle discussed in the episode shows the approach can work: slicing UK Biobank into five datasets of about 100,000 people each recapitulates the prediction achieved on the whole dataset.
The remaining barrier is governance. Jeff describes it as a human, ethical, and legal challenge that LLMs will not speed up. Solving it would let models learn from a few million people at once, instead of five separate groups of 500,000.
For most of his genomics career, Jeff has held that additive linear models get you almost all the way there. In UK Biobank today, complex multitask residual neural networks barely beat lasso regression. At the scale of a few million people, with a wider range of data modalities, he is starting to change his mind about the importance of nonlinear effects.
Part of the evidence is RAVEN, an EHR model trained on a New York dataset of about 1.3 million people. Its performance did not saturate as the data grew. Jeff's new lab is applying these ideas to three disease areas:
Finland's health registries add another layer. FinnGen covers about 500,000 people. A separate secondary-use registry covers up to 7 million people and includes clinician notes, with identifiers scrubbed by Findata.
Those notes record why a clinician switched a patient's treatment, which is direct evidence of how that patient responded. Jeff's colleague Andrea Ganna has built a proof of concept that uses LLMs to extract this information. Linked to genetic and molecular data, treatment response at registry scale is the kind of outcome data that can tie a molecular subtype to a clinical decision.
Jeff uses AI heavily in his own work, including Claude for reading papers, and he draws clear lines. He treats writing as thinking, so he writes anything he wants another person to spend time reading himself. He likens handing that work to AI to bringing a forklift into the gym. For code, he concentrates his review on the tight analysis code.
Training is his larger concern. Trainees should be allowed to use Claude, he argues, but they then need to be "in a room with no computer for an hour" to discuss with someone who knows the subject. He also expects the bottleneck in many fields to move from computation to experimental data generation, where lab expertise remains the differentiator.
The episode closes on Jeff's time at Sanger during the pandemic. The institute sequenced about 2.5 million SARS-CoV-2 genomes, the largest such effort globally, with a swab-to-sequence turnaround of 48–72 hours. That experience shapes his view of AI risk. Viral evolution is hard to predict, so he considers the idea that an LLM could simply invent a super virus naive for now.
Jeff's new lab is hiring PhD students and postdocs. Listen to the full episode below.