My plenary keynote at the 2026 World Congress of Psychiatric Genetics, Glasgow, 1 October 2026. Full recorded run-through, slides and podcast below.

Let me start with a file.

A genome file, handed to the pharmacogenomics tool of an AI agent. The tool read it, ran its analysis and returned a clean, confident report: all normal, recommended doses for 51 drugs.

The file was empty. There was no DNA in it. Not a single genotype.

We found this while testing agent tools for our Perspective on agentic genomics in Cell Genomics, and it has since been fixed. I keep coming back to it because it is the purest form of the problem I took to Glasgow. Not a wrong answer about a person. A confident answer about a person who was never in the data.

The record is uneven, and the unevenness compounds

Everything an AI system knows, it learned from our cohorts, our trials and our papers. That record represents some people far better than others. My group calls the pattern compounding neglect: a disease or population neglected at the first stage of research meets a barrier at every stage after it. No cohort, no discovery. No discovery, no trial. No trial, no treatment.

The evidence in the talk, in brief:

  • Biobanks. Across 70 major biobanks in 29 countries (38,595 papers since 2000), Africa carries 24% of the world’s disease burden and its six biobanks produced 0.7% of the papers. Europe and North America hold 45 of the 70 cohorts and produced 88%. Lymphatic filariasis affects around 120 million people; we found not a single paper on it across all 70.
  • Psychiatric evidence. Across five large health-record biobanks, UK Biobank accounts for 83% of the mapped depression papers, 79% for anxiety disorders and 87% for bipolar disorder. About 95% of its participants describe themselves as white.
  • Dosing rules. In PharmGKB, half of study participants had no classifiable ancestry. Of the rest, 64% were European, under 4% African and one in a thousand Indigenous American.
  • Interpretation. In the Peruvian Genome Project (736 people, 28 populations), 94% of 1,210 high-impact rare variants have no entry in ClinVar. We could read those genomes. We could not say what they meant.
  • The shape of knowledge. Mapping 13.1 million PubMed search results across 175 diseases, conduct disorder has the single most isolated literature, and ADHD, anxiety disorders and self-harm sit among the most isolated fifth. Disconnected knowledge is exactly what a machine finds hardest to reach.
  • The models. Across 10,500 questions to six language models, the better studied a disease, the more closely the answers resemble its literature (correlation about 0.66). Among diseases with the same amount of literature, those that weigh most on Africa are served less well.

Agentic genomics lowers the barrier to producing analyses, not to judging them

AI models are no longer only answering questions. They are doing the work. In our Cell Genomics Perspective with Heinner Guio and Segun Fatumo, we called this agentic genomics: agents that plan, run and refine entire genomic analyses. Our central argument fits in one sentence. Agentic genomics lowers the barrier to generating analyses. It does not lower the barrier to evaluating them.

So we stopped asking the agent to improvise the science. We write the scientific decisions down as a skill, an executable methods section anyone can open, read, challenge and run. Then we built a test that could prove us wrong.

What the pharmacogenomics benchmark showed

Eight models, 110 guideline cases, every case three times, more than 13,000 answers. Answering from memory, 63% matched our reference answers. Handed the guideline to read, agreement fell to 55%. With expert-written rules, it rose to 97%, and the spread between models narrowed from 21 points to 6.

Then the uncomfortable part. In one RYR1 case our answer key was wrong. Every model that read the guideline gave the correct answer, uncertain, and our benchmark marked all 24 wrong. Every configuration that followed our rules reproduced our mistake, and the benchmark marked all 24 right. Determinism is not truth. A machine that follows a wrong rule perfectly is precisely wrong, every time, and invisible because it agrees with the answer key. When we deliberately corrupted the rules, the agents followed the corruption in 87 of 90 responses.

On real genomes, agreement with the reference caller fell from 97% on curated cases to 64% in my own family, 61% in Iberian samples, 55% in Peru and 38% in Uganda, where the models stayed silent on four in ten genetic states. Two things changed at once here: real genomes carry wider, less textbook states, and the real-genome runs gave the models no rule table, while the curated cases did. Our expert rules cover 16% to 22% of the distinct states we found in the first three cohorts, and 6.6% in Uganda. The safest part of the system reaches least far exactly where the knowledge was already thinnest. That is compounding neglect, rebuilt in software.

Three things to do next

  1. Measure the silence. Count who receives no answer at all, population by population, the way a financial audit reports every account.
  2. Audit the answer key. Before trusting a 97%, ask whose truth it was measured against.
  3. Write the missing rules, by and for the populations the current rules miss. Six point six per cent coverage in Uganda and 94% of high-impact Peruvian variants with no clinical record are not just findings. They are the to-do list.

When the system does not know, it must say so. And then the gap should become our next question, never a verdict on the person in front of us.

Watch, listen, download

Corrections to the recording. At 15:07 I say “13 million abstracts and papers”; it is 13.1 million PubMed search results, which count an abstract once per matching disease. At 25:35, “97% accuracy” means agreement with our reference answers, and the models did not become identical: their spread narrowed from 21 points to 6. At 28:38 the Peruvian cohort is 736 people. At 30:11 the framing test used European, Latin American and East African descriptions.

Papers. Corpas, Guio and Fatumo, Agentic genomics: from pipeline automation to autonomous validation, Cell Genomics 2026 (doi:10.1016/j.xgen.2026.101305). Corpas, Iacoangeli, Bourdenx et al., pharmacogenomics benchmark, Cell Genomics (accepted); code and data at doi:10.5281/zenodo.22806813. Everything shown is open at clawbio.ai, where ClawBio now has 1,149 GitHub stars and 98 skills.

Podcast also available on PocketCasts, SoundCloud, Spotify, Google Podcasts, Apple Podcasts, and RSS.

Leave a comment

Personal Genomics Zone podcast

About the podcast

Personal Genomics Zone is my podcast on agentic genomics, AI agents in the lab, and who genomic medicine actually serves. Talks, live demos and long conversations with the people building the field.

Spotify | Apple Podcasts | YouTube

Read Latest Blog Entries