Dr Manuel Corpas


Dr Manuel Corpas is a globally recognized genomicist and health data scientist whose work has advanced the frontiers of equity in precision medicine. His research spans population genomics, pharmacogenomics, and biobanking, with a longstanding commitment to underserved and underrepresented populations. As President of the Spanish Congress of Genomic Medicine, he leads the largest Spanish-speaking platform for genomic health equity.

He is a driving force behind major sequencing efforts such as the Peruvian Genome Project, which expands global reference datasets to include diverse Indigenous and Latin American populations. He has contributed to widely adopted clinical and open-source tools, including DECIPHER for rare disease diagnosis and BioJS for genomic data visualization.

Dr Corpas’s leadership spans cross-sectoral impact: from co-developing GA4GH’s global standards for equity and inclusion, to contributing to landmark initiatives such as Deciphering Developmental Disorders (13,000 trios), ELIXIR-UK, and GOBLET. As a Fellow of both the Alan Turing Institute and the Software Sustainability Institute, he is advancing the use of AI to convert health data into actionable, inclusive insight.

A prolific author with over 80 peer-reviewed publications and a sought-after international speaker, Dr Corpas blends rigorous science with visionary public engagement. His mission: to harness the full potential of AI and genomic data to build equitable, globally scalable health solutions, with a special focus on Latin America and the Global South.


Dr Corpas’ research sits at the intersection of genomics, AI, and health equity, building measurable systems that close representation gaps, especially across Latin America. He leads population‑scale projects using long‑read sequencing (ONT) and emerging multi‑omics to characterise novel human variation and translate findings into clinically actionable and pharmacogenomic insights, with robust community engagement and governance.

A second strand develops tools for equity assurance in data‑driven medicine. He is creating the Health Equity Informative Marker (HEIM), a quantitative framework to audit and benchmark representation, governance, and benefit sharing across biobanks and AI pipelines. His group applies LLMs and advanced analytics to EHR‑linked biobanks (e.g., UK Biobank), maps field trends via large‑scale MeSH/semantic clustering, and evaluates platform bias (arrays vs WGS) and global population diversity (1000G, SGDP).

Outputs span open datasets, reproducible methods, and policy‑relevant evidence aligned with GA4GH and FAIR principles, aimed at standards that ensure precision medicine benefits all populations.



Sustainable Development Goals
In brief

Research areas

Human Genomics, Bioinformatics, Next Generation Sequencing, Computational Genomics and Personal Genome Interpretation

Skills / expertise

Programming, Data Analysis, Pipeline Development, UNIX and Cloud Computing

Supervision interests

Polygenic Risk Scores, Statistical Genetics, Whole Genome/Exome Analysis, Population Genomics, Diversity Genomics, Artificial Intelligence and Data Science