1 min readHealth & Medicine

Scientists rethink how we sequence human genomes

Professor Euan Ashley breaks down why the field’s 20-year-old reference genome falls short – and how a more complete “pangenome” could sharpen precision medicine.

Euan Ashley smiles at the camera in a lab, while another researcher in a white coat works in the background.
Euan Ashley | Jim Gensheimer

When scientists strung together the first human genome more than 20 years ago, it was a historic moment for science and humanity. Fast-forward to today’s clinic – sequencing a human genome is no longer a special occasion. The once-arduous undertaking is an increasingly ubiquitous resource in medicine, having helped researchers and physicians parse the genetic foundation of millions of people worldwide.

An individual’s genome holds near-innumerable clues to their present and future health. Mapping the human genome broadly has created opportunities across biomedicine to understand how genes (which are composed of long stretches of DNA) give rise to our biological selves – including aberrations that can lead to certain diseases or conditions.

A sequenced genome is one of the most dependable harbingers of health and a valuable resource for a variety of different applications in precision medicine: It can help detect fetal abnormalities early, reveal rare diseases, and advance cancer care – from personalized cancer vaccines designed around a patient’s unique tumor mutations to detecting initial and recurring tumors earlier and treating them more precisely.

Yet, even as the utility and prevalence of genome sequencing expand and evolve, the technologies and standards used to parse and interpret the genome have stagnated. In a new scientific paper, Euan Ashley, MB ChB, DPhil, and a team of leading multidisciplinary genomics experts discussed why new thinking around genome sequencing is warranted, and why sorting out these details is crucial to precision medicine.

Ashley, the Arthur L. Bloomfield Professor in Medicine and the Roger and Joelle Burnell Professor in Genomics and Precision Health, explained the most important aspects of that work.

Let’s start with genome sequencing. What is it and how does it work?

Genome sequencing allows us to read a person’s complete DNA sequence, which can tell us many things about their health. But, contrary to what people might think, when you sequence a person’s genome, what’s produced isn’t one long readout of every letter of DNA in a certain order.

Sequencing machines use bits and pieces of DNA, often collected from blood or saliva, to spell out chunks of the genome. Initially, what you have is a collection of snippets of DNA. Then an algorithm matches up those pieces with something called a reference genome to create the full picture of an individual’s genomic self.

The reference genome is a representation of the human genome that many scientists use as a guide when piecing together individual genomes. I often liken the process of sequencing a human genome to doing a complex jigsaw puzzle. Pieces of DNA that most closely resemble certain parts of the reference genome are matched to those specific sections, and bit by bit, that matching exercise eventually results in a person’s complete sequence. In a very simplified example, if a location on the reference genome has a string of DNA that’s AAGTTTC, and an individual’s DNA snippet is AAGTTTC, that’s likely where that piece of the sequence belongs.

When the genome is complete, researchers can compare the individual against the reference and look for mutations, or genetic errors, that might be causing a disease. The reference is a map and a comparison tool wrapped into one, but it comes with its own complexities and drawbacks – it’s something that I, along with many colleagues in the genome sequencing field, are actively trying to rethink.

Why is there a need to rethink the reference genome?

The reference genome was made by cobbling together a few anonymous genomes about 20 years ago. And while gene sequencing technology has moved on since then, our standards – the way we analyze and interpret genomes, which largely depends on the reference genome – have not.

Scientists still compare an individual’s DNA to this one reference, but it’s not a perfect representation of the human genome. First, it’s incomplete – the reference genome captures the coding regions of DNA, but it leaves out a small but significant chunk of the genome.

There are also known genetic variations that aren’t well represented in the reference, which makes them harder or near impossible to detect in a real population. Inevitably, that means some rare but clinically important variants (mutations that contribute to diseases) are missed.

Lastly, the reference is haploid, meaning it only represents half of the human genome. But humans are diploid, meaning we have two copies of every chromosome: one from mom and one from dad.

So, what’s needed is a new version of the reference genome that captures the full picture: both copies of our chromosomes and the genetic variation between and within genes – rather than comparing everyone to a single, outdated, incomplete reference.

What does an ideal solution look like?

Researchers have been working toward something called the pangenome, which is essentially an expansion of how we represent the human genome. It’s a collection of genomes – not just one reference – and each one is a real genome from a real person. If you think of the reference genome as a single map, the pangenome is an atlas. It’s a broad collection of genomic diversity that creates a fuller representation of humanity’s diverse genetic makeup.

In addition, the best pangenomes capture the entire genome. We call it “T2T” or “telomere-to-telomere” to signal that the sequencing information includes the entirety of every chromosome. (Telomeres are the molecular caps on either end of a chromosome.)

If you think of the reference genome as a single map, the pangenome is an atlas. It’s a broad collection of genomic diversity that creates a fuller representation of humanity’s diverse genetic makeup.

We’ve talked about what’s needed to more effectively interpret and build a human genome. But what about altering the genome? Does gene sequencing play a role in gene therapy?

Absolutely. Gene sequencing is often the first step to determine whether someone is a candidate for gene therapy. Sequencing is also a sort of quality control mechanism for gene editing tools, such as CRISPR. Before these tools are used in patients, researchers need to confirm they’re making the right DNA edit. So, we test them first on a group of cells in a dish and then use sequencing to see whether the editor made the right changes, or if it also edited other parts of the genome erroneously – so-called off-target effects.

Currently, when scientists test these editing technologies, we sort of go looking under the lamppost. There are algorithms that predict where the most likely off-target edits might occur, but as sequencing technology has improved, we should be looking across the entire genome. Doing so could uncover unexpected changes that older approaches might miss, while also helping improve the algorithms that predict where off-target edits are most likely to pop up.

How does genome sequencing fit into precision medicine?

Over the past several years, we’ve seen the core technologies for sequencing become so affordable that its prevalence has shot up, both at the level of individual sequencing and in population-level genetic discovery programs like All of Us and the U.K. Biobank. To date, millions of people worldwide have had their genomes sequenced.

Whether it’s to diagnose patients with rare diseases, personalize cancer treatments, or screen babies for conditions, each of those applications uses the same premise – sequence a person’s genome and analyze for mutations, but each field currently follows a bit of a different path. That’s why it’s so important to bring together experts from different biomedical fields and figure out how we can harmonize our approaches. That way we’re not recreating the wheel every time.

Ultimately, the goal is to create a common foundation for genomic medicine to help ensure every patient benefits from the most accurate, consistent, and reliable interpretation of their genome.

For more information

This story was originally published by Stanford Medicine.

Writer

Hanae Armitage

Share this story