
The reference sequence that most genomic analyses are compared against may be steering automated software toward false alarms, according to new research from Semmelweis University in Hungary.
Whole-genome sequencing could one day let doctors identify a person's risk for specific diseases before symptoms appear, and is a central key to “personalized medicine”. But accurately reading a genome depends on what it's compared with. The reference widely used and accepted internationally is GRCh38, a single, linear coordinate system. Although it contains genetic data from multiple individuals, 70% of it is based on the genome of one male from Buffalo, New York. It isn't an “average” human genome that applies equally well to everyone.
For the study, published in GeroScience, the researchers compared whole genomes from 20 healthy Hungarians with GRCh38. In all 20 cases, the automated analysis identified the same severe genetic variant. After expert curation using additional, representative regions of the genome, the variant turned out to be a false positive—none of the 20 people had a true, medically relevant variant.
The problem arises because we rely on automated software to complete whole genome analysis
“If the system encounters a healthy genetic variant in a patient that is listed in the reference as an exceptionally rare gene variant, it may trigger a false alarm,” said Gyula Richárd Nagy, corresponding author of the article.
In other words, if the reference doesn't capture the full range of human genetic diversity, then without expert oversight, healthy genetic traits could be mistaken for defects.
The researchers stress that this isn't a problem in current practice, since experts manually review the discrepancies the software identifies using other population databases. But, that correction step would become an unsustainable bottleneck if the number of screenings grew substantially, and could stand in the way of expanding population-wide preventive genomics.
The team points to a graph-based pan-genome as a global solution. Rather than comparing every person's genome to a single reference sequence, a pan-genome builds a network-like map from the data of many genetically diverse individuals, depicting common and divergent segments of multiple genomes in a branching structure. This approach would cover human diversity more comprehensively and eliminate this type of error “by design.” A pan-genome graph already exists as a computational reference, but it isn't yet used routinely in clinical practice. The main obstacle is computing power.
“The key question for the future is to what extent information technology will be able to keep pace with advances in medicine,” Győrffy said. “Current everyday IT systems simply cannot support the daily use of such massive databases.”