Send Message

Pan-Genomes: Why One Reference Genome Is No Longer Enough for Crop Improvement

For two decades crop genomics has run on a single reference genome per species, one plant standing in for millions. Long-read sequencing has shown that this convenient assumption discards a large share of a crop's genetic variation. A pan-genome, the combined gene content of many individuals of a species, reveals that only a fraction of genes occurs in every plant while the rest come and go between varieties: in potato, just a quarter of 51,401 pangene clusters are present in all 45 accessions examined. These presence-absence and structural variants underlie disease resistance, stress tolerance, yield and quality traits, yet they are largely invisible to a single linear reference. Graph pan-genomes, which hold every version of a region in one searchable structure, recover much of this hidden variation; in tomato they raised the average estimated heritability of 20,323 traits from 0.33 to 0.41. This article explains what a pan-genome is, how a graph reference differs from a linear one, and what it offers plant breeders working on crops from rice to chickpea.