GRCh37 vs GRCh38 vs T2T-CHM13: Choosing a Human Reference Genome

Language: اردو ورژن پڑھیں
Scope: Educational and research guidance. Always follow the reference build required by your validated pipeline, laboratory, database and reporting standard.
Last reviewed: 5 October 2026.

GRCh37 vs GRCh38 vs T2T-CHM13: Choosing a Human Reference Genome

A genomic coordinate is incomplete without its reference assembly. The same chromosome and position can point to a different base—or no equivalent position at all—when the assembly changes. That affects alignment, variant calling, annotation, primer design, database lookup and clinical reporting.

This guide compares GRCh37, GRCh38, T2T-CHM13 and the human pangenome, then gives a practical decision workflow. It complements the Sci Chores WSL2 NGS guide and the FASTQ and Phred quality guide.

What is a reference genome?

A reference genome is a coordinate system and sequence representation used to organize genomic data. It is not an “average person” and it does not contain every human allele. Linear references such as GRCh37 and GRCh38 provide one primary path for each chromosome plus additional sequences, patches and alternate loci.

The assembly name, patch level and sequence dictionary belong in the analysis record. “Human genome” alone is not reproducible.

GRCh37: legacy compatibility

GRCh37 was released in 2009 and remains embedded in older cohorts, databases, capture designs and validated pipelines. You may also see the related UCSC name hg19. They are closely related, but file naming, mitochondrial sequence and auxiliary contigs can differ; do not treat every GRCh37 and hg19 FASTA as byte-identical.

Use GRCh37 when a study must remain compatible with a legacy dataset or when a validated workflow explicitly requires it. Starting a new project on GRCh37 merely because older publications used it usually creates future conversion work.

GRCh38: the practical linear-reference default

GRCh38 was released in 2013 and corrected or improved many regions represented in GRCh37. The Genome Reference Consortium currently lists GRCh38.p14 as the latest patch release and has postponed a coordinate-changing GRCh39 while newer reference models are evaluated.

For many new short-read human analyses, GRCh38 is the most practical default because aligners, variant callers, annotation resources and public databases broadly support it. The exact FASTA still matters: primary assembly only, analysis set, decoy sequences, alternate contigs and HLA sequences can change alignment behaviour and contig naming.

What do patch releases mean?

A patch release adds correction or alternate sequences without changing the chromosome coordinates of the major assembly. For example, GRCh38.p14 does not mean that ordinary chromosome positions were renumbered fourteen times. Patch scaffolds must still be handled consistently by the aligner, reference index and annotation files.

T2T-CHM13: a much more complete linear assembly

The Telomere-to-Telomere Consortium published T2T-CHM13 in 2022. The assembly added nearly 200 million bases compared with prior references and resolved many centromeric, segmentally duplicated and repeat-rich regions. The original complete assembly covered the autosomes and chromosome X; a separate complete Y-chromosome assembly followed later.

T2T-CHM13 is valuable when the research question depends on regions poorly represented in GRCh38, long-read assembly, structural variation or repeat biology. It is not automatically the best replacement for every routine pipeline: tool support, gene annotation, population databases, truth sets, capture targets and historical coordinates may still be centred on GRCh38.

The human pangenome: multiple haplotypes instead of one path

The Human Pangenome Reference Consortium published a draft pangenome in 2023 using diverse, haplotype-resolved assemblies. A graph representation can retain alternative sequence paths that a single linear reference cannot show, reducing reference bias and improving representation of structural diversity.

Pangenome workflows are advancing quickly, but they are not a drop-in substitution for every BAM/VCF-based pipeline. Graph-aware mapping, variant representation and benchmarking require compatible software and carefully defined outputs. For production work, confirm that downstream annotation, quality control and reporting understand the chosen representation.

Why reference mismatches cause errors

  • A VCF coordinate may be interpreted against the wrong base.
  • Reads may align differently because decoys, alternate loci or contig sets differ.
  • Annotation may assign the wrong transcript consequence.
  • A primer may be designed against sequence that does not match the reported coordinate system.
  • Variant databases may return no record or the wrong record.
  • Lift-over may fail in rearranged, repetitive or assembly-specific regions.

GRCh37 to GRCh38 conversion: lift-over is not reanalysis

Coordinate lift-over maps intervals between assemblies. It does not realign reads, reproduce the original caller or guarantee an equivalent allele representation. Indels and complex variants often need normalization and allele-aware validation after conversion.

If raw reads are available and the result matters, re-aligning and re-calling against the target reference is generally more defensible than changing VCF coordinates alone. That is a practical recommendation, not a universal rule; validated pipelines may require a specific controlled procedure.

A practical decision table

Use case Reasonable starting choice Main caution
New short-read human project GRCh38 Use one documented FASTA, contig set and matching resources.
Legacy cohort or validated assay Required GRCh37 build Do not silently mix hg19/GRCh37 resource bundles.
Repeat-rich or previously unresolved regions T2T-CHM13 Check annotation, benchmark and downstream compatibility.
Graph-aware population or structural-variation research Human pangenome Use compatible graph tools and define variant representation.
Cross-study comparison One harmonized build Record conversion failures and validate alleles after lift-over.

Reference files that must match

The FASTA is only the start. Keep the following resources on the same assembly and compatible contig naming:

  • aligner indexes and sequence dictionaries;
  • known-variant resources used for recalibration;
  • gene and transcript annotation files;
  • interval lists and capture targets;
  • population-frequency and clinical databases;
  • blacklists, decoys and benchmark truth sets.

A mixture such as chr1 in one file and 1 in another is a warning. Renaming contigs may solve syntax but not biological incompatibility.

Primer design and reference builds

For PCR or Sanger work, record the assembly, transcript accession/version and HGVS description before retrieving flanking sequence. Pronto Primer can help generate candidates around a known variant, but the input build must match the variant description and every primer pair still requires specificity checks and laboratory validation.

Minimum reproducibility record

  • assembly name and patch level;
  • FASTA source, URL, release date and checksum;
  • contig naming and included alternate/decoy sequences;
  • annotation release and transcript set;
  • reference-dependent databases and versions;
  • alignment and variant-calling tool versions;
  • any lift-over chain, failed intervals and allele-normalization steps.

Bottom line

Choose the reference that fits the question and the complete toolchain. GRCh38 is usually the practical linear-reference starting point for a new conventional human NGS workflow. GRCh37 remains necessary for some legacy and validated datasets. T2T-CHM13 adds difficult genomic regions, while the pangenome represents multiple haplotypes and reduces reliance on a single linear path. None of them removes the need for compatible resources, benchmarking and explicit documentation.

Authoritative sources

  1. Genome Reference Consortium: Human genome overview
  2. NCBI human genome resources
  3. Nurk et al. The complete sequence of a human genome. Science, 2022
  4. Liao et al. A draft human pangenome reference. Nature, 2023

Don’t miss out on Science!

We don’t spam! Read our privacy policy for more info.

Leave a Comment

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Scroll to Top