Language: اردو ورژن پڑھیں
Scope: Educational and research use. Quality scores are one part of sequencing QC; they do not by themselves validate a clinical result.
Last reviewed: 2 October 2026.
FASTQ and Phred Quality Scores: A Practical NGS QC Guide
FASTQ files carry two things for every sequencing read: the nucleotide sequence and a quality character for each base. Those quality characters are an estimate of base-call reliability. Reading them correctly helps you decide whether a dataset is ready for alignment, trimming, or a deeper investigation.
This guide follows the FASTQ specification and current Illumina documentation. It complements the Sci Chores WSL2 NGS workflow; it is not a substitute for a validated laboratory or clinical pipeline.
What is inside a FASTQ record?
A conventional FASTQ record has four lines:
- An identifier beginning with
@. - The nucleotide sequence.
- A
+separator line, optionally repeating the identifier. - A quality string with one character per nucleotide.
The sequence and quality string must have the same length. A mismatch usually indicates file corruption, an incorrect conversion, or a parsing problem that should be fixed before analysis.
How Phred quality scores work
A Phred score is defined as Q = -10 log10(P), where P is the estimated probability that the base call is wrong. Therefore:
- Q10 corresponds to about a 1 in 10 error probability, or 90% expected accuracy.
- Q20 corresponds to about a 1 in 100 error probability, or 99% expected accuracy.
- Q30 corresponds to about a 1 in 1,000 error probability, or 99.9% expected accuracy.
These are per-base error estimates, not a promise that every read or every variant is correct. Chemistry, cycle, library preparation, platform, and downstream evidence all matter.
Why the quality string looks like symbols
FASTQ stores quality values compactly as ASCII characters. In the commonly used Phred+33 encoding, the quality value is the character’s ASCII code minus 33. For example, ! represents Q0, while higher printable characters represent higher scores. Do not assume an old Phred+64 encoding unless the instrument or file documentation explicitly says so.
Some modern workflows use quality-score binning. Binning reduces the number of distinct quality values while preserving useful information for many analyses. Always record the platform and conversion settings in your run manifest.
What to check before alignment
- File integrity: run
gzip -t sample.fastq.gzfor compressed files and verify checksums when files were transferred. - Record structure: confirm that every record has four lines and that sequence and quality lengths agree.
- Pair consistency: for paired-end data, verify that R1 and R2 exist, have compatible names, and contain matching read counts.
- Per-base quality: inspect whether quality declines sharply toward the end of reads or differs between cycles.
- Adapters and overrepresented sequences: investigate adapter contamination and unexpected library sequences before choosing trimming parameters.
- GC and duplication patterns: interpret them in the context of the assay. Amplicon, RNA-seq, targeted panels, and whole-genome libraries have different expected profiles.
FastQC and MultiQC: a small reproducible check
With FastQC and MultiQC installed, a basic check can be recorded as:
mkdir -p qc/fastqc qc/multiqc logs
gzip -t raw_data/*.fastq.gz
fastqc --threads 4 --outdir qc/fastqc raw_data/*.fastq.gz 2>&1 | tee logs/fastqc.log
multiqc qc/fastqc --outdir qc/multiqc 2>&1 | tee logs/multiqc.log
Keep the tool versions, command lines, input checksums, reference build, and QC reports with the project. A red or amber FastQC module is a signal to investigate, not an automatic reason to discard the entire dataset.
Quality score is not the same as coverage or mapping quality
Base quality describes confidence in an individual base call. Coverage describes how many reads support a location. Mapping quality describes confidence in where a read aligned. Variant review needs these signals together, along with strand balance, allele fraction, local sequence context, platform artefacts, and an appropriate reference. A high average Q score cannot rescue poor alignment or contamination.
Should you trim every low-quality base?
No universal Q20 or Q30 trimming rule is valid for every assay. Trimming can remove useful sequence and may alter read-length distributions. Compare the pre- and post-trimming reports, use parameters justified for the library and downstream tool, and retain the exact command and version. For some modern aligners and workflows, adapter handling and quality-aware alignment may be preferable to aggressive trimming.
Common mistakes
- Reading the fourth line as a numeric score without accounting for ASCII encoding.
- Using a Phred+64 assumption on a Phred+33 file.
- Calling a dataset “good” from one average score while ignoring cycle-specific failures.
- Confusing Q score with read depth, mapping quality, or variant quality.
- Running clinical samples on an unmanaged personal computer without the required privacy, retention, and audit controls.
Practical checklist
- Record the sequencer, run software, conversion settings, and quality encoding.
- Verify checksums and paired-read counts.
- Run FastQC and combine reports with MultiQC.
- Investigate adapters, quality decay, GC bias, duplication, and overrepresented sequences in assay context.
- Document any trimming or filtering and compare reports before and after.
- Keep the original FASTQ files unchanged and work from a traceable copy.
Sources and further reading
- Cock et al., The Sanger FASTQ file format
- Illumina BaseSpace: Quality-score encoding
- Illumina: Sequencing quality scores
- NCBI SRA file-format guide
- FastQC documentation
- MultiQC documentation
