Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2023 Jan 1;39(1):btad019.
doi: 10.1093/bioinformatics/btad019.

SAMStat 2: quality control for next generation sequencing data

Affiliations

SAMStat 2: quality control for next generation sequencing data

Timo Lassmann. Bioinformatics. .

Abstract

Motivation: SAMStat is an efficient program to extract quality control metrics from fastq and SAM/BAM files. A distinguishing feature is that it displays sequence composition, base quality composition and mapping error profiles split by mapping quality. This allows users to rapidly identify reasons for poor mapping including the presence of untrimmed adapters or poor sequencing quality at individual read positions.

Results: Here, we present a major update to SAMStat. The new version now supports paired-end and long-read data. Quality control plots are drawn using the ploty javascript library.

Availability and implementation: The source code of SAMStat and code to reproduce the results are found here: https://github.com/timolassmann/samstat.

PubMed Disclaimer

Figures

Fig. 1.
Fig. 1.
Selection of SAMStat QC plots. (a) Base composition of a polyA+ library. (b) Length distribution of a PacBio library, split by mapped and unmapped reads. (c) Mean base quality of a 10× single cell library

References

    1. Andrews S. et al. (2010) Fastqc: a quality control tool for high throughput sequence data. https://www.bioinformatics.babraham.ac.uk/projects/fastqc/.
    1. Bonfield J.K. et al. (2021) HTSlib: C library for reading/writing high-throughput sequencing data. GigaScience, 10, giab007. - PMC - PubMed
    1. Ewels P. et al. (2016) MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinformatics, 32, 3047–3048. - PMC - PubMed
    1. Graubert A. et al. (2021) RNA-SeQC 2: efficient RNA-seq quality control and quantification for large cohorts. Bioinformatics, 37, 3048–3050. - PMC - PubMed
    1. Lassmann T. et al. (2011) SamStat: monitoring biases in next generation sequencing data. Bioinformatics, 27, 130–131. - PMC - PubMed

Publication types