Concepedia

Publication | Open Access

ANGSD: Analysis of Next Generation Sequencing Data

3.1K

Citations

34

References

2014

Year

TLDR

High‑throughput DNA sequencing technologies generate vast amounts of data, creating a need for fast, flexible, and memory‑efficient implementations to analyze thousands of samples simultaneously. The study introduces ANGSD, a multithreaded program suite designed to facilitate large‑scale sequencing data analysis. ANGSD is a multithreaded, open‑source C/C++ program that calculates summary statistics, performs association mapping and population genetic analyses directly on raw sequencing data or genotype likelihoods, supports multiple input formats such as BAM and imputed Beagle genotype probability files, is.

Abstract

BackgroundHigh-throughput DNA sequencing technologies are generating vast amounts of data. Fast, flexible and memory efficient implementations are needed in order to facilitate analyses of thousands of samples simultaneously.ResultsWe present a multithreaded program suite called ANGSD. This program can calculate various summary statistics, and perform association mapping and population genetic analyses utilizing the full information in next generation sequencing data by working directly on the raw sequencing data or by using genotype likelihoods.ConclusionsThe open source c/c++ program ANGSD is available at http://www.popgen.dk/angsd. The program is tested and validated on GNU/Linux systems. The program facilitates multiple input formats including BAM and imputed beagle genotype probability files. The program allow the user to choose between combinations of existing methods and can perform analysis that is not implemented elsewhere.

References

YearCitations

Page 1