Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Compared with other tools

Fastmash runs GNU datamash commands with portable results, and extends that workflow with quoted CSV, weighted means, record selection, table health and dataset comparison. Other good tools cover much more ground. This page says honestly where each one fits. Facts about the other tools come from their own documentation (checked September 2026); tell us if something has changed.

At a glance

FastmashGNU datamashMillerqsvcsvtk
Command languagedatamash’sdatamash’sIts own verbs and DSLIts own subcommandsIts own subcommands
Written inRustCGoRustGo
LicenceMIT or Apache-2.0GPL-3.0-or-laterBSD-2-ClauseMITMIT
InputDelimited text and explicit quoted CSVTab or single-byte delimited textCSV (RFC 4180), TSV, JSON and moreCSV, TSV, Excel, Parquet and moreCSV and TSV
Quoted CSV fieldsYes, in supported calculation, selection and comparison modesNoYesYesYes
Grouped statisticsYesYesYes (stats1 -g)Through pivotp or sqlpYes (summary -g)
Joins, filters, reshapingNoNoYesYesYes
PlatformsLinux x86-64Linux, macOS, Windows and othersLinux, macOS, Windows, BSDLinux, macOS, WindowsLinux, macOS, Windows, BSD

GNU datamash

The original, and the reference Fastmash is measured against. It runs almost everywhere and is packaged by every Linux distribution, conda-forge and Homebrew.

Choose Fastmash when you run datamash commands on Linux and want them faster (up to 4 times on the published benchmarks), want the same digits on every machine, or want clear errors where datamash 1.9 aborts or reads past its data. Your commands don’t change; the differences are listed on one page.

Stay with GNU datamash when you need macOS, native Windows or a non-x86 CPU, or when you need a GNU-maintained tool. Some disk-spilling jobs still favor GNU on native Linux; the benchmarks show the costs and measurement conditions.

Miller

A complete toolkit for record-oriented data: it reads and writes CSV (with quoting), TSV, JSON and several other formats, and has a full expression language for filtering, computing new fields, joining, sorting, sampling and reshaping. Its stats1 verb computes grouped statistics (count, sum, mean, median and any percentile, mode, variance, skewness and more) without sorting the input first.

Choose Miller when your data is JSON, or the job needs more than statistics: filtering, joining or transforming records. Choose Fastmash when the job is a datamash-style summary of delimited text and speed or exact datamash output matters.

qsv

A very fast, very broad CSV toolkit in Rust: indexing, joins (through Polars), SQL queries, validation, sampling, format conversion to Parquet, Excel and databases, and whole-column statistics. Grouped aggregation goes through pivotp (count, sum, mean, median, quantiles, first, last) or sqlp.

Choose qsv when you work with large CSV files and want one tool for exploring, querying and converting them. Choose Fastmash when you want datamash’s operations and syntax, for example sample and population standard deviations, skewness and kurtosis, mode, collapse or crosstab per group, with datamash-identical output.

csvtk

A cross-platform CSV/TSV toolkit popular in bioinformatics, with subcommands for selecting, filtering, joining, sorting, reshaping, converting and plotting. csvtk summary computes grouped statistics (count, sum, mean, median, quartiles, standard deviation, first, last, unique values, collapse), printing two decimals by default.

Choose csvtk when you want a single, familiar toolkit for everyday table work, including quoted CSV and compressed files. Choose Fastmash when you need datamash’s full set of statistics and exact datamash output, or its speed on large inputs.

Using them together

These tools combine well. Fastmash can read quoted CSV directly:

fastmash --csv -H -s -g region mean price < data.csv

Use another tool to filter or join records, then Fastmash to summarize them. For ordinary text workflows, a quote-aware conversion remains useful:

mlr --icsv --otsv cat data.csv | fastmash -H -s -g region mean price