Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Fastmash

Fast command-line statistics for delimited text.

Fastmash summarizes, groups and reshapes delimited text and quoted CSV. It accepts the command language of GNU datamash, so if you already use datamash, you already know how to use Fastmash.

It also reads quoted CSV directly, gives your result columns stable names, calculates weighted means and selects the highest or lowest complete records. Inspect a table’s health or compare summaries from two exports with the same field selectors and portable arithmetic.

$ cat readings.tsv
site	reading
west	2
east	4
west	6
west	7
$ fastmash -H -s -g site count reading median reading < readings.tsv
GroupBy(site)	count(reading)	median(reading)
east	1	4
west	3	6

Install it with one command, or see other ways:

curl -fsSL https://fastmash.io/install.sh | sh

Fastmash is faster than GNU datamash*

Up to four times faster, on real jobs: quantiles, grouped summaries and millions of decimals.

How many times faster Fastmash is than GNU datamash on the core jobs, per host Intel laptop, native Linux AMD desktop, native Linux 0× 1× 2× 3× 4× 5× Exon-count quantiles (RefGene) Exon-count quantiles (RefGene), Intel laptop, native Linux: GNU 112.8 ms, Fastmash 26.6 ms 4.2× Exon-count quantiles (RefGene), AMD desktop, native Linux: GNU 50.3 ms, Fastmash 13.1 ms 3.8× Transcripts per gene (RefGene) Transcripts per gene (RefGene), Intel laptop, native Linux: GNU 112.0 ms, Fastmash 61.0 ms 1.8× Transcripts per gene (RefGene), AMD desktop, native Linux: GNU 45.6 ms, Fastmash 26.5 ms 1.7× Exon statistics per gene (RefGene) Exon statistics per gene (RefGene), Intel laptop, native Linux: GNU 159.4 ms, Fastmash 83.1 ms 1.9× Exon statistics per gene (RefGene), AMD desktop, native Linux: GNU 62.2 ms, Fastmash 36.5 ms 1.7× Many small groups (genes) Many small groups (genes), Intel laptop, native Linux: GNU 160.4 ms, Fastmash 105.0 ms 1.5× Many small groups (genes), AMD desktop, native Linux: GNU 55.0 ms, Fastmash 38.7 ms 1.4× Gene example (datamash manual) Gene example (datamash manual), Intel laptop, native Linux: GNU 7.3 ms, Fastmash 3.8 ms 1.9× Gene example (datamash manual), AMD desktop, native Linux: GNU 3.0 ms, Fastmash 1.6 ms 1.9× Sum and mean, 1M decimals Sum and mean, 1M decimals, Intel laptop, native Linux: GNU 324.7 ms, Fastmash 122.5 ms 2.7× Sum and mean, 1M decimals, AMD desktop, native Linux: GNU 105.9 ms, Fastmash 46.8 ms 2.3× Sum and mean, 100k decimals Sum and mean, 100k decimals, Intel laptop, native Linux: GNU 36.8 ms, Fastmash 14.6 ms 2.5× Sum and mean, 100k decimals, AMD desktop, native Linux: GNU 10.9 ms, Fastmash 5.3 ms 2.1× Grouped decimals, 100k rows Grouped decimals, 100k rows, Intel laptop, native Linux: GNU 69.9 ms, Fastmash 38.5 ms 1.8× Grouped decimals, 100k rows, AMD desktop, native Linux: GNU 25.6 ms, Fastmash 14.8 ms 1.7× One dominant group One dominant group, Intel laptop, native Linux: GNU 21.9 ms, Fastmash 11.9 ms 1.8× One dominant group, AMD desktop, native Linux: GNU 9.0 ms, Fastmash 4.5 ms 2.0× Wine quality by grade Wine quality by grade, Intel laptop, native Linux: GNU 6.4 ms, Fastmash 3.2 ms 2.0× Wine quality by grade, AMD desktop, native Linux: GNU 2.8 ms, Fastmash 1.4 ms 2.0× Large sort with disk spill Large sort with disk spill, Intel laptop, native Linux: GNU 572.4 ms, Fastmash 694.1 ms 0.8× Large sort with disk spill, AMD desktop, native Linux: GNU 180.0 ms, Fastmash 309.7 ms 0.6× Times faster than GNU datamash 1.9, by median elapsed time. The dashed line is the same speed; shorter bars are slower. 0.1.0 release binary 1f1fec94, 2026-10-05/06

* In 0.1.0, some disk-spilling sorts and the distinct-value geometric-mean job are slower on native Linux. All results and the method.

And better in other ways

FeatureGNU datamash 1.9Fastmash
ResultsDepend on the CPU and the C math libraryIdentical on every machine: all arithmetic in software
LocalesNeed locale data installed; silently fall back to C without itBuilt in: numbers in 318 locales, sorting in 194
perc:100Can read past the end of the valuesReturns the largest value
rmdup on several key fieldsCrashes on an assertionA clear refusal
Malformed parameters such as strbin:2.0Misleading errors or undefined behaviorAn accurate error message
crosstab labels over 511 bytesTruncatedKept in full

Everything else works the same: over 70 operations and modes, the same options, field selectors, grouping, headers and output. For most scripts, switching means changing one word. Migrating from datamash · All differences

Get started

Fastmash is open source under the MIT or Apache-2.0 license. Source on GitHub.