Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Benchmarks

Speed is Fastmash’s reason to exist, so every release is measured against GNU datamash 1.9 on the same machines, with the same data, and the results are published in full: wins and losses.

Fastmash 0.1.0

How many times faster Fastmash is than GNU datamash on the core jobs, per host Intel laptop, native Linux AMD desktop, native Linux 0× 1× 2× 3× 4× 5× Exon-count quantiles (RefGene) Exon-count quantiles (RefGene), Intel laptop, native Linux: GNU 112.8 ms, Fastmash 26.6 ms 4.2× Exon-count quantiles (RefGene), AMD desktop, native Linux: GNU 50.3 ms, Fastmash 13.1 ms 3.8× Transcripts per gene (RefGene) Transcripts per gene (RefGene), Intel laptop, native Linux: GNU 112.0 ms, Fastmash 61.0 ms 1.8× Transcripts per gene (RefGene), AMD desktop, native Linux: GNU 45.6 ms, Fastmash 26.5 ms 1.7× Exon statistics per gene (RefGene) Exon statistics per gene (RefGene), Intel laptop, native Linux: GNU 159.4 ms, Fastmash 83.1 ms 1.9× Exon statistics per gene (RefGene), AMD desktop, native Linux: GNU 62.2 ms, Fastmash 36.5 ms 1.7× Many small groups (genes) Many small groups (genes), Intel laptop, native Linux: GNU 160.4 ms, Fastmash 105.0 ms 1.5× Many small groups (genes), AMD desktop, native Linux: GNU 55.0 ms, Fastmash 38.7 ms 1.4× Gene example (datamash manual) Gene example (datamash manual), Intel laptop, native Linux: GNU 7.3 ms, Fastmash 3.8 ms 1.9× Gene example (datamash manual), AMD desktop, native Linux: GNU 3.0 ms, Fastmash 1.6 ms 1.9× Sum and mean, 1M decimals Sum and mean, 1M decimals, Intel laptop, native Linux: GNU 324.7 ms, Fastmash 122.5 ms 2.7× Sum and mean, 1M decimals, AMD desktop, native Linux: GNU 105.9 ms, Fastmash 46.8 ms 2.3× Sum and mean, 100k decimals Sum and mean, 100k decimals, Intel laptop, native Linux: GNU 36.8 ms, Fastmash 14.6 ms 2.5× Sum and mean, 100k decimals, AMD desktop, native Linux: GNU 10.9 ms, Fastmash 5.3 ms 2.1× Grouped decimals, 100k rows Grouped decimals, 100k rows, Intel laptop, native Linux: GNU 69.9 ms, Fastmash 38.5 ms 1.8× Grouped decimals, 100k rows, AMD desktop, native Linux: GNU 25.6 ms, Fastmash 14.8 ms 1.7× One dominant group One dominant group, Intel laptop, native Linux: GNU 21.9 ms, Fastmash 11.9 ms 1.8× One dominant group, AMD desktop, native Linux: GNU 9.0 ms, Fastmash 4.5 ms 2.0× Wine quality by grade Wine quality by grade, Intel laptop, native Linux: GNU 6.4 ms, Fastmash 3.2 ms 2.0× Wine quality by grade, AMD desktop, native Linux: GNU 2.8 ms, Fastmash 1.4 ms 2.0× Large sort with disk spill Large sort with disk spill, Intel laptop, native Linux: GNU 572.4 ms, Fastmash 694.1 ms 0.8× Large sort with disk spill, AMD desktop, native Linux: GNU 180.0 ms, Fastmash 309.7 ms 0.6× Times faster than GNU datamash 1.9, by median elapsed time. The dashed line is the same speed; shorter bars are slower. 0.1.0 release binary 1f1fec94, 2026-10-05/06

Mean of the two sessions’ median elapsed times, in milliseconds; lower is better. Measured with the 0.1.0 release binaries themselves (fastmash sha256 1f1fec94…), the same bytes you download. The chart is generated from core-jobs.tsv by scripts/benchmark_chart.py.

The animated demo replays the Intel laptop’s RefGene quartile job below. Playback is slowed 20 times so that both runs are visible; its clocks show the measured times.

JobDataIntel laptop, native Linux
GNU → Fastmash
AMD desktop, native Linux
GNU → Fastmash
Quartiles of exon countsRefGene annotations112.8 → 26.6 (4.2×)50.3 → 13.1 (3.8×)
Transcripts per geneRefGene annotations112.0 → 61.0 (1.8×)45.6 → 26.5 (1.7×)
Exon statistics per geneRefGene annotations159.4 → 83.1 (1.9×)62.2 → 36.5 (1.7×)
Many small groupsSynthetic grouping data160.4 → 105.0 (1.5×)55.0 → 38.7 (1.4×)
Gene example from the datamash manualGene annotations7.3 → 3.8 (1.9×)3.0 → 1.6 (1.9×)
Sum and mean of a million decimalsSynthetic324.7 → 122.5 (2.7×)105.9 → 46.8 (2.3×)
Sum and mean of 100,000 decimalsSynthetic36.8 → 14.6 (2.5×)10.9 → 5.3 (2.1×)
Grouped decimals, 100,000 rowsSynthetic69.9 → 38.5 (1.8×)25.6 → 14.8 (1.7×)
One dominant groupSynthetic21.9 → 11.9 (1.8×)9.0 → 4.5 (2.0×)
Wine quality by gradeUCI Wine Quality6.4 → 3.2 (2.0×)2.8 → 1.4 (2.0×)
Large sort with disk spillSynthetic572.4 → 694.1 (0.82×)180.0 → 309.7 (0.58×)

Across 71 combinations of jobs and settings, 28 on the Intel laptop and 19 on the native AMD desktop meet the full speed-win rule: at least 20% and 5 ms saved in both sessions, with consistent paired runs. Both hosts have wins in all four non-startup workload families. Every job gives the expected output. The named slower and inconclusive cases below retain their individual dispositions. Full tables and the rules are on the method page.

Where GNU datamash is still faster

These jobs cross the material elapsed slowdown rule in both sessions. Each cell shows the two session medians in milliseconds:

HostSet and jobGNU datamashFastmashExtra time per run
AMD desktopCore, disk spill176.0 / 184.0313.4 / 306.0122–137 ms
AMD desktopDedicated disk spill177.0 / 175.6303.8 / 319.5127–144 ms
AMD desktopLarger disk spill264.6 / 261.5457.8 / 477.3193–216 ms
Intel laptopCore, disk spill573.7 / 571.0694.5 / 693.8121–123 ms
Intel laptopMany keys788.6 / 772.0959.4 / 970.9171–199 ms
Intel laptopGeometric mean, 200,000 distinct values37.3 / 37.544.7 / 44.97.4 ms

These individual costs are accepted for 0.1.0. The AMD spill medians reach 1.83 times GNU’s time on these invocations. Lower CPU work and charged peak memory in the spill measurements are separate observations; the elapsed verdicts remain regressions. The geometric mean keeps the existing numerical precision and range checks.

Three further comparisons are inconclusive under the two-session rule:

HostSet and jobGNU datamash (ms)Fastmash (ms)
AMD desktopMany keys248.6 / 380.9423.2 / 388.5
Intel laptopDedicated disk spill580.7 / 961.9714.8 / 832.3
Intel laptopLarger disk spill1943.6 / 892.41088.4 / 1069.7

These uncertain results are also accepted as named limitations. They establish neither a GNU competitiveness pass nor a win. All samples remain in the record, including an Intel dedicated-spill Fastmash run of 4.887 seconds. Session medians give no tail-latency guarantee or upper bound for other input sizes.

The decision retains the measured implementation: every predecessor elapsed and CPU guard passes, both hosts meet the workload-family and small-command rules, and all checked outputs match. The numbers concern the recorded warm-cache, disk-backed, 512 MiB command cgroup conditions on these two native Linux hosts. Untimed observations confirm actual spilling and correct output; they do not establish an I/O bottleneck.

Run them yourself

The benchmark kit runs eight of these jobs on your machine with both programs, checks that they agree, and prints a table you can share:

curl -fsSLO https://raw.githubusercontent.com/pederbe/fastmash/main/bench/fastmash-bench.py
python3 fastmash-bench.py

The best benchmark is your own job. Compare both programs on your data:

export LC_ALL=C.UTF-8
time datamash -s -g 1 median 2 < your-data.tsv > /dev/null
time fastmash -s -g 1 median 2 < your-data.tsv > /dev/null

For stable numbers, repeat each command several times, alternate the two programs, and use a quiet machine. hyperfine does this for you:

hyperfine --warmup 2 \
  'datamash -s -g 1 median 2 < your-data.tsv' \
  'fastmash -s -g 1 median 2 < your-data.tsv'

If Fastmash is slower on a job you care about, please tell us: slow real-world jobs are what we optimize next.