Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Operations

Operations summarize a field over all records, or over each group. Operation names are case-insensitive. Parameters follow a colon: perc:90, trimmean:0.1.

For operations that transform each record instead, see Per-row operations and modes.

Counting and text

OperationResult
countNumber of values, including text and empty fields
countuniqueNumber of distinct values (case-sensitive unless -i)
unique, uniqSorted, comma-separated list of distinct values
collapseComma-separated list of all values, in input order
firstFirst value
lastLast value
randOne value chosen at random (reproducible with --seed)

-c X (--collapse-delimiter) changes the separator used by unique and collapse.

Basic statistics

OperationResult
sumSum
meanArithmetic mean
min, maxMinimum, maximum
absmin, absmaxValue with the smallest or largest magnitude, keeping its sign
rangemax minus min

Other means

OperationResult
geomeanGeometric mean
harmmeanHarmonic mean
msMean of the squared values
rmsRoot mean square
trimmean[:P]Mean after removing the fraction P of values from each end (default 0; 0.5 gives the median)
wmean VALUE:WEIGHTWeighted mean, with finite nonnegative contribution weights

wmean takes one value/weight pair per result, for example wmean price:units with named fields or wmean 2:3 with positions. It supports text or quoted CSV, headers, custom result names and adjacent or sorted groups. A nonempty calculation needs positive retained weight. --narm omits whole pairs after checking supplied partners. See Weighted mean for examples, missing values and numerical range limits.

Quantiles and order statistics

OperationResult
medianMiddle value, or the mean of the two middle values
q1, q3First and third quartiles
iqrInterquartile range, q3 minus q1
perc[:N]Nth percentile, for integer N from 1 to 100 (default 95)
modeMost frequent value (the smallest, on ties)
antimodeLeast frequent value (the smallest, on ties)

Quartiles and percentiles use linear interpolation between order statistics (Hyndman and Fan’s type 7, the default in R and NumPy).

Dispersion

OperationResult
pvar, svarPopulation and sample variance
pstdev, sstdevPopulation and sample standard deviation
madMedian absolute deviation, scaled by 1.4826
madrawMedian absolute deviation, unscaled

Shape and normality

OperationResult
pskew, sskewPopulation and sample skewness
pkurt, skurtPopulation and sample excess kurtosis
jarqueJarque–Bera normality test p-value
dpoD’Agostino–Pearson omnibus normality test p-value

Paired statistics

These take a pair of fields, LEFT:RIGHT:

OperationResult
pcov, scovPopulation and sample covariance
ppearson, spearsonPopulation and sample Pearson correlation coefficient
dotprodDot product
$ printf '1\t2\n2\t4\n3\t7\n' | fastmash scov 1:2 spearson 1:2
2.5	0.99339926779878

spearson is the sample Pearson coefficient, not Spearman’s rank correlation.

Small samples and missing data

  • jarque and dpo keep GNU datamash’s formulas, including its tail cancellation: very small p-values can print as 0.
  • Sample statistics need enough values: svar and sstdev give nan for a single value; sskew needs at least three and skurt at least four.
  • When every value in a group is missing (with --narm), count and sum give 0; mean, median and most statistics give nan; min and max give -inf and inf; first and last give N/A.
  • An empty input produces no output line.

Memory use

Most operations keep a running total. Quantiles, modes, dispersion, shape, normality and paired statistics must keep every value of a group, and text operations like unique and collapse keep their text. Memory then grows with the size of the largest group, and this storage is not spilled to disk. Related operations on the same field share storage: quantiles, mode and antimode share one sorted copy, the variance family another, and the moment and normality statistics another.