Grouping and sorting
Groups
-g KEYS (also written groupby KEYS, grouping KEYS or gb KEYS) makes
Fastmash summarize each group of records that share the same key values.
Output rows start with the key values, followed by the results:
$ printf 'A\tx\t1\nA\tx\t2\nA\ty\t5\nB\ty\t4\n' | fastmash -g 1,2 sum 3
A x 3
A y 5
B y 4
A group is a run of adjacent records with equal keys. When the key changes, the group ends and its results are printed. This makes grouping a single pass over the input, with memory proportional to one group.
If equal keys are not adjacent, the same key appears in more than one group:
$ printf 'b\t2\na\t1\nb\t3\n' | fastmash -g 1 sum 2
b 2
a 1
b 3
Sorting first: -s
-s sorts the records by their keys before grouping, so every key forms one group:
$ printf 'b\t2\na\t1\nb\t3\n' | fastmash -s -g 1 sum 2
a 1
b 5
The sort is stable: records with equal keys keep their input order, so
first, last and collapse give predictable results.
Key order depends on the locale. In C and C.UTF-8,
keys are ordered byte by byte. In the supported language locales, such as
en_US.UTF-8, fr_FR.UTF-8 or pl_PL.UTF-8, keys are ordered alphabetically,
as GNU sort orders them, using built-in Unicode collation rules and glibc’s
table of the punctuation and symbols it ignores.
If your input is already sorted, leave out -s: the result is the same and
the job is faster.
Large inputs
When standard input is a file (fastmash -s -g 1 sum 2 < data.tsv) and its
keys repeat, Fastmash does not sort the records at all: it collects each
group’s records as they arrive and sorts only the groups, which is faster and
needs memory only for the groups (hash grouping). The result is the same as
sorting’s. Where it could differ, for example when a record lacks a key field
or a value is invalid, or where the groups turn out to be too many, Fastmash
reads the file again and sorts it, so the output and any error are exactly
what the sort gives.
Input from a pipe (cut -f 1,3 data.tsv | fastmash -s -g 1 sum 2) is grouped
the same way in a language locale such as en_US.UTF-8: Fastmash keeps what
it reads in memory, as the sort would, so that it can sort it if it has to. A
job that gives up late, near the end of a large input, then takes somewhat
longer than sorting would have. In C and C.UTF-8, piped input is sorted,
which needs less memory there. FASTMASH_PIPE_GROUPING=sort sorts piped
input in every locale, and =hash groups it in every locale where Fastmash
sorts in process. With -W, input whose columns are separated by the same
blanks for each value is grouped this way too; where the same key comes with
different blanks before it, Fastmash sorts. --vnlog and rand always sort;
FASTMASH_GROUPING=sort makes every job sort.
Otherwise Fastmash sorts in memory, and when the data is large it writes sorted chunks
to temporary files in TMPDIR (default /tmp) and merges them. The files are
anonymous: the operating system removes them automatically, even if Fastmash
is interrupted.
A chunk holds each record’s grouping keys and the fields its operations use
(the whole record with --full, in language locales, or when the field
separator is a letter, a digit or the decimal point, which a number could run
on into) and 32 bytes per record for sorting. A sort starts with a 64 MiB
chunk. When it fills, Fastmash holds the whole input in memory instead, if
that fits: within the buffer GNU sort would use for an input file, a quarter
of the available memory, a share of what is free under any cgroup memory limit
(one per processor) or on the host (one per running Fastmash), and with less
than half the swap in use. For piped input it doubles the chunk each time it
fills, within the same bounds. Otherwise it spills. In language locales the
records are prepared for sorting on several threads, whose batches take a
sixteenth of the first 64 MiB.
FASTMASH_SORT_MEMORY_BYTES fixes the chunk’s memory target instead, and
FASTMASH_SORT_TRACE (set to anything) reports each decision on standard
error. Neither is a limit on Fastmash’s total memory use.
In language locales, records are prepared for sorting on up to 8 threads,
chosen as GNU sort chooses its default: the processors available to
Fastmash, or OMP_NUM_THREADS if it is set, capped by OMP_THREAD_LIMIT.
Set OMP_NUM_THREADS=1 to sort on one thread.
In the C, POSIX and C.UTF-8 locales, some operations still use the
system sort command, through Fastmash’s sort supervisor: rand, geomean,
harmmean, ms, rms, the skewness and kurtosis operations, jarque,
dpo, the paired statistics, and rmdup with -s. This route writes its
temporary files in a private directory in TMPDIR; see
System requirements for what it needs. In language locales
every operation uses Fastmash’s own sorter.
Ignoring case: -i
-i compares keys ignoring ASCII letter case, for grouping and sorting. The
output shows each key as it first appeared. -i also affects countunique
and unique. Under the Turkic character types (tr_TR, az_AZ and a few
others), where glibc does not fold i and I, -i is refused; set
LC_CTYPE=C.UTF-8 to fold ASCII letters.
Pivot tables: crosstab
crosstab KEY1,KEY2 (or ct) builds a two-way table with one row per value of the first
key and one column per value of the second. By default each cell is a count;
you can name one operation instead:
$ printf 'a\tx\t1\na\ty\t2\nb\tx\t3\n' | fastmash -s crosstab 1,2 sum 3
x y
a 1 2
b 3 N/A
Missing combinations show the --filler text (default N/A). Rows and columns
are ordered byte by byte. Use -s unless each combination’s records are already
adjacent.