Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Quoted CSV and custom result names

Fastmash supports quoted CSV input and output, and optional names for calculated output columns. CSV is selected explicitly; ordinary delimited text stays the default.

WorkflowCSV input/outputResult names
Aggregate and per-row reportsYes, including supported grouping and full copied fieldsCalculated output columns with an explicit output header
Weighted meanYes, including adjacent and sorted groupsOne name per value/weight result
Top-N selectionYes, including adjacent and sorted groupsNo: selection copies the source fields
Dataset comparisonYes; both sources use the same input formatOne name throughout each summary’s before/after/change columns
Crosstab, field modes and legacy table modesNoNo
Table healthNo; ordinary text input and readable or TSV outputNo

Custom result names

Give calculated output columns stable names in your reports. Select Output headers with --header-out or -H, then repeat --result-name=INDEX:NAME for the labels you want to replace. The separate form --result-name INDEX:NAME works too.

For example, a report can give its sum a stable schema name while retaining the generated label for its mean:

printf 'reading\n2\n4\n' | LC_ALL=C fastmash -H --result-name=1:total sum reading mean reading
total	mean(reading)
6	3

Indices follow Operation-produced output order after Selector expansion, starting at 1. A list or range contributes each expanded result, repeated Operations contribute their results again, and a paired Operation contributes one result. Each field selected by cut also contributes a result. Grouping keys and the copied fields from --full do not consume indices or receive replacement labels.

In this example, sum 2,1-2 produces three results in the order 2, 1, 2. Names can be supplied out of order, and unnamed results keep their generated labels:

printf '1\t2\n3\t4\n' | LC_ALL=C fastmash --header-out --result-name=3:again --result-name=1:total sum 2,1-2
total	sum(field-1)	again
6	4	6

Names replace labels only. They do not change Selector lookup, required-field checks, calculation values, output order, or when headers appear. Empty input still produces no invented header; with -H, header-only input retains its normal Output header without a result row. Later input or output failures can leave partial output, as with ordinary generated headers.

INDEX must be an unsigned, nonzero decimal integer within the expanded result count. Repeating a target, omitting a name, or supplying an empty name is an error. Duplicate name values are allowed. Split at the first colon, so a name can itself contain colons. Name bytes are preserved as supplied by the operating system. For text output, names cannot contain the effective Output delimiter or Record terminator. CSV output allows comma, quote, colon, CR and LF bytes in names and encodes them as one field. Invalid names fail before any output.

Result names work for ordinary aggregate and Per-row operations, including adjacent and sorted groups. Crosstab, table modes (transpose, check, rmdup) and field modes (reverse, noop) refuse naming before reading input. Only the exact --result-name spelling is accepted; existing GNU option abbreviations retain their normal meanings. --vnlog alone does not satisfy the requirement for explicit Output headers; add --header-out or -H.

GNU datamash 1.9 generates labels such as sum(reading). Stable custom report labels otherwise need a wrapper or a downstream renaming step. Naming inside the Command avoids that extra step and follows the existing header timing. This is a convenience for report schemas, with no claim of faster computation.

CSV output from delimited text

Use --csv-out to produce quoted CSV from ordinary delimited-text input. It works with aggregate and Per-row operations, including adjacent or sorted groups, explicit headers and copied --full fields. Repeating the switch is idempotent. It does not enable either Input or Output headers. Top-N selection also supports CSV output from complete selected text Records, including sorted Groups; it adds no calculation columns.

For example, tab-separated observations can become a CSV report with a stable result label, including keys and names containing commas or quotation marks:

printf 'place\treading\nlab,west\t2\nlab,west\t4\n' | LC_ALL=C fastmash --csv-out -H -g place --result-name='1:total,"daily"' sum reading
GroupBy(place),"total,""daily"""
"lab,west",6

Every logical output field is encoded separately: complete generated labels, custom names, Grouping keys, formatted results and each copied --full field. Commas separate fields and LF ends each emitted Record. A field containing a comma, double quote, CR or LF is quoted, with internal double quotes doubled. An empty field prints as "". Embedded line breaks retain their bytes. Output is canonical; it does not retain an input field’s quoting style or line ending.

Decimal-comma formatted numbers and collapse or unique lists stay within one quoted result field. The Collapse delimiter is still the list’s internal separator; CSV adds no escaping scheme inside those lists. Encoding preserves the result bytes, including non-UTF-8 and NUL bytes, under the existing Operation and header display rules. For example, raw keys and copied fields retain NUL, while selected text results and generated field-name displays stop at NUL as they do in ordinary output. A leading UTF-8 BOM remains field data.

Output-only CSV allows text-input -t, -W and -C. An explicit --output-delimiter, -z or --vnlog conflicts with CSV output and fails with status 1, regardless of option order or a redundant delimiter value. Crosstab, legacy Table modes and Field modes refuse CSV with status 77 before input. Health reports remain text-only and reject CSV switches using their status-1 calculation-control diagnostic. Only the exact CSV switch spellings are accepted; existing GNU option abbreviations retain their meanings.

CSV output retains empty-input and header-only timing and the normal failure policy. A later input, calculation or transport failure can leave partial output.

Calculations on quoted CSV

Use --csv-in for strict quoted CSV input with ordinary text output, or --csv for CSV input and output together. --csv is equivalent to --csv-in --csv-out. Repeating these exact-only switches is idempotent; none enables headers. Top-N selection supports decoded CSV input for whole-dataset, adjacent-Group and sorted-Group selection with either output format. Sorted selection retains complete decoded Records and original-input ties through bounded Spill. Its copied Output header uses the original source header or first accepted source Record, independently of sorted encounter order. Its ranking Field uses the decoded field-local numerical contract, and copied data preserves complete bytes. Aggregate and Per-row calculations support their existing positional lists, numeric ranges, escaped named Selectors, paired Selectors and parameters. Aggregate grouping and --full work with decoded CSV fields. Add -s to bring equal decoded Grouping keys together, using bounded Spill when the input exceeds the maintained sort-memory target. Per-row operations retain their ordinary conflict with Grouping keys. With no Grouping keys, -s keeps the streaming workflow and source order, including Per-row and --full reports. Output-only CSV retains sorted text-input workflows.

For example, a quoted export can be summed directly, without a quote-aware preprocessing step, while preserving a comma in its decoded column name:

printf '"amount,raw",note\r\n"2","two\nlines"\n4,"a""b"' | LC_ALL=C fastmash --csv -H --result-name='1:total,"daily"' sum 'amount\,raw'
"total,""daily"""
6

Grouping keys compare decoded values, so west and "west" belong to the same adjacent Group. Repeated keys in a later run stay separate without sorting:

printf '"region,name",value\n"west,coast",2\n"west,coast",4\neast,3\n"west,coast",5\n' | LC_ALL=C fastmash --csv -H -g 'region\,name' --result-name=1:total sum value
"GroupBy(region,name)",total
"west,coast",6
east,3
"west,coast",5

Multiple and named keys retain their normal request order, exact name lookup and first-duplicate-name rules. Adjacent equality is byte-based, with ASCII folding under -i where the selected locale supports it. Equal-length keys compare up to NUL, while displayed keys retain their complete byte spans. Locale-aware numeric input and output keep their existing rules; a language locale does not silently merge differently spelled adjacent keys.

Sorting the same export merges the separate west,coast runs:

printf '"region,name",value\n"west,coast",2\n"west,coast",4\neast,3\n"west,coast",5\n' | LC_ALL=C fastmash --csv -H -s -g 'region\,name' --result-name=1:total sum value
"GroupBy(region,name)",total
east,3
"west,coast",11

Sorting compares complete decoded keys with the existing locale and case policy, then preserves source order for tied keys. Group equality remains distinct from locale ordering: differently spelled keys which the locale ties can still form separate adjacent Groups. Headers bind before sorting; without an Input header, a generated Output header takes its field width from the first sorted Record. Copied Records keep their actual widths and original logical Record/physical-line locations through reordering and representative selection.

CSV sorting uses the same sort-memory policy as ordinary text, including FASTMASH_SORT_MEMORY_BYTES, and spills sorted runs to temporary files when needed. Complete decoded fields and their original input locations survive sorting. The memory target covers sort chunks, not the whole process; an oversized Record, merge heads and retained statistical or text samples have their own storage needs. Samples are not spilled. Temporary I/O errors fail with status 1, and checked allocation failures refuse with status 77. Later failures may leave earlier output.

Full reports copy each decoded source field before the calculation results. For example, a Per-row report can retain a comma in a source field while selecting its numeric field with a custom result label:

printf 'place,value\n"west,coast",2\neast,4\n' | LC_ALL=C fastmash --csv -H --full --result-name=1:reading cut value
place,value,reading
"west,coast",2,2
east,4,4

For an aggregate, the representative starts as the Group’s first Record; first, last, rand and improving extrema may select another Record under their existing rules. Extrema ties retain the earlier selection. With multiple Operations the representative need not supply every result. Aggregate --full retains its existing GNU compatibility warning; Per-row --full does not warn. Copied-field labels replace GroupBy labels, and Result-name indices still count only Operation-produced results. Without sorting, header width comes from the first logical Record; each copied data Record keeps its actual width, without padding or filler. Required fields must still exist, including with --no-strict or --filler. Copied fields preserve complete bytes, including NUL, independently of an Operation’s display rules.

A double quote opens quoting only at the start of a field. Inside quotes, commas, CR and LF are data, and two double quotes decode to one. After a closing quote, only comma, LF, CRLF or end of input is allowed. Quotes within unquoted fields, stray bytes after closing quotes, unfinished quoted fields and bare CR outside quotes are errors, including in unused fields. Spaces are data, with no implicit trimming. LF and CRLF endings may be mixed; a complete final Record may omit its ending.

An empty file has no Records. Each blank LF or CRLF Record has one empty field; quoted and unquoted empty fields decode identically. A trailing comma supplies a final empty field, while a terminal Record ending supplies no extra Record. Different widths are permitted when all requested fields exist. An absent field still errors; an existing empty field is valid text but invalid numeric input. --narm retains the existing exact, case-insensitive NA/N/A/NaN matching and independent paired-side compaction. It does not skip empty fields or trim spaces.

Decoding preserves bytes, including tabs, NUL, non-UTF-8 and internal CRLF. Operations keep their existing display and transformation rules: cut stops display at NUL, while base64 and checksums consume the complete decoded field. The active numerical locale and presentation settings apply to each isolated decoded numeric field. An unused numeric neighbor cannot affect conversion. Ordinary text numerical conversion keeps its existing behavior.

A leading UTF-8 BOM is data. Before an unquoted Input header it becomes part of the name and generated label, following the normal byte-escaped Selector rules. Before a quoted first field it is data followed by an illegal quote, so that export fails. There is no BOM stripping or transcoding. This dialect has stated byte and blank-Record extensions; it is not a claim of full RFC 4180 conformance.

With --csv-in, text results use the normal tab default and accept an explicit --output-delimiter. Decoded delimiters or line breaks in result fields may make that text report ambiguous; use --csv when field boundaries must survive output. CSV input conflicts with explicit -t, -W and -C, including their GNU abbreviations and redundant comma delimiters. Either CSV direction conflicts with -z and --vnlog, in either option order, with Error status 1. Legacy modes refuse CSV before input; Health rejects all CSV switches and remains text-only.

Unsorted input is read incrementally, with checked growth for decoded Record and Field storage. Adjacent Grouping and Per-row execution retain owned decoded Records without losing field boundaries or original locations when representative and spare storage swap. Statistical sample retention keeps its existing memory policy; parsing does not make samples spillable.

CSV validation errors identify the original logical Record number, counting an Input header, and its starting physical line. Syntax errors also identify the detected physical line and byte position, counting raw bytes including quotes; unfinished quotes identify unexpected end of input. LF inside quotes advances the physical line; CRLF counts as one line break. Header lookup errors include the header’s location when it exists. I/O failures keep their I/O context without inventing a Record or turning a read failure into a quoted-field EOF error. Malformed CSV uses status 1, checked Capacity refusals use 77, and output finalization retains normal precedence. Later failures can leave partial output.