Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Highest and lowest records

Fastmash supports Top-N selection from ordinary delimited text or quoted CSV, across the whole dataset or separately within adjacent or sorted Groups. Select one numeric Field with top:N FIELD for the highest values or bottom:N FIELD for the lowest. N is mandatory, accepts leading zeros, and must be an unsigned decimal integer from 1 through 18446744073709551615. The Field must be one positional or escaped named selector, without lists, ranges or pairs. Named Fields require -H or --header-in. Selection cannot be mixed with calculations or other Modes.

For example, given these tab-separated Records:

A	8	first observation
B	10	second observation
C	10	third observation
D	9	fourth observation

fastmash top:2 2 emits B then C with all their Fields. fastmash bottom:2 2 emits A then D. Output is in numeric rank order. Equal values keep original input order, and earlier Records win ties at the cutoff. Duplicates remain separate observations; ties never expand the cutoff beyond N. Fewer eligible Records yield fewer output Records, without padding.

Ranking uses the active Numerical profile and Supported locale, including its binary80 precision. Numbers that convert to the same value tie, regardless of their spelling. Both zero signs tie. Infinities sort beyond finite values in their numeric direction. Any unskipped NaN is an Error because it is unordered. Only the ranking Field must be numeric; accompanying Fields are copied without numeric checks.

With --narm, Records whose ranking Field is exactly NA, N/A or NaN, in any letter case, are omitted. Matching does not trim whitespace. Empty or absent Fields, malformed numbers, range failures, signed NaNs and NaN payloads remain Errors. Empty and all-missing datasets emit no data Records. Every accepted Record is checked, including late Records that cannot enter the selection.

Selection copies complete logical Fields in their original order, retaining source numeric spelling, trailing empty Fields, ragged widths and arbitrary accompanying bytes. It adds no score, rank or result column. Ordinary separators, whitespace input, output delimiters, comment filtering and NUL-terminated Records use the existing controls. Whitespace separator runs become the effective output delimiter, and a final unterminated Record receives an output terminator. Ordinary output does not escape embedded output delimiters or Record terminators in copied Fields. Use --csv-out to encode selected text Records as CSV, including sorted text Groups. Use --csv-in to select decoded CSV Records with ordinary output, or --csv for CSV input and output together. These exact-only switches are idempotent and do not enable headers.

CSV input uses the existing strict comma/double-quote grammar, including doubled quotes, multiline Fields, LF or CRLF Record endings, blank Records and a final complete Record without a terminator. It does not trim values, strip a BOM, transcode bytes or normalize decoded newlines. Ranking reads only the decoded numeric Field, while ordinary text keeps its existing conversion across Field boundaries. Every CSV Field is structurally validated, including accompanying Fields on omitted Records and candidates that cannot win.

CSV output encodes each complete logical Field independently, preserving numeric spelling, commas, quotes, CR/LF, NUL and non-UTF-8 bytes. It doubles quotes, quotes empty Fields as "", and emits LF Record endings. Source quote spelling and CRLF framing are not copied. Copied data remains complete even where the separate header-label display convention stops at NUL. --full is accepted with the same output and no deprecation warning. -s without Grouping keys leaves this whole-dataset selection unchanged.

With -g FIELD[,FIELD...], each consecutive Group has its own cutoff. Groups appear in encounter order, and rank order applies within each Group. For example, fastmash -H -g category top:2 reading selects the highest two Records in each adjacent category, binding category and reading from the Input header. A later repeated category remains a separate Group. Selected Records retain every Field once, without an extra Grouping-key prefix. Multiple keys and -i follow the existing Group equality and Supported locale controls; no-key -i leaves numeric ranking unchanged. Empty key text is a key value, while an absent required key is an Error. Every required key and Group boundary is checked before --narm omission, so an all-missing intervening Group still separates equal-key runs on either side. CSV keys compare decoded Field values, so different quote spellings of the same text belong to the same adjacent Group.

Add -s to prepare interleaved categories under the existing sorted Group workflow, for example fastmash -s -H -g category top:2 reading. Groups then follow the existing key order, with ranking applied inside each Group. Sorting keeps complete Records and original input identity through memory and Spill, so earlier source Records still win equal-rank ties. Sorting order and Group equality retain their existing distinct meanings: whitespace key separators, case controls, language-locale spellings and NUL-terminated input can affect the prepared runs without introducing global key normalization. Selection uses the existing native sort for text and decoded CSV sort for quoted CSV. Pipe and regular-file input use the same selection route for each format.

For example, fastmash --csv -s -H -g category top:2 reading selects the highest two Records per category from an interleaved CSV export. Sorting uses decoded key bytes, so quoted and unquoted spellings of the same key sort together. Each selected Record keeps its complete decoded Fields, including multiline notes, through sorting and Spill; CSV output encodes them again. Locale ordering still does not merge differently spelled keys that existing Group equality treats separately. No global key normalization is introduced.

--header-in consumes the first accepted Record as the Input header and excludes it from ranking. Names resolve to the first matching label when a header contains duplicates. --header-out emits one copied-field Output header for the whole Command, without operation wrappers or appended Group labels; -H selects both controls. With an Input header, copied labels follow the existing display convention, stopping at the first NUL byte. Without one, labels are field-1, field-2 and so on, using the first accepted source data Record’s width even when that Record is omitted or does not win. Later ragged Records do not revise the header, and sorting does not change its source schema. Positional ranking Fields may extend beyond the Input header’s labels when the data supplies them. Empty input invents no header; header-only input with both controls and all-missing nonempty input can emit a requested header without data. Name lookup fails before header output.

The entire input must be inspected before ungrouped data is emitted. Completed adjacent Groups can emit as the next Group begins. A read failure or malformed Record does not emit the incomplete current selection, but earlier Groups and the Output header may remain. Output writes and finalization must succeed; a later output failure can leave partial bytes. Sorted grouping finishes input intake before emitting selected data; a read or temporary-sort failure emits no incomplete sorted Group. Later ranking or key Errors can leave already completed Groups. Diagnostics retain original accepted Record numbers, including the Input header and excluding filtered comments. CSV input Errors identify the original logical Record and its starting physical line; syntax Errors also retain their detected location.

At most N candidate Records are retained for the current Group or whole dataset, with storage growing as Records arrive and candidate state released between Groups. A large requested N does not allocate N Records in advance. N limits candidate count, not bytes or total process memory: large Records or a large N can still exhaust memory. Checked resource failures refuse with status 77; Fastmash does not truncate Fields, reduce N or approximate the winners. Sorted grouping has a separate sort-memory target, controlled by FASTMASH_SORT_MEMORY_BYTES, and Spills input when needed under the maintained sort policy. Candidate retention remains bounded by N for the active Group and does not Spill. Neither limit is a guarantee of total process memory.

CSV input conflicts with explicit text-input separators and comment filtering. Any CSV format conflicts with NUL termination and --vnlog; CSV output also conflicts with an explicit Output delimiter. Format conflicts are Errors with status 1, independently of option order. CSV input with ordinary output can use an explicit Output delimiter, subject to ordinary framing limits. A named Field without -H or --header-in is a command Error. Explicit result names, numeric presentation controls, Collapse delimiter, Filler, random Seed, --no-strict, --vnlog and --sort-cmd are unsupported for selection and refuse with status 77. Malformed options and conflicting CSV controls retain Error status 1. Ordinary Commands keep their existing behavior.