You needed the first column of a CSV and typed awk '{print $1}'. That works — and it is often the wrong first move. cut, sort, uniq, and tr exist to shape text in a pipeline before you write a program: slice fields, order keys, count repeats, and translate characters. Reach for awk when the job mixes field math, BEGIN/END, or a condition on a column. Until then, these four tools are enough. Start with a delimiter and a field list, then layer on sort keys, unique counts, and character translation as the stream gets noisier.
Warm-up: a small checkout log
Give yourself a playground that is columnar but not the space-separated access log from the awk article. The file below is comma-separated: user, channel, method, path, status, and response size in bytes.
Create it once, then reuse it for every example in this article:
cat > checkout.csv <<'EOF'
alice,web,GET,/cart,200,128
bob,api,POST,/checkout,201,2048
alice,web,GET,/health,200,64
carol,api,GET,/orders,404,256
bob,web,POST,/login,401,512
alice,api,GET,/orders,200,4096
carol,web,GET,/cart,200,128
bob,api,POST,/checkout,201,2048
dave,web,GET,/about,200,96
EOF
Nine rows, a repeated checkout line, and one user (dave) who appears only once — enough to see slice, sort, and count behave differently.
Slice columns with cut
cut extracts fields. -d sets the delimiter; -f picks one or more fields. Print every username:
cut -d, -f1 checkout.csv
alice
bob
alice
carol
bob
alice
carol
bob
dave
Keep more than one field by listing them. User and path — cut keeps the comma between the fields it prints:
cut -d, -f1,4 checkout.csv
alice,/cart
bob,/checkout
alice,/health
carol,/orders
bob,/login
alice,/orders
carol,/cart
bob,/checkout
dave,/about
A hyphen is a range: cut -d, -f3-5 prints method, path, and status together. When a colleague already filtered with grep, cut can finish the extraction:
grep ',404,' checkout.csv | cut -d, -f1,4
carol,/orders
Note: cut splits on a single delimiter character, not on whitespace runs. The default delimiter is a tab, not a space — echo 'alice web GET' | cut -f1 prints the whole line. Use -d' ' for spaces, and know that two spaces in a row create an empty field (unlike awk, which collapses whitespace). cut has no $NF; the last field on a ragged line is an awk job. -c is the other mode: character positions, not fields — cut -c1-5 on bob,api,… prints bob,a.
Order rows with sort
sort reorders lines. With no flags it compares the whole line as text, so checkout.csv groups by username. Text order is the wrong order for numbers: 10 sorts before 2 because '1' comes before '2'. -n compares numerically:
printf '10\n2\n100\n' | sort
10
100
2
printf '10\n2\n100\n' | sort -n
2
10
100
On a delimited file, -t sets the field separator and -k names the key. Sort by status (field 5) as a number. -k5,5 means “start at field 5, stop at field 5” so later columns do not leak into the key:
sort -t, -k5,5n checkout.csv
alice,api,GET,/orders,200,4096
alice,web,GET,/cart,200,128
alice,web,GET,/health,200,64
carol,web,GET,/cart,200,128
dave,web,GET,/about,200,96
bob,api,POST,/checkout,201,2048
bob,api,POST,/checkout,201,2048
bob,web,POST,/login,401,512
carol,api,GET,/orders,404,256
Add r for reverse and a second key to break ties. Highest status first, then username:
sort -t, -k5,5nr -k1,1 checkout.csv
The 404 and 401 rows rise to the top; the five 200 rows then sort by user (alice, then carol, then dave).
-u keeps unique lines (the first of each duplicate group after sorting). The sample has one repeated checkout row, so sort -u checkout.csv prints eight lines instead of nine.
Note: sort -u is unique whole lines. When you need unique values of one column, slice first (cut -d, -f1 | sort -u). When you also need counts, skip -u and use uniq -c in the next section.
Count uniques with uniq
uniq drops consecutive duplicate lines. Consecutive is the whole contract: unsorted input only collapses neighbors.
Cut usernames and run uniq with no sort:
cut -d, -f1 checkout.csv | uniq
alice
bob
alice
carol
bob
alice
carol
bob
dave
Nothing collapsed — no two identical names sit next to each other. Note: uniq without sort is a silent lie. Always sort first unless you already know the stream is grouped.
The histogram you actually want is sort | uniq -c. Leading spaces are GNU uniq padding the count:
cut -d, -f1 checkout.csv | sort | uniq -c
3 alice
3 bob
2 carol
1 dave
Pipe that into sort -nr to rank by frequency. Equal counts reverse as text, so bob can appear above alice:
cut -d, -f1 checkout.csv | sort | uniq -c | sort -nr
3 bob
3 alice
2 carol
1 dave
-d keeps values that appeared more than once; -u keeps values that appeared exactly once:
cut -d, -f1 checkout.csv | sort | uniq -d
alice
bob
carol
cut -d, -f1 checkout.csv | sort | uniq -u
dave
The same shape works on any column: cut -d, -f4 checkout.csv | sort | uniq -c | sort -nr ranks paths.
Note: sort | uniq and sort -u agree on unique lines. Prefer sort -u when you do not need counts; prefer sort | uniq -c when the number of repeats is the answer.
Translate characters with tr
tr reads bytes from stdin and writes a translated stream. It does not take a filename — pipe into it. Three modes cover almost every daily use: translate, delete, squeeze.
Uppercase every username with character classes (safer than a-z / A-Z if the locale is not C):
cut -d, -f1 checkout.csv | tr '[:lower:]' '[:upper:]'
ALICE
BOB
ALICE
CAROL
BOB
ALICE
CAROL
BOB
DAVE
-d deletes a set. Strip digits from a status-bearing snippet:
echo 'GET /cart 200' | tr -d '0-9'
GET /cart
-s squeezes repeated characters down to one. Messy CSV exports and double-spaced columns are the usual targets:
echo 'alice,,,web,,GET' | tr -s ','
alice,web,GET
printf 'bob api\n' | tr -s ' '
bob api
Note: Squeeze first when delimiters are doubled — then cut field numbers are trustworthy. tr is a character mapper, not a regex engine; substituting a word belongs to sed.
When to hand off to awk or sed
These four tools stop at shape. They do not filter on a numeric field, sum a column, or rewrite a capture group. Hand off when the next step is a real program:
- Field filters,
$NF,BEGIN/END, and arithmetic — awk. The awk article already notes thatcutis often enough; this page is that “enough.” - Substitutions, addresses, in-place edits — sed.
- Search and invert — grep. One
grep | cutpipe is a filter plus a slice, not a grep tutorial. - Paths into commands — find, xargs, and sed together.
If you can name the delimiter and the field list, stay with cut / sort / uniq / tr. If you need “field 5 greater than 400, then add field 6,” open awk.
Quick reference card
Keep this nearby until the flags become muscle memory:
| Goal | Command |
|---|---|
| One field | cut -d, -f1 file |
| Several fields | cut -d, -f1,4 file |
| Field range | cut -d, -f3-5 file |
| Character range | cut -c1-8 file |
| Sort as text | sort file |
| Numeric sort | sort -n file |
| Sort by key | sort -t, -k5,5n file |
| Reverse / unique lines | sort -r file / sort -u file |
| Unique values | cut -d, -f1 file | sort -u |
| Count uniques | … | sort | uniq -c |
| Rank by count | … | sort | uniq -c | sort -nr |
| Duplicates only | sort | uniq -d |
| Singletons only | sort | uniq -u |
| Translate | tr '[:lower:]' '[:upper:]' |
| Delete chars | tr -d ',' |
| Squeeze repeats | tr -s ' ' |
Practice drills
Use the sample checkout.csv (recreate it from the warm-up if you changed it) and try these without peeking. The point is to pick the tool with intent, not to memorize flags under pressure.
- Print only the path column.
- Sort the file by response size (field 6) numerically, smallest first.
- Count requests per user and rank the counts high to low.
- List users who appear more than once, and the user who appears only once — two commands.
- Uppercase every username with
cutandtr.
When you are ready to compare, here are solid answers — not the only ones, but clear and portable:
cut -d, -f4 checkout.csv
sort -t, -k6,6n checkout.csv
cut -d, -f1 checkout.csv | sort | uniq -c | sort -nr
cut -d, -f1 checkout.csv | sort | uniq -d
cut -d, -f1 checkout.csv | sort | uniq -u
cut -d, -f1 checkout.csv | tr '[:lower:]' '[:upper:]'
If you can work through those five comfortably, you already cover most real text-prep work: slice a column, order a key, count what repeats, and clean characters in the pipe. Start with cut -d -f, then add sort, uniq -c, and tr only when the first pass is still messy — and stop before you reinvent awk.