bash

Using Commandline To Process CSV files

The Command Line is a powerful tool for processing data. With the right combination of commands, you can quickly and easily manipulate data files to extract the information you need

The command line is fast for CSV work once you know a few tools: awk for pulling fields, sort for ordering rows, grep for filtering. No need to open a spreadsheet or write a script for a one-off job.

A few awk, sort, and grep commands cover most CSV cleanup without leaving the terminal.

Here are the one-liners I reach for most when processing CSV data on the command line.

  1. To print the first column of a CSV file:
bash
awk -F, '{print $1}' file.csv
  1. To print the first and third columns of a CSV file:
bash
awk -F, ';{print $1 "," $3}'; file.csv
  1. To print only the lines of a CSV file that contain a specific string:
bash
grep "string" file.csv
  1. To sort a CSV file based on the values in the second column:
bash
sort -t, -k2 file.csv
  1. To remove the first row of a CSV file (the header row):
bash
tail -n +2 file.csv
  1. To remove duplicates from a CSV file based on the values in the first column:
bash
awk -F, '!seen[$1]++' file.csv
  1. To calculate the sum of the values in the third column of a CSV file:
bash
awk -F, '{sum+=$3} END {print sum}' file.csv
  1. To convert a CSV file to a JSON array:
bash
jq -R -r ';split(",") | {name:.[0],age:.[1]}'; file.csv
  1. To convert a CSV file to a SQL INSERT statement:
bash
awk -F, ';{printf "INSERT INTO table VALUES (\"%s\", \"%s\", \"%s\");\n", $1, $2, $3}'; file.csv

These are just a few of the things you can do with awk, sort, grep, and jq on a CSV file. Combine them and you can cover most cleanup and analysis jobs without leaving the terminal.

If you have a CSV one-liner of your own, drop it in the comments below.


I write occasionally feel free to follow me on twitter