bash

WebScraping in Bash

A Bash script that scrapes links and titles from a webpage into a CSV file, using curl, awk, and sed.

Python usually gets the spotlight for web scraping, with libraries like BeautifulSoup and Scrapy doing the heavy lifting. Bash can do it too, and this post walks through a script that pulls links and titles from a webpage and writes them to a CSV file.

I spend most of my workday in the terminal writing Bash automation scripts, so as a change of pace I tried web scraping with it. Bash isn’t built for this the way Python is, but it turned out to handle the job well enough that I wanted to share it.

Bash isn’t built for web scraping, but curl, awk, and sed together are enough to pull links and titles from a page.

The Bash Script

bash
#!/bin/bash# Define the URL to scrapebase_url="https://lite.cnn.com"url="https://lite.cnn.com/"# Create a CSV file and add a headerecho "Link,Title" >; cnn_links.csv# Extract links and titles and save them to the CSV filelink_array=($(curl -s "$url" | awk -F ';href="' '/<a/{gsub(/".*/, "", $2); print $2}';))for link in "${link_array[@]}"; do    full_link="${base_url}${link}"    title=$(curl -s "$full_link" | grep -o ';<title[^>;]*>;[^<]*</title>;' | sed -e 's/<title>;//g' -e 's/<\/title>;//g';)    echo "\"$full_link\",\"$title\"" >> cnn_links.csvdoneecho "Scraping and CSV creation complete. Links and titles saved to 'cnn_links.csv'."

How it works

This Bash script does the following:

  1. It defines the base URL and the URL of the webpage you want to scrape.
  2. It creates a CSV file named cnn_links.csv with a header row containing “Link” and “Title” columns.
  3. Using curl, it fetches the HTML content of the specified webpage and extracts all the links found within anchor tags (<a>) using awk.
  4. It then iterates through the array of links and extracts the page titles by making additional curl requests to each link.
  5. Finally, it appends the extracted links and titles to the CSV file in the desired format.

Breaking it down further

  1. grep -o '<title[^>]*>[^<]*</title>' extracts the page title from the HTML content using regular expressions:
    1. -o tells grep to output only the matched part of the input text.
    2. <title[^>]*> matches the opening <title> tag and any attributes (e.g., <title attribute="value">), if present.
    3. [^<]* matches any characters that aren’t < (the text inside the <title> tag).
    4. </title> matches the closing </title> tag.
  2. sed -e 's/<title>//g' -e 's/<\/title>//g' removes the <title> and </title> tags from the extracted title:
    1. -e lets you specify multiple commands for sed to run.
    2. 's/<title>//g' replaces every <title> with an empty string (removing the opening tag).
    3. 's/<\/title>//g' replaces every </title> with an empty string (removing the closing tag).

Combining these commands:

  1. grep extracts the text within the <title> and </title> tags.
  2. sed then removes the tags themselves, leaving only the text content of the title.

This command also uses awk to extract URLs from an HTML document. Let’s break it down step by step:

  1. awk -F 'href="':

    • awk is a text processing tool that operates on text files or input streams.
    • -F 'href="' sets the field separator to 'href="'. This means awk will treat 'href="' as the delimiter for splitting input lines into fields.
  2. '/<a/{gsub(/".*/, "", $2); print $2}':

    • /<a/ is a pattern that specifies a condition: lines containing <a>. This ensures that the following actions are only applied to lines containing anchor tags.
    • gsub(/".*/, "", $2) is an awk function that globally substitutes (gsub) everything from the first double quote (") to the end of the field ($2) with an empty string. In this case, it effectively removes the opening ", and the result is the URL.
    • print $2 prints the modified field (the extracted URL).

So, this awk command looks for lines containing anchor tags (<a>) and extracts the URLs by removing everything before the first double quote (") in the href attribute. The extracted URLs are then printed as output.

Wrapping up

This is a simple web scraper built from awk, sed, grep, and curl. Since Bash ships on most Linux systems, it can scrape web pages without installing anything extra. It’s not as capable as Python for this, but it handles simple tasks like pulling links and titles from a page well. I wouldn’t reach for it on anything more complex.

This was a fun script to build while learning more Bash. I’d still call myself a beginner at it.

If you have bash tips, or have used bash for something similar, I’d love to hear about it in the comments.


If you liked this post, you can support my work by buying me a coffee. It would mean a lot. You can also follow me on Twitter, and tag me @muhammad_o7 if you share this post there.