Python usually gets the spotlight for web scraping, with libraries like BeautifulSoup and Scrapy doing the heavy lifting. Bash can do it too, and this post walks through a script that pulls links and titles from a webpage and writes them to a CSV file.
I spend most of my workday in the terminal writing Bash automation scripts, so as a change of pace I tried web scraping with it. Bash isn’t built for this the way Python is, but it turned out to handle the job well enough that I wanted to share it.
Bash isn’t built for web scraping, but curl, awk, and sed together are enough to pull links and titles from a page.
The Bash Script
#!/bin/bash # Define the URL to scrape base_url = "https://lite.cnn.com" url = "https://lite.cnn.com/" # Create a CSV file and add a header echo "Link,Title" > ; cnn_links.csv # Extract links and titles and save them to the CSV file link_array =( $( curl -s "$url" | awk -F ' ; href = "' '/<a/{gsub(/" .*/, "" , $2 ) ; print $2 }' ; ) ) for link in "${link_array[@]}" ; do full_link = "${base_url}${link}" title = $( curl -s "$full_link" | grep -o ' ; < title[^> ; ]*> ; [^ < ]* < /title> ; ' | sed -e ' s/ < title> ; //g ' -e ' s/ < \/title> ; //g ' ; ) echo "\" $full_link \ ",\" $title \ "" >> cnn_links.csv done echo "Scraping and CSV creation complete. Links and titles saved to 'cnn_links.csv'." How it works
This Bash script does the following:
- It defines the base URL and the URL of the webpage you want to scrape.
- It creates a CSV file named
cnn_links.csvwith a header row containing “Link” and “Title” columns. - Using
curl, it fetches the HTML content of the specified webpage and extracts all the links found within anchor tags(<a>)usingawk. - It then iterates through the array of links and extracts the page titles by making additional
curlrequests to each link. - Finally, it appends the extracted links and titles to the CSV file in the desired format.
Breaking it down further
grep -o '<title[^>]*>[^<]*</title>'extracts the page title from the HTML content using regular expressions:-otells grep to output only the matched part of the input text.<title[^>]*>matches the opening<title>tag and any attributes (e.g.,<title attribute="value">), if present.[^<]*matches any characters that aren’t<(the text inside the<title>tag).</title>matches the closing</title>tag.
sed -e 's/<title>//g' -e 's/<\/title>//g'removes the<title>and</title>tags from the extracted title:-elets you specify multiple commands forsedto run.'s/<title>//g'replaces every<title>with an empty string (removing the opening tag).'s/<\/title>//g'replaces every</title>with an empty string (removing the closing tag).
Combining these commands:
grepextracts the text within the<title>and</title>tags.sedthen removes the tags themselves, leaving only the text content of the title.
This command also uses awk to extract URLs from an HTML document. Let’s break it down step by step:
awk -F 'href="':awkis a text processing tool that operates on text files or input streams.-F 'href="'sets the field separator to'href="'. This meansawkwill treat'href="'as the delimiter for splitting input lines into fields.
'/<a/{gsub(/".*/, "", $2); print $2}':/<a/is a pattern that specifies a condition: lines containing<a>. This ensures that the following actions are only applied to lines containing anchor tags.gsub(/".*/, "", $2)is anawkfunction that globally substitutes (gsub) everything from the first double quote (") to the end of the field ($2) with an empty string. In this case, it effectively removes the opening", and the result is the URL.print $2prints the modified field (the extracted URL).
So, this awk command looks for lines containing anchor tags (<a>) and extracts the URLs by removing everything before the first double quote (") in the href attribute. The extracted URLs are then printed as output.
Wrapping up
This is a simple web scraper built from awk, sed, grep, and curl. Since Bash ships on most Linux systems, it can scrape web pages without installing anything extra. It’s not as capable as Python for this, but it handles simple tasks like pulling links and titles from a page well. I wouldn’t reach for it on anything more complex.
This was a fun script to build while learning more Bash. I’d still call myself a beginner at it.
If you have bash tips, or have used bash for something similar, I’d love to hear about it in the comments.
If you liked this post, you can support my work by buying me a coffee. It would mean a lot. You can also follow me on Twitter, and tag me @muhammad_o7 if you share this post there.