Natural Language Processing (NLP) is the branch of computer science that lets computers work with human language. It powers practical things like chatbots, spam filtering, spell check, and better search engines.
Stemming is fast but crude; lemmatization is slower but returns real words.
This guide covers the core NLP concepts worth knowing:
- Tokenization: breaking text into meaningful units
- Stemming: reducing words to their root form
- Lemmatization: vocabulary-aware word normalization
- POS tags: parts of speech tagging
- Named entity recognition: identifying entities in text
- Chunking: extracting meaningful phrases
Tokenization
Tokenization is the first step in the NLP process: splitting text into minimal meaningful units a machine can work with. Further reading
Stemming (Stemmation)
Stemming (sometimes called stemmation) is the process of reducing words to their root form. For example, words like Plays, Played, and Playing all reduce to the root word Play.
Stemming works by stripping prefixes and suffixes from words. The most common algorithm is the Porter Stemmer, which uses a series of rules to systematically remove word endings.
from nltk.stem import PorterStemmer
stemmer = PorterStemmer()
words = ["playing", "played", "plays", "player"]
for word in words:
print(f"{word} -> {stemmer.stem(word)}")
# Output: playing -> play, played -> play, plays -> play, player -> playerFurther reading:
Lemmatization
Lemmatization is a more sophisticated technique compared to stemming. While stemming simply chops off word endings, lemmatization uses vocabulary and morphological analysis to return the base dictionary form of a word (called a lemma).
The key difference: stemming might reduce “better” to “bett”, but lemmatization correctly returns “good”.
from nltk.stem import WordNetLemmatizer
lemmatizer = WordNetLemmatizer()
# Lemmatization considers parts of speech
print(lemmatizer.lemmatize("running", pos="v")) # -> run
print(lemmatizer.lemmatize("better", pos="a")) # -> good
print(lemmatizer.lemmatize("geese")) # -> gooseStemming vs. lemmatization:
- Stemming: fast, rule-based, may produce non-words (
studies→studi) - Lemmatization: slower, dictionary-based, always produces real words (
studies→study)
Further reading:
POS tagging (parts of speech)
POS tagging assigns grammatical categories (noun, verb, adjective, and so on) to each word in a sentence. It’s essential for understanding sentence structure and meaning.
import nltk
sentence = "The quick brown fox jumps over the lazy dog"
tokens = nltk.word_tokenize(sentence)
pos_tags = nltk.pos_tag(tokens)
print(pos_tags)
# [('The', 'DT'), ('quick', 'JJ'), ('brown', 'JJ'), ('fox', 'NN'), ...]Common POS tags: NN (noun), VB (verb), JJ (adjective), RB (adverb), DT (determiner).
Named entity recognition (NER)
Named Entity Recognition identifies and classifies named entities in text into predefined categories like person names, organizations, locations, dates, and more.
import nltk
sentence = "Mark Zuckerberg is the CEO of Facebook in California"
tokens = nltk.word_tokenize(sentence)
pos_tags = nltk.pos_tag(tokens)
entities = nltk.ne_chunk(pos_tags)
# Identifies: Mark Zuckerberg (PERSON), Facebook (ORGANIZATION), California (GPE)Chunking
Chunking extracts meaningful phrases from text rather than individual words. For example, “South Africa” should be treated as a single entity, not two separate words.
# Chunking groups related words together
# "New York City" -> single location chunk
# "machine learning" -> single noun phraseChunking is useful for:
- Extracting noun phrases for keyword analysis
- Information extraction from documents
- Building knowledge graphs
Further reading:
Putting it together
Tokenization, stemming, lemmatization, POS tagging, NER, and chunking are the foundation for building text processing applications. Whether you’re building a chatbot, a search engine, or a sentiment analyzer, these concepts are essential.
To get started with NLP in Python, install NLTK:
pip install nltk python -c "import nltk; nltk.download('punkt'); nltk.download('averaged_perceptron_tagger'); nltk.download('wordnet')"