python

NLP Basics: Stemming, Lemmatization, Tokenization & More

Learn the fundamentals of Natural Language Processing including stemming (stemmation), lemmatization, tokenization, POS tagging, named entity recognition, and chunking with practical examples.

Natural Language Processing (NLP) is the branch of computer science that lets computers work with human language. It powers practical things like chatbots, spam filtering, spell check, and better search engines.

Stemming is fast but crude; lemmatization is slower but returns real words.

This guide covers the core NLP concepts worth knowing:

  • Tokenization: breaking text into meaningful units
  • Stemming: reducing words to their root form
  • Lemmatization: vocabulary-aware word normalization
  • POS tags: parts of speech tagging
  • Named entity recognition: identifying entities in text
  • Chunking: extracting meaningful phrases

Tokenization

Tokenization is the first step in the NLP process: splitting text into minimal meaningful units a machine can work with. Further reading

Stemming (Stemmation)

Stemming (sometimes called stemmation) is the process of reducing words to their root form. For example, words like Plays, Played, and Playing all reduce to the root word Play.

Stemming works by stripping prefixes and suffixes from words. The most common algorithm is the Porter Stemmer, which uses a series of rules to systematically remove word endings.

python
from nltk.stem import PorterStemmer
stemmer = PorterStemmer()

words = ["playing", "played", "plays", "player"]
for word in words:
    print(f"{word} -> {stemmer.stem(word)}")
# Output: playing -> play, played -> play, plays -> play, player -> player

Further reading:

Lemmatization

Lemmatization is a more sophisticated technique compared to stemming. While stemming simply chops off word endings, lemmatization uses vocabulary and morphological analysis to return the base dictionary form of a word (called a lemma).

The key difference: stemming might reduce “better” to “bett”, but lemmatization correctly returns “good”.

python
from nltk.stem import WordNetLemmatizer
lemmatizer = WordNetLemmatizer()

# Lemmatization considers parts of speech
print(lemmatizer.lemmatize("running", pos="v"))  # -> run
print(lemmatizer.lemmatize("better", pos="a"))   # -> good
print(lemmatizer.lemmatize("geese"))             # -> goose

Stemming vs. lemmatization:

  • Stemming: fast, rule-based, may produce non-words (studies → studi)
  • Lemmatization: slower, dictionary-based, always produces real words (studies → study)

Further reading:

POS tagging (parts of speech)

POS tagging assigns grammatical categories (noun, verb, adjective, and so on) to each word in a sentence. It’s essential for understanding sentence structure and meaning.

python
import nltk
sentence = "The quick brown fox jumps over the lazy dog"
tokens = nltk.word_tokenize(sentence)
pos_tags = nltk.pos_tag(tokens)
print(pos_tags)
# [('The', 'DT'), ('quick', 'JJ'), ('brown', 'JJ'), ('fox', 'NN'), ...]

Common POS tags: NN (noun), VB (verb), JJ (adjective), RB (adverb), DT (determiner).

Named entity recognition (NER)

Named Entity Recognition identifies and classifies named entities in text into predefined categories like person names, organizations, locations, dates, and more.

python
import nltk
sentence = "Mark Zuckerberg is the CEO of Facebook in California"
tokens = nltk.word_tokenize(sentence)
pos_tags = nltk.pos_tag(tokens)
entities = nltk.ne_chunk(pos_tags)
# Identifies: Mark Zuckerberg (PERSON), Facebook (ORGANIZATION), California (GPE)

Chunking

Chunking extracts meaningful phrases from text rather than individual words. For example, “South Africa” should be treated as a single entity, not two separate words.

python
# Chunking groups related words together
# "New York City" -> single location chunk
# "machine learning" -> single noun phrase

Chunking is useful for:

  • Extracting noun phrases for keyword analysis
  • Information extraction from documents
  • Building knowledge graphs

Further reading:


Putting it together

Tokenization, stemming, lemmatization, POS tagging, NER, and chunking are the foundation for building text processing applications. Whether you’re building a chatbot, a search engine, or a sentiment analyzer, these concepts are essential.

To get started with NLP in Python, install NLTK:

bash
pip install nltkpython -c "import nltk; nltk.download('punkt'); nltk.download('averaged_perceptron_tagger'); nltk.download('wordnet')"