By the end of this lesson, you will be able to build a basic extractive text summarizer using Python and NLTK that identifies key sentences based on word frequency.
What it is
Text summarization reduces a large document into a shorter version while preserving its core meaning. There are two main approaches: extractive (selecting existing sentences) and abstractive (generating new sentences). This lesson focuses on extractive summarization, which is computationally cheaper and easier to implement than abstractive methods using Large Language Models (LLMs). The mental model relies on the assumption that important words appear frequently; therefore, sentences containing these high-frequency words are likely to be central to the topic. Related terms include TF-IDF (Term Frequency-Inverse Document Frequency), stop words, and sentence segmentation.Why it matters
- Efficiency: Quickly condense long reports, news articles, or research papers for human review.
- Cost Reduction: Extractive methods require far less computational power than running LLMs for every summary request.
- Transparency: Since summaries use original text, users can verify facts directly against the source.
- Preprocessing: Useful for cleaning data before feeding it into more complex NLP pipelines.
Syntax or steps
The process involves four distinct steps: 1. Tokenization: Split the text into sentences and words. 2. Cleaning: Remove punctuation and common "stop words" (e.g., "the", "is") that add no semantic value. 3. Scoring: Calculate the frequency of each remaining word. Assign a score to each sentence based on the sum of its word frequencies. 4. Selection: Choose the top-scoring sentences to form the summary.Example
import nltk
from nltk.corpus import stopwords
from nltk.tokenize import sent_tokenize, word_tokenize
# Download necessary resources (run once)
nltk.download('punkt')
nltk.download('stopwords')
def summarize_text(text, num_sentences=3):
# 1. Tokenize sentences and words
sentences = sent_tokenize(text)
words = word_tokenize(text.lower())
# 2. Filter out stop words and non-alphabetic tokens
stop_words = set(stopwords.words('english'))
filtered_words = [w for w in words if w.isalpha() and w not in stop_words]
# 3. Calculate word frequencies
freq_table = {}
for word in filtered_words:
freq_table[word] = freq_table.get(word, 0) + 1
# Normalize frequencies by max count
max_freq = max(freq_table.values()) if freq_table else 1
for word in freq_table:
freq_table[word] /= max_freq
# 4. Score sentences
sentence_scores = {}
for sent in sentences:
sent_words = word_tokenize(sent.lower())
score = sum(freq_table.get(w, 0) for w in sent_words if w.isalpha() and w not in stop_words)
sentence_scores[sent] = score
# Select top sentences
sorted_sents = sorted(sentence_scores.items(), key=lambda x: x[1], reverse=True)
summary = [sent for sent, _ in sorted_sents[:num_sentences]]
return ' '.join(summary)
text = """Natural language processing is a subfield of linguistics, computer science,
and artificial intelligence concerned with the interactions between computers and human language.
In particular, how to program computers to process and analyze large amounts of natural language data.
The goal is a computer capable of understanding the contents of documents, including the contextual nuances
of the language within them."""
print(summarize_text(text))
Explanation: The code first splits the input into sentences. It then creates a frequency table for significant words, ignoring common stop words. Each sentence is scored by summing the normalized frequencies of its words. Finally, the highest-scoring sentences are joined to create the output.
Common mistakes
- Ignoring Stop Words: Including words like "the" or "and" skews frequency counts toward grammatical structure rather than content. Always filter them out.
- Lack of Normalization: If one word appears 100 times and another 50, the first dominates. Dividing by the maximum frequency ensures balanced scoring.
- Punctuation Issues: Failing to remove punctuation from tokens causes "word." and "word" to be treated as different keys in your frequency dictionary.
- Order Preservation: Sorting by score often scrambles the narrative flow. For better readability, re-sort the selected sentences by their original position in the text.
When to use it
| Method | Best For | Limitations |
|---|---|---|
| Extractive (This Lesson) | News, technical docs, quick previews | Can lack cohesion; cannot paraphrase |
| Abstractive (LLMs) | Creative writing, conversational agents | High cost; risk of hallucination |