Back to Data Science Notes
Topic #89

Building a Text Summarizer

By the end of this lesson, you will be able to build a basic extractive text summarizer using Python and NLTK that identifies key sentences based on word frequency.

What it is

Text summarization reduces a large document into a shorter version while preserving its core meaning. There are two main approaches: extractive (selecting existing sentences) and abstractive (generating new sentences). This lesson focuses on extractive summarization, which is computationally cheaper and easier to implement than abstractive methods using Large Language Models (LLMs). The mental model relies on the assumption that important words appear frequently; therefore, sentences containing these high-frequency words are likely to be central to the topic. Related terms include TF-IDF (Term Frequency-Inverse Document Frequency), stop words, and sentence segmentation.

Why it matters

  • Efficiency: Quickly condense long reports, news articles, or research papers for human review.
  • Cost Reduction: Extractive methods require far less computational power than running LLMs for every summary request.
  • Transparency: Since summaries use original text, users can verify facts directly against the source.
  • Preprocessing: Useful for cleaning data before feeding it into more complex NLP pipelines.

Syntax or steps

The process involves four distinct steps: 1. Tokenization: Split the text into sentences and words. 2. Cleaning: Remove punctuation and common "stop words" (e.g., "the", "is") that add no semantic value. 3. Scoring: Calculate the frequency of each remaining word. Assign a score to each sentence based on the sum of its word frequencies. 4. Selection: Choose the top-scoring sentences to form the summary.

Example

import nltk
from nltk.corpus import stopwords
from nltk.tokenize import sent_tokenize, word_tokenize

# Download necessary resources (run once)
nltk.download('punkt')
nltk.download('stopwords')

def summarize_text(text, num_sentences=3):
    # 1. Tokenize sentences and words
    sentences = sent_tokenize(text)
    words = word_tokenize(text.lower())
    
    # 2. Filter out stop words and non-alphabetic tokens
    stop_words = set(stopwords.words('english'))
    filtered_words = [w for w in words if w.isalpha() and w not in stop_words]
    
    # 3. Calculate word frequencies
    freq_table = {}
    for word in filtered_words:
        freq_table[word] = freq_table.get(word, 0) + 1
        
    # Normalize frequencies by max count
    max_freq = max(freq_table.values()) if freq_table else 1
    for word in freq_table:
        freq_table[word] /= max_freq
        
    # 4. Score sentences
    sentence_scores = {}
    for sent in sentences:
        sent_words = word_tokenize(sent.lower())
        score = sum(freq_table.get(w, 0) for w in sent_words if w.isalpha() and w not in stop_words)
        sentence_scores[sent] = score
        
    # Select top sentences
    sorted_sents = sorted(sentence_scores.items(), key=lambda x: x[1], reverse=True)
    summary = [sent for sent, _ in sorted_sents[:num_sentences]]
    
    return ' '.join(summary)

text = """Natural language processing is a subfield of linguistics, computer science, 
and artificial intelligence concerned with the interactions between computers and human language. 
In particular, how to program computers to process and analyze large amounts of natural language data. 
The goal is a computer capable of understanding the contents of documents, including the contextual nuances 
of the language within them."""

print(summarize_text(text))
Explanation: The code first splits the input into sentences. It then creates a frequency table for significant words, ignoring common stop words. Each sentence is scored by summing the normalized frequencies of its words. Finally, the highest-scoring sentences are joined to create the output.

Common mistakes

  • Ignoring Stop Words: Including words like "the" or "and" skews frequency counts toward grammatical structure rather than content. Always filter them out.
  • Lack of Normalization: If one word appears 100 times and another 50, the first dominates. Dividing by the maximum frequency ensures balanced scoring.
  • Punctuation Issues: Failing to remove punctuation from tokens causes "word." and "word" to be treated as different keys in your frequency dictionary.
  • Order Preservation: Sorting by score often scrambles the narrative flow. For better readability, re-sort the selected sentences by their original position in the text.

When to use it

MethodBest ForLimitations
Extractive (This Lesson)News, technical docs, quick previewsCan lack cohesion; cannot paraphrase
Abstractive (LLMs)Creative writing, conversational agentsHigh cost; risk of hallucination
Use extractive methods when speed and factual accuracy are paramount. Use abstractive methods when fluency and synthesis are required.

Practice

Guided Exercise: Modify the `summarize_text` function to accept a parameter `min_length`. Exclude any sentence shorter than this character count from being selected as a summary line. Challenge: Implement a simple penalty system where sentences that share too many words with already selected sentences receive a lower score (to reduce redundancy). Hint: Compare the set of words in the current candidate against the union of words in previously selected sentences.

Quick check

Question: Why do we divide word frequencies by the maximum frequency? Answer: To normalize the scores so that no single extremely frequent word disproportionately dominates the sentence ranking compared to other relevant terms.

Summary

Extractive summarization offers a lightweight, transparent way to condense text by leveraging word frequency statistics. While it lacks the generative flair of LLMs, it remains an essential tool for efficient data preprocessing and rapid information retrieval in data science workflows.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Building a Text Summarizer – FAQs

Quick answers about learning Building a Text Summarizer in Data Science.

This free note from Coding Now Tech Institute explains Building a Text Summarizer in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Building a Text Summarizer, is 100% free with no signup required.
With focused practice, most students grasp Building a Text Summarizer in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now