By the end of this lesson, you will understand how Naive Bayes classifiers use probability theory to categorize data quickly and efficiently, even with limited training examples.
What it is
Naive Bayes is a family of probabilistic machine learning algorithms based on applying Bayes' theorem with strong (naive) independence assumptions between the features. The core mental model is simple: instead of calculating the complex joint probability of all features occurring together, the algorithm assumes each feature contributes independently to the probability of a class. This makes calculations extremely fast and scalable. Key related terms includePrior Probability (initial belief before seeing data), Likelihood (probability of evidence given a hypothesis), and Posterior Probability (updated belief after seeing evidence).
Why it matters
- Speed: It trains and predicts almost instantly, making it ideal for real-time systems.
- Small Data Performance: It works surprisingly well with small datasets where other models might overfit or fail to converge.
- High Dimensionality: It handles text classification tasks with thousands of features (words) effectively.
- Simplicity: The mathematical foundation is straightforward, making it easy to interpret and debug.
Syntax or steps
The algorithm follows three main steps: 1. Calculate the prior probability of each class from the training data. 2. For each feature, calculate the likelihood of that feature appearing in each class. 3. Apply Bayes' theorem to compute the posterior probability for each class given the input features, then select the class with the highest probability.Example
Here is a minimal Python example using thescikit-learn library to classify emails as spam or not spam based on word counts.
from sklearn.naive_bayes import MultinomialNB
from sklearn.feature_extraction.text import CountVectorizer
# Sample data: 4 emails and their labels
emails = [
"free money now",
"meeting tomorrow at 10",
"win cash prize",
"project deadline update"
]
labels = ["spam", "ham", "spam", "ham"]
# Step 1: Convert text to numerical features (word counts)
vectorizer = CountVectorizer()
X = vectorizer.fit_transform(emails)
# Step 2: Initialize and train the Naive Bayes classifier
clf = MultinomialNB()
clf.fit(X, labels)
# Step 3: Predict a new email
new_email = ["free meeting tomorrow"]
new_features = vectorizer.transform(new_email)
prediction = clf.predict(new_features)
print(f"Prediction: {prediction[0]}")
Explanation:
CountVectorizer converts raw text into a matrix of token counts. MultinomialNB is chosen because it is suitable for discrete counts (like word frequencies). The fit method learns the probabilities, and predict applies them to new data.
Common mistakes
- Ignoring Zero Probabilities: If a word never appears in a specific class during training, its likelihood becomes zero, crashing the calculation. Use Laplace smoothing (built into most implementations) to fix this.
- Assuming True Independence: In reality, words like "New" and "York" are dependent. While the "naive" assumption simplifies math, be aware it can reduce accuracy in highly correlated feature sets.
- Using Wrong Variant: Do not use Gaussian NB for text data; use Multinomial or Bernoulli NB. Conversely, do not use Multinomial for continuous sensor data; use Gaussian NB.
- Feature Scaling: Unlike many ML models, Naive Bayes does not require feature scaling (normalization) because it relies on relative probabilities, not distances.
When to use it
Compare Naive Bayes with Logistic Regression, another common linear classifier.| Criterion | Naive Bayes | Logistic Regression |
|---|---|---|
| Training Speed | Extremely Fast | Slower (iterative optimization) |
| Data Size | Good for small data | Better for large data |
| Interpretability | High (probabilistic) | Medium (coefficients) |
| Accuracy | Often lower if features correlate | Often higher if features correlate |
Practice
Guided Exercise: Modify the example above to predict the label for the phrase "cash prize deadline". What do you expect the output to be?Challenge: Add a fifth email "urgent free project" with label "spam". Retrain the model and predict "urgent cash". Does the prediction change? Why?
Hint: Look at which words appear in both classes and how the priors shift.