Back to Data Science Notes
Topic #56

Probability & Bayes' Theorem

By the end of this lesson, you will be able to calculate conditional probabilities and apply Bayes' Theorem to update beliefs based on new evidence.

What it is

Probability quantifies uncertainty. Conditional probability, denoted as $P(A|B)$, measures the likelihood of event A occurring given that event B has already occurred. Bayes' Theorem provides a mathematical framework for updating the probability of a hypothesis as more evidence or information becomes available. It connects prior beliefs with observed data to produce posterior beliefs.

Key terms include:

  • Prior ($P(H)$): Initial belief about a hypothesis before seeing data.
  • Likelihood ($P(E|H)$): Probability of observing evidence if the hypothesis is true.
  • Evidence ($P(E)$): Total probability of observing the evidence under all hypotheses.
  • Posterior ($P(H|E)$): Updated belief after considering the evidence.

Why it matters

  • Spam Filtering: Updates the probability an email is spam based on specific words appearing in the text.
  • Medical Diagnosis: Refines the probability of a disease given a positive test result, accounting for false positives.
  • Machine Learning: Forms the basis of Naive Bayes classifiers used in text classification and recommendation systems.
  • Risk Assessment: Helps financial models adjust risk probabilities as market conditions change.

Syntax or steps

The core formula for Bayes' Theorem is:

P(H|E) = (P(E|H) * P(H)) / P(E)

To solve a problem using this theorem:

  1. Identify the Hypothesis ($H$) and Evidence ($E$).
  2. Determine the Prior probability $P(H)$.
  3. Determine the Likelihood $P(E|H)$.
  4. Calculate the total probability of Evidence $P(E)$ using the Law of Total Probability: $P(E) = P(E|H)P(H) + P(E|\neg H)P(\neg H)$.
  5. Plug values into the formula to find the Posterior $P(H|E)$.

Example

Consider a rare disease affecting 1% of the population ($P(Disease) = 0.01$). A test detects the disease correctly 99% of the time ($P(Pos|Disease) = 0.99$) but gives a false positive 5% of the time ($P(Pos|No Disease) = 0.05$). If a patient tests positive, what is the probability they actually have the disease?

# Define probabilities
prior_disease = 0.01
prior_no_disease = 1 - prior_disease

likelihood_pos_given_disease = 0.99
likelihood_pos_given_no_disease = 0.05

# Calculate total probability of testing positive (Evidence)
total_prob_positive = (likelihood_pos_given_disease * prior_disease) + \
                      (likelihood_pos_given_no_disease * prior_no_disease)

# Apply Bayes' Theorem
posterior_disease_given_pos = (likelihood_pos_given_disease * prior_disease) / total_prob_positive

print(f"Probability of disease given positive test: {posterior_disease_given_pos:.4f}")

Explanation: Despite a highly accurate test, the low prevalence (prior) means many positives are false alarms. The calculation shows the actual probability is roughly 17%, not 99%. This highlights the importance of priors.

Common mistakes

  • Ignoring the Base Rate: Assuming $P(H|E) \approx P(E|H)$. Always account for how common the hypothesis is initially.
  • Misinterpreting False Positives: Forgetting that even small error rates can dominate when the condition is rare.
  • Confusing Independence: Applying Bayes' theorem incorrectly when events are not independent without adjusting the model.
  • Division by Zero: Failing to check if $P(E)$ is zero, which makes the posterior undefined.

When to use it

Compare Bayesian updating with frequentist hypothesis testing.

FeatureBayesian ApproachFrequentist Approach
PhilosophyUpdates belief degreesTests against null hypothesis
Data UsageIncorporates prior knowledgeRelies solely on current data
OutputPosterior distributionp-value / Confidence Interval
Best ForSmall data, sequential updatesLarge samples, fixed designs

Practice

Guided Exercise: Modify the code above to calculate the probability of having the disease given a negative test result. Hint: Use $P(Neg|Disease) = 1 - 0.99$ and $P(Neg|No Disease) = 1 - 0.05$.

Challenge: Write a function `bayes_update(prior, likelihood_true, likelihood_false)` that returns the posterior probability given a positive observation. Test it with the disease example.

Quick check

Question: Why might a medical test with 99% accuracy still yield a low probability of disease after a positive result?

Answer: Because the disease is rare (low prior), so the number of healthy people receiving false positives outweighs the few sick people receiving true positives.

Summary

Bayes' Theorem allows us to rationally update our beliefs by combining prior knowledge with new evidence. It emphasizes that context (base rates) is crucial for interpreting statistical results accurately.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Probability & Bayes' Theorem – FAQs

Quick answers about learning Probability & Bayes' Theorem in Data Science.

This free note from Coding Now Tech Institute explains Probability & Bayes' Theorem in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Probability & Bayes' Theorem, is 100% free with no signup required.
With focused practice, most students grasp Probability & Bayes' Theorem in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now