By the end of this lesson, you will understand how to use Git and GitHub to track changes in your data science code, enabling collaboration and ensuring your analysis is reproducible.
What it is
Git is a distributed version control system that records snapshots of your project files over time. GitHub is a web-based hosting service for Git repositories. In data science, this combination allows you to save different versions of your scripts, notebooks, and documentation without manually creating files like analysis_v1.py, analysis_final.py, or analysis_FINAL_real.py. Key terms include repository (the project folder), commit (a saved snapshot with a message), and branch (an experimental copy of the code).
Why it matters
- Reproducibility: You can revert to a previous state if new code breaks existing results.
- Collaboration: Multiple team members can work on the same project simultaneously without overwriting each other’s work.
- Documentation: Commit messages serve as a chronological log of why changes were made.
- Backup: Your code is stored remotely on GitHub, protecting against local hardware failure.
Syntax or steps
The basic workflow involves three stages: staging changes, committing them locally, and pushing them to GitHub.
git add <filename>: Moves modified files to the staging area.git commit -m "message": Saves the staged changes to the local history.git push origin main: Uploads the local commits to the remote GitHub repository.
Example
# Initialize a new repository in the current directory
git init
# Create a simple Python script
echo "print('Hello Data Science')" > hello_ds.py
# Stage the file for tracking
git add hello_ds.py
# Commit the change with a descriptive message
git commit -m "Initial commit: added hello_ds.py"
# Link local repo to a remote GitHub repository (replace URL)
git remote add origin https://github.com/username/repo.git
# Push the commit to GitHub
git push -u origin main
This sequence creates a tracked file, saves its state, and uploads it to the cloud. The -u flag sets the upstream branch, so future pushes only require git push.
Common mistakes
- Committing large datasets: Never commit raw CSVs or model weights directly. Use
.gitignoreto exclude these files to keep the repository lightweight. - Vague commit messages: Avoid messages like "update" or "fix." Instead, write "Fixed null handling in cleaning step."
- Working directly on main: For complex projects, create a new branch (
git checkout -b feature-name) to isolate experiments before merging. - Ignoring conflicts: If two people edit the same line, Git pauses. Resolve conflicts manually in the editor before committing again.
When to use it
| Scenario | Use Git/GitHub | Alternative |
|---|---|---|
| Team collaboration on code | Yes | Emailing zip files |
| Tracking notebook iterations | Yes (with caution) | Manual file naming |
| Storing sensitive credentials | No | Environment variables / Secrets manager |
| Large binary data storage | No | DVC (Data Version Control) or Cloud Storage |
Practice
Guided Exercise: Create a text file named notes.txt containing "Step 1: Load data." Run git add notes.txt and git commit -m "Add initial notes". Then modify the file to say "Step 2: Clean data," stage it, and commit again. Check your history using git log --oneline.
Challenge: Try to undo the last commit without losing your changes. Hint: Look up git reset --soft HEAD~1.
Quick check
Question: Why should you avoid committing Jupyter Notebook outputs directly to Git?
Answer: Outputs often contain random seeds, timestamps, or large plots that cause frequent merge conflicts and bloat the repository size. It is better to clear outputs before committing or use tools like jupytext to convert notebooks to plain text scripts.
Summary
Git and GitHub provide essential infrastructure for professional data science by managing code history and facilitating teamwork. Mastering the basic add-commit-push cycle ensures your work is safe, traceable, and shareable.