By the end of this lesson, you will be able to create and customize line plots in Python using Matplotlib to visualize trends in continuous data.
What it is
A line plot connects individual data points with straight lines, making it ideal for showing how a variable changes over time or across an ordered sequence. The mental model is simple: each point represents a measurement at a specific position on the x-axis, and the line illustrates the trajectory between those measurements. Key related terms include x-axis (independent variable), y-axis (dependent variable), and marker (optional symbols at data points).
Why it matters
- Trend Identification: Quickly reveals upward, downward, or cyclical patterns in time-series data like stock prices or temperature readings.
- Comparison: Multiple lines can be overlaid to compare different datasets, such as sales figures for two products.
- Continuity: Unlike bar charts, line plots emphasize the flow and connection between discrete observations.
- Simplicity: Requires minimal configuration to produce clear, publication-ready visuals.
Syntax or steps
The core function is plt.plot(). You must import matplotlib.pyplot, prepare your data arrays, call the plotting function, add labels, and display the result. Always ensure your x and y data have the same length.
Example
import matplotlib.pyplot as plt
# Sample data: months and corresponding sales
months = ['Jan', 'Feb', 'Mar', 'Apr', 'May']
sales = [150, 200, 180, 250, 300]
# Create the line plot
plt.figure(figsize=(8, 5))
plt.plot(months, sales, marker='o', linestyle='-', color='blue')
# Add titles and labels
plt.title("Monthly Sales Trend")
plt.xlabel("Month")
plt.ylabel("Units Sold")
# Display the plot
plt.show()
Explanation: First, we define lists for months and sales. We initialize a figure with plt.figure() to control size. The plt.plot() command draws the line; marker='o' adds circles at each data point, while linestyle='-' ensures a solid line. Finally, plt.title() and axis labels provide context, and plt.show() renders the window.
Common mistakes
- Mismatched Data Lengths: If
xandyarrays differ in size, Matplotlib raises an error. Always verify lengths before plotting. - Forgetting
plt.show(): In scripts, omitting this prevents the plot from appearing. Note that in Jupyter notebooks, it may auto-display, but explicit calling is safer. - Overcrowding Axes: Plotting too many lines without distinct colors or legends makes the chart unreadable. Use
plt.legend()when comparing multiple series. - Ignoring Scale: Line plots can exaggerate small fluctuations if the y-axis range is too narrow. Adjust limits with
plt.ylim()if necessary.
When to use it
Line plots are best for continuous or ordinal data where order matters. Compare them with scatter plots, which are better for showing relationships between two variables without implying a direct sequence.
| Chart Type | Best For | Data Relationship |
|---|---|---|
| Line Plot | Trends over time/sequence | Ordered, connected points |
| Scatter Plot | Correlation between variables | Independent, unconnected points |
Practice
Guided Exercise: Modify the example above to plot two lines: one for "Product A" sales and another for "Product B" sales. Use different colors and add a legend.
Challenge: Generate random data for 10 days using numpy.random.rand(10) and plot it as a line graph with red markers. Hint: Import numpy and use np.arange(10) for the x-axis.
Quick check
Question: What parameter in plt.plot() allows you to change the symbol displayed at each data point?
Answer: The marker parameter (e.g., marker='o' for circles).
Summary
Line plots are essential for visualizing trends in sequential data. By mastering plt.plot() and its customization options like markers and colors, you can transform raw numbers into clear, actionable insights.