Can five data points predict a sixth? A student who studied 1 to 5 hours scored 52, 55, 61, 64, and 70 on practice tests. Linear regression turns those five scores into one line that predicts about 73.9 after 6 hours. This guide shows how that line is built, how to read it, and when to distrust it.
| What it does | Fits the straight line y = mx + b that best matches paired data |
|---|---|
| How it fits | Least squares: the smallest total of squared vertical misses |
| Slope formula | m = sum of (x – mean x)(y – mean y) / sum of (x – mean x)^2 |
| Intercept formula | b = mean y – m x mean x |
| Fit score | R-squared, from 0 (no linear pattern) to 1 (perfect line) |
| Main risk | Outliers and predictions far outside the measured range |
Linear Regression Quick Reference
Every simple regression uses the same small set of pieces. The table below lists each one with its value from the study-hours data. Keep it handy while you read the rest of the guide.
| Piece | Symbol | Study-hours value |
|---|---|---|
| Predictor (input) | x | Hours studied, 1 to 5 |
| Response (output) | y | Test score, 52 to 70 |
| Mean of x and mean of y | x-bar, y-bar | 3 and 60.4 |
| Slope | m | 4.5 points per hour |
| Intercept | b | 46.9 |
| Fitted line | y-hat = mx + b | y-hat = 4.5x + 46.9 |
| Coefficient of determination | R-squared | 0.9868, or about 98.7% |
| Prediction at x = 6 | y-hat | 73.9 |
The little hat on y-hat marks a predicted value, not a measured one. That difference drives the whole method, as the next section shows.
What Does Linear Regression Actually Do?
Linear regression draws the single straight line that sits closest to a cloud of paired points. It summarizes how one variable tends to change as another changes, then uses that line to make predictions.
No straight line passes through every real data point. Each point misses the line by some vertical gap, called a residual. The method picks the line that makes the sum of those squared gaps as small as possible.
Why square the gaps instead of just adding them? Positive and negative misses would cancel out, so a bad line could still total zero. Squaring makes every miss count. It also punishes one big miss more than several small ones, since a gap of 4 counts as 16.
The method is old and widely trusted. The NIST statistics handbook calls it the most widely used modeling method, dating to work by Gauss and Legendre around 1800.
- Residual
- The actual y minus the predicted y. A point above the line has a positive residual.
- Least squares
- The rule that chooses the line with the smallest sum of squared residuals.
- Predictor and response
- The predictor, x, is the input you know. The response, y, is the result you want to estimate.
- Interpolation
- Predicting inside the range of x values you measured. It is the safe use of a regression line.
- Extrapolation
- Predicting beyond the measured range. The line may not hold out there.
How Is the Best-Fit Line Calculated?
You find the means, measure each point’s distance from them, then divide two sums. The slope comes first, and the intercept follows from it.
| x | y | x – 3 | y – 60.4 | Product | (x – 3)^2 |
|---|---|---|---|---|---|
| 1 | 52 | -2 | -8.4 | 16.8 | 4 |
| 2 | 55 | -1 | -5.4 | 5.4 | 1 |
| 3 | 61 | 0 | 0.6 | 0 | 0 |
| 4 | 64 | 1 | 3.6 | 3.6 | 1 |
| 5 | 70 | 2 | 9.6 | 19.2 | 4 |
| Sums | 45 | 10 | |||
The slope is 45 divided by 10, which equals 4.5. The intercept is 60.4 minus 4.5 times 3, which gives 46.9. The finished line is y-hat = 4.5x + 46.9.
Notice that the line always passes through the point of means, here (3, 60.4). That fact makes a quick sanity check on any hand calculation. For longer lists, the linear regression calculator returns the slope, intercept, and fit from pasted x and y values.
How Do You Read the Slope and Intercept?
The slope is the predicted change in y for each one-unit rise in x. The intercept is the predicted y when x equals zero. Both only make sense in the units of your data.
Here the slope says each extra hour of study goes with about 4.5 more points. A negative slope would mean y falls as x rises. A slope near zero means x tells you little about y.
The intercept of 46.9 is the predicted score after zero hours of study. That reading is fair here, because zero hours is close to the data. For many data sets, x = 0 sits far outside the measured range. Then treat the intercept as a fitting constant, not a real-world fact.
The slope also links to spread. It equals the correlation r times the standard deviation of y, divided by the standard deviation of x. With r = 0.9934, a y spread of 7.16, and an x spread of 1.58, you get 4.5 again. Our guide to how standard deviation measures spread covers that second ingredient in depth.
What Does R-Squared Tell You About the Fit?
R-squared is the share of the variation in y that the line explains. A value of 0.9868 means the line accounts for about 98.7% of the ups and downs in the scores.
The calculation compares two totals. The total variation is the sum of squared distances from each y to the mean, which is 205.2 here. The leftover variation is the sum of squared residuals, which is 2.7. One minus 2.7 divided by 205.2 gives 0.9868.
A high R-squared does not prove the model is right. A curved pattern can still score high with a straight line. Always look at the residuals as well as the single number.
When Does a Regression Line Point You the Wrong Way?
A regression line misleads when one stray point, a curved pattern, or a far-off prediction breaks its assumptions. Four checks catch most problems before they reach a decision.
One Outlier Can Flatten the Line
Change the last score from 70 to 50 and refit. The slope drops from 4.5 to 0.5, and R-squared falls from 0.99 to 0.02. One point out of five erased the trend, so check odd values before you trust any fit.
Predictions Beyond the Data Break Down
The line predicts 73.9 at 6 hours, just past the data. At 12 hours it predicts 100.9, which is impossible on a 100-point test. Scores level off in real life, and a straight line never does.
Curves Need a Different Model
Plot the points before fitting. Residuals that form a U shape or an arch signal a curve, and a straight line will miss it in a steady way.
A Strong Fit Is Not a Cause
A tight line shows the two variables move together. It does not show that x drives y. Read why correlation does not prove causation before you claim one variable causes the other.
Paste your x and y lists into the Linear Regression Calculator to get the best-fit line, slope, intercept, and R-squared in one step.
Questions People Ask About Linear Regression
What Is Linear Regression in Simple Terms?
It is a method that draws the straight line closest to a set of paired points. The line shows how y tends to change as x changes, and it lets you estimate y for a new x inside your data range.
How Is Linear Regression Different From Correlation?
Correlation gives one number, r, for how tightly two variables move together. Regression gives a full equation with a slope and intercept, so you can make predictions. The slope equals r times the spread of y divided by the spread of x.
What Is a Good R-Squared Value?
No single cutoff works for every field. Tightly controlled measurements tend to score higher than noisy data about people or markets. A high value still needs a residual check, because a curved pattern can also score well.
How Many Data Points Do You Need for Linear Regression?
The math needs at least two points with different x values. Two points always give a perfect line, so they say nothing about fit. More points spread across a wide range of x give a steadier slope.
Can Linear Regression Predict Values Outside My Data?
It can produce a number, but that number grows less reliable the farther you go. In the study example, 12 hours predicts 100.9 points, which is impossible on a 100-point test. Keep predictions close to the measured range.
Where These Numbers Come From
References Used in This Article
This article explains simple linear regression with one predictor for general learning. The worked numbers use a small illustrative data set, and all values were recomputed for accuracy. Reviewed for accuracy by Prof. Dr. Khalil Mudassar, PhD. Last updated September 27, 2026.
Author
Shakeel Muzaffar is the Founder and Editor-in-Chief of MultiCalculators.com, bringing over 15 years of experience in digital publishing, product strategy, and online tool development. He leads the platform's editorial vision, ensuring every calculator meets strict standards for accuracy, usability, and real-world value. Shakeel personally oversees content quality, formula verification workflows, and the platform's commitment to publishing tools that are genuinely useful for students, professionals, and everyday users worldwide.




