Each dot is a London rental listing. Move the sliders to draw the line you think fits best. Every vertical gap between a dot and your line is an error; least squares picks the line that makes the total area of the squares as small as possible.
The computer finds the line with the smallest possible sum of squared errors. Here is how, and what R² measures.
A good model leaves no pattern: the dots should scatter evenly around zero.
Regression needs numbers, so a category becomes one or more 0/1 columns called dummies. Each dummy's coefficient is a difference from a base category.
Why not include a dummy for every zone? Try it.
So far every dummy shifted the line by the same amount everywhere. An interaction lets the effect of one variable depend on the value of another.
Add and remove predictors, read the coefficient table, and compare models. Every model you fit is saved below so you can compare any two.
Synthetic data generated for teaching, so we know the true relationships. The 20 listings used in tab 1 are highlighted. Download the CSV.