What a line of best fit is — and is not
A scatterplot shows one point for each individual in a data set. The line of best fit is the straight line that comes closest to those points overall; it is a model, not a rule the data obeys. Almost no point lies exactly on it, and that is expected.
For the scatterplot above, the line of best fit is $\hat y=8.5+2.4x$, where $x$ is the number of years since $2010$ and $\hat y$ is the predicted number of subscribers, in thousands. Interpret both numbers.
Slope $2.4$: for each additional year, the model predicts an increase of $\boxed{2.4\text{ thousand subscribers}}$ — that is, $2400$ subscribers per year.
Intercept $8.5$: at $x=0$, which is the year $2010$, the model predicts $\boxed{8.5\text{ thousand subscribers}}$.
The slope is not “$2.4$ subscribers” and not “$2.4$ years”: a slope always carries the unit of $y$ divided by the unit of $x$.
Đơn vị của hệ số góc luôn là “đơn vị của $y$ trên mỗi đơn vị của $x$”. Ở nhóm Xử lý số liệu, câu hỏi hay nhất không phải “hệ số góc bằng bao nhiêu” mà “hệ số góc nghĩa là gì” — và bốn phương án thường chỉ khác nhau ở đơn vị hoặc ở chỗ đảo vai trò $x$ với $y$. Đọc kỹ trục trước khi đọc phương án.
Predicting inside the data, and outside it
Substituting a value of $x$ gives a predicted $y$; solving the equation for $x$ works just as well in the other direction.
A line of best fit passes through the points $(2,15)$ and $(10,39)$ on a scatterplot. Write its equation, then predict $y$ when $x=15$.
Slope $=\dfrac{39-15}{10-2}=\dfrac{24}{8}=3$. Using $(2,15)$: $15=b+3(2)$, so $b=9$ and $\hat y=9+3x$.
At $x=15$: $\hat y=9+45=\boxed{54}$.
Using the same model $\hat y=9+3x$, for what value of $x$ does the model predict $y=60$? The scatterplot's data covered $x$ from $1$ to $12$ only. Comment on the prediction.
Solve $9+3x=60$, so $3x=51$ and $x=\boxed{17}$.
Because $17$ lies well outside the interval $1\le x\le 12$ where the data were collected, the model has never been tested there. Nothing guarantees the pattern continues, so this prediction is far less trustworthy than one made inside the data range.
Dự đoán trong vùng dữ liệu gọi là nội suy, dự đoán ngoài vùng dữ liệu gọi là ngoại suy. Đề SAT rất hay hỏi “vì sao dự đoán này có thể không đáng tin” và đáp án đúng gần như luôn là “giá trị $x$ nằm ngoài khoảng dữ liệu đã dùng để dựng đường thẳng”. Các phương án nhiễu thường nói về độ dốc hoặc về việc đường thẳng không đi qua mọi điểm — đúng nhưng không phải lý do.
Units, and the two-step slope question
The slope is stated in the units of the axes; the question is often asked in different units. Convert after using the model, never before.
A scatterplot shows a subscription's cost $y$, in dollars, against time $x$, in months, and the line of best fit has slope $3$. According to the model, by how many dollars does the cost increase over $4$ years?
The slope means $3$ dollars per month. Four years is $4\times 12=48$ months, so the predicted increase is $3\times 48=\boxed{144}$ dollars.
Nhân hệ số góc với $4$ vì đề nói “4 năm” trong khi trục $x$ tính bằng tháng. Ở nhóm này, đề gần như luôn cố tình cho đơn vị trục khác đơn vị câu hỏi — gam với kilôgam, giây với phút, tháng với năm. Ghi rõ đơn vị của trục ra bên cạnh trước khi tính.
Residuals
For a data point $(x,y)$, the residual is \[ \text{residual}=\text{actual }y-\text{predicted }\hat y . \]
Residual $>0$: the point lies above the line (the model under-predicts). Residual $<0$: the point lies below the line (the model over-predicts). Residual $=0$: the point is exactly on the line.
For the model $\hat y=8.5+2.4x$ of the first section, the data point at $x=7$ has an actual value of $27.1$ thousand subscribers. Find the residual and say what it means.
Predicted: $\hat y=8.5+2.4(7)=8.5+16.8=25.3$. Residual: $27.1-25.3=\boxed{1.8}$ thousand subscribers.
The residual is positive, so that year's actual figure sat $1800$ subscribers above what the model predicted — the point is plotted above the line.
Tính ngược thành “dự đoán trừ thực tế”. Dấu sẽ lật, và câu hỏi “điểm nằm trên hay dưới đường thẳng” sẽ trả lời sai. Thứ tự đúng là thực tế trừ dự đoán — đọc theo nghĩa “dữ liệu vượt mô hình bao nhiêu”.
What a scatterplot licenses you to conclude
A scatterplot shows that two quantities move together. It does not show why.
- Observational study — the researcher only records what already happens. Conclusion allowed: an association between the two quantities, for the group studied. A cause-and-effect claim is not allowed.
- Randomly selected sample — individuals are chosen at random from a population. Conclusion allowed: the finding may be generalised to that population.
- Random assignment to treatment groups — the researcher decides who gets what. Conclusion allowed: a cause-and-effect claim, for the individuals in the study.
Random selection buys generalisation; random assignment buys causation. They are different words and they buy different things.
Researchers recorded the weekly hours of exercise and the resting heart rate of $500$ adults who volunteered for a health screening. The scatterplot shows lower resting heart rates for adults who exercise more. Which conclusion is supported?
The researchers assigned nothing — they only recorded. So no cause-and-effect claim is available: adults who already exercise more may differ from the others in diet, age or genetics. The volunteers were also not selected at random from all adults, so the finding cannot be extended beyond them. The supported conclusion is exactly this: among these $500$ adults, more weekly exercise is associated with a lower resting heart rate.
Đây là chỗ mất điểm nhiều nhất của học sinh Việt Nam ở nhóm này, vì chương trình phổ thông không dạy. Học thuộc hai câu: chỉ quan sát thì chỉ được nói “có liên hệ”; phải có phân nhóm ngẫu nhiên mới được nói “gây ra”. Và một câu thứ ba: muốn suy rộng cho cả tổng thể thì mẫu phải được chọn ngẫu nhiên từ tổng thể đó — người tình nguyện không phải mẫu ngẫu nhiên.
When a line is the wrong model
A straight line encodes constant change per unit. Data that grow by a constant percentage per unit curve upward and need an exponential model $\hat y=a\,b^{x}$.
At $x=0,1,2,3,4$ a quantity takes the values $5,\;15,\;45,\;135,\;405$. Is a linear or an exponential model appropriate?
Differences: $10,30,90,270$ — not constant, so not linear. Ratios: $\frac{15}{5}=\frac{45}{15}=\frac{135}{45}=\frac{405}{135}=3$ — constant, so the model is $\boxed{\hat y=5\cdot 3^{x}}$, exponential growth.
A quantity is modelled by $\hat y=40(1.15)^{x}$, where $x$ is measured in years. What percentage increase does the model predict from $x=0$ to $x=2$?
The predicted values are $40$ and $40(1.15)^2=40(1.3225)=52.9$. The increase is $\dfrac{52.9-40}{40}=0.3225$, that is $\boxed{32.25\%}$ — not $15\%+15\%=30\%$, because the second year's growth applies to a base that already grew.
Phần trăm cộng dồn không phải phép cộng. Tăng $15\%$ hai lần cho hệ số $1.15\times 1.15=1.3225$, tức $32.25\%$, chứ không phải $30\%$. Mọi câu “phần trăm của phần trăm” ở nhóm này đều xử lý bằng cách nhân các hệ số rồi trừ $1$.