Sample, population, and the one job of a sample
A population is every member of the group the question is really about. A sample is the part of it that was actually measured. The number computed from the sample — a proportion, a mean — is an estimate of the corresponding number for the population.
A sample is only worth anything when it was selected at random from the population, and the conclusion may only reach as far as the group that random selection covered. Randomly selecting $200$ students from one school gives an estimate for that school and for nothing wider.
Câu hỏi đầu tiên phải hỏi khi đọc đề dạng này không phải “tính gì” mà là “mẫu được rút ngẫu nhiên từ đâu”. Cụm từ đó nằm ngay trong đề (randomly selected from all residents of the city, randomly selected from the members of the club…) và nó khoá luôn phạm vi của kết luận. Chỗ này chương trình phổ thông Việt Nam gần như không dạy, nên phần lớn học sinh bỏ qua nó và mất câu vì chọn phương án nói về một nhóm rộng hơn nhóm được lấy mẫu.
From a sample proportion to a population count
\[ \text{estimated population count}\;=\;\frac{\text{count in the sample}}{\text{sample size}}\times\text{population size}. \] The fraction must be formed inside the sample first; only then does the population size enter.
A random sample of $250$ of the $48{,}000$ residents of a city was surveyed, and $160$ of them said they use the city library. Based on the sample, what is the best estimate of the number of residents who use the library?
Sample proportion: $\dfrac{160}{250}=0.64$. Estimate: $0.64\times48{,}000=\boxed{30{,}720}$ residents.
A random sample of $45$ of the $3{,}600$ crates in a warehouse was weighed; the mean mass of the sampled crates was $18.4$ kilograms. Based on the sample, estimate the total mass of all crates in the warehouse, in metric tons. (1 metric ton $=1{,}000$ kilograms.)
Estimated total mass: $3{,}600\times18.4=66{,}240$ kilograms.
Convert once, at the end: $66{,}240\div1{,}000=\boxed{66.24}$ metric tons.
A random sample of $300$ of the $7{,}500$ members of an association was surveyed. Of those sampled, $24\%$ had attended the annual conference, and $\tfrac58$ of the attenders renewed their membership. Estimate the number of members of the association who both attended and renewed.
Attended, as a share of the association: $24\%$. Attended and renewed: $0.24\times\tfrac58=0.15$, that is $15\%$ of all members.
Estimate: $0.15\times7{,}500=\boxed{1{,}125}$ members.
The sample size $300$ never enters the arithmetic — it only justifies using the sample percentages for the whole association.
Nhân tỉ lệ thứ hai với cỡ mẫu thay vì với cỡ tổng thể, hoặc dùng cỡ mẫu $300$ ở đâu đó trong phép nhân. Cỡ mẫu chỉ có một vai trò: cho phép đem tỉ lệ của mẫu áp sang tổng thể. Sau khi đã có tỉ lệ rồi, con số duy nhất được nhân vào là cỡ tổng thể.
Margin of error — an interval of plausible values
A study rarely reports one number. It reports an estimate together with a margin of error at some confidence level: \[ \text{estimate}\pm\text{margin of error}\quad\longrightarrow\quad \text{interval of plausible values for the \emph{population} value.} \]
- It is about the population parameter — the mean or proportion for everyone.
- It is not about individual members: the interval does not say that most individuals fall inside it.
- It is not about the sample: the sample value is known exactly, it is the middle of the interval.
- It does not say the population value is definitely inside; it says the values inside are the plausible ones at that confidence level.
A random sample of the students at a university reported a mean of $12.4$ hours of study per week, with a margin of error of $0.8$ hour at the $95\%$ confidence level. Which conclusion is appropriate?
The interval is $12.4-0.8=11.6$ to $12.4+0.8=13.2$ hours.
Appropriate: it is plausible that the mean study time of all students at the university is between $11.6$ and $13.2$ hours per week.
Not appropriate: “$95\%$ of students study between $11.6$ and $13.2$ hours” (that is a statement about individuals), or “the mean is exactly $12.4$ hours” (that discards the margin entirely), or any statement about students at a different university.
Dịch sát nghĩa: “biên sai số $0.8$” nghĩa là con số thật của cả trường có thể là bất kỳ giá trị nào trong khoảng $11.6$–$13.2$. Nó không nói gì về một sinh viên cụ thể, và cũng không nói $12.4$ là sai. Phương án đúng của SAT ở dạng này gần như luôn chứa chữ plausible và một khoảng hai đầu, còn phương án sai thì nói về từng người hoặc nói chắc chắn.
What narrows the interval
At a fixed confidence level, a larger random sample gives a smaller margin of error. Less variability in the data also narrows it. Raising the confidence level widens it. Repeating the study with the same sample size does not narrow it, and neither does surveying a different population.
Town A has $40{,}000$ residents and Town B has $400{,}000$. A pollster surveys $1{,}000$ randomly selected residents of each town and reports each estimate at the $95\%$ confidence level. How do the two margins of error compare?
They are essentially the same. The margin of error is driven by how many people were asked and how much their answers vary, not by how many people the town contains. Town B being ten times larger does not force a wider interval, and surveying $1\%$ of one town versus $0.25\%$ of the other is not the quantity that matters.
Nhầm “cỡ mẫu lớn hơn” với “tổng thể lớn hơn”. Biên sai số phụ thuộc vào số người được hỏi, không phụ thuộc vào số người trong thành phố. Hỏi $1{,}000$ người ở một thành phố nửa triệu dân cho biên sai số gần y hệt như hỏi $1{,}000$ người ở một thành phố năm triệu dân.
Standard deviation, read without computing
Standard deviation measures how far the values sit from their mean. The exam asks for comparisons and for the direction of a change, never for the value itself.
- Values packed tightly around the mean $\rightarrow$ small standard deviation; values pushed out towards the extremes $\rightarrow$ large.
- Adding the same number to every value shifts the mean by that number and leaves the standard deviation unchanged.
- Removing a value far from the mean decreases the standard deviation; adding one far from the mean increases it.
- Adding a value equal to the mean leaves the mean alone but decreases the standard deviation — the same total spread is now shared among more values.
Set M is $10,20,30,40,50$ and Set N is $28,29,30,31,32$. Compare their means and their standard deviations.
Both means are $30$. In Set N every value sits within $2$ of the mean; in Set M the values reach $20$ away from it. So the means are equal and Set M has the greater standard deviation.
The values $3,7,8,10,12$ are each multiplied by $3$. What happens to the mean and to the standard deviation?
The mean was $8$ and becomes $24$: it is multiplied by $3$. Every gap between two values is also multiplied by $3$, so the spread is stretched by the same factor — the standard deviation is multiplied by $3$ as well. (Contrast this with a shift: moving every value the same distance in the same direction changes no gap at all.)
Cách nghĩ an toàn: độ lệch chuẩn đo khoảng cách giữa các số với nhau, không đo chúng nằm ở đâu trên trục. Dời cả tập sang phải $6$ đơn vị thì mọi khoảng cách giữ nguyên nên độ lệch chuẩn giữ nguyên. Nhưng kéo giãn tập (nhân mọi giá trị với $2$) thì mọi khoảng cách gấp đôi, độ lệch chuẩn gấp đôi theo.
Which conclusion the study allows
A researcher randomly selected $200$ members of a university's chess club and found that they slept a mean of $6.8$ hours per night. Which conclusion is appropriate?
The random selection was made from the chess club, so the estimate applies to the chess club. A conclusion about all students at the university, or about students in general, reaches beyond the group that was sampled — chess club members may differ from other students in exactly the way being measured.