Mục lục khoá họcĐang ở M3-09
← Khoá Math
0/46 bài đã học xong
Xử lý số liệu · Bài M3-09 · Bài 9/10 của chương · bài 37/46 của Toán

Samples, Inference and Margin of Error

Mẫu, suy luận ra tổng thể, biên sai số

Học xong bài này bạn làm được gì
Đi từ một con số đo trên mẫu ra một ước lượng cho tổng thể, kể cả khi đề bắt đổi đơn vị hai lần hoặc chồng tỉ lệ lên tỉ lệ. Đọc được một khoảng “ước lượng $\pm$ biên sai số” và nói đúng nó nói về ai — về tổng thể, không phải về từng cá nhân, cũng không phải về mẫu. Biết cái gì làm khoảng đó hẹp lại, cái gì không. So sánh độ lệch chuẩn của hai tập bằng mắt, và nói được một thay đổi làm nó tăng hay giảm — SAT không bắt tính bằng công thức, nhưng bắt trả lời chắc chắn.

Bài này gỡ cho bạn chuyện gì

Câu hỏi đầu tiên phải đặt khi đọc đề dạng này không phải "tính gì", mà là: "mẫu được rút ngẫu nhiên từ đâu?"

Cụm từ đó nằm ngay trong đề — randomly selected from… — và nó quyết định kết luận đi được xa tới đâu. Rút ngẫu nhiên $200$ sinh viên từ một trường thì chỉ nói được về trường đó, dù cỡ mẫu có lớn cỡ nào.

Còn biên sai số thì dịch sát nghĩa như sau: "biên sai số $0{,}8$" nghĩa là con số thật của cả nhóm có thể là bất kỳ giá trị nào trong khoảng $\pm 0{,}8$. Nó không nói gì về một cá nhân cụ thể.

Cần nắm trước khi vào bài

Thuật ngữ của bài — cầm sẵn rồi hãy đọc

Thuật ngữ trong đềTiếng ViệtNhớ thế này
populationtổng thểtoàn bộ nhóm muốn kết luận về
samplemẫunhóm nhỏ thật sự được khảo sát
randomly selected fromchọn ngẫu nhiên từcụm quyết định phạm vi kết luận
sample proportiontỉ lệ trong mẫutính trong mẫu trước, rồi mới nhân cỡ tổng thể
margin of errorbiên sai sốkhoảng $\pm$ quanh ước lượng
plausible valuescác giá trị hợp lýý nghĩa thật của khoảng tin cậy
confidence levelmức tin cậynâng mức tin cậy thì khoảng rộng ra
parametertham số tổng thểcon số thật của cả nhóm — thứ khoảng tin cậy nói về

Sample, population, and the one job of a sample

A population is every member of the group the question is really about. A sample is the part of it that was actually measured. The number computed from the sample — a proportion, a mean — is an estimate of the corresponding number for the population.

A sample is only worth anything when it was selected at random from the population, and the conclusion may only reach as far as the group that random selection covered. Randomly selecting $200$ students from one school gives an estimate for that school and for nothing wider.

Giải thích tiếng Việt

Câu hỏi đầu tiên phải hỏi khi đọc đề dạng này không phải “tính gì” mà là “mẫu được rút ngẫu nhiên từ đâu”. Cụm từ đó nằm ngay trong đề (randomly selected from all residents of the city, randomly selected from the members of the club…) và nó khoá luôn phạm vi của kết luận. Chỗ này chương trình phổ thông Việt Nam gần như không dạy, nên phần lớn học sinh bỏ qua nó và mất câu vì chọn phương án nói về một nhóm rộng hơn nhóm được lấy mẫu.

From a sample proportion to a population count

The one line of arithmetic

\[ \text{estimated population count}\;=\;\frac{\text{count in the sample}}{\text{sample size}}\times\text{population size}. \] The fraction must be formed inside the sample first; only then does the population size enter.

Ví dụ mẫu

A random sample of $250$ of the $48{,}000$ residents of a city was surveyed, and $160$ of them said they use the city library. Based on the sample, what is the best estimate of the number of residents who use the library?

Thử tự làm ra giấy trước đã.
Ví dụ mẫu

A random sample of $45$ of the $3{,}600$ crates in a warehouse was weighed; the mean mass of the sampled crates was $18.4$ kilograms. Based on the sample, estimate the total mass of all crates in the warehouse, in metric tons. (1 metric ton $=1{,}000$ kilograms.)

Thử tự làm ra giấy trước đã.
Ví dụ mẫu

A random sample of $300$ of the $7{,}500$ members of an association was surveyed. Of those sampled, $24\%$ had attended the annual conference, and $\tfrac58$ of the attenders renewed their membership. Estimate the number of members of the association who both attended and renewed.

Thử tự làm ra giấy trước đã.
Lỗi thường gặp

Nhân tỉ lệ thứ hai với cỡ mẫu thay vì với cỡ tổng thể, hoặc dùng cỡ mẫu $300$ ở đâu đó trong phép nhân. Cỡ mẫu chỉ có một vai trò: cho phép đem tỉ lệ của mẫu áp sang tổng thể. Sau khi đã có tỉ lệ rồi, con số duy nhất được nhân vào là cỡ tổng thể.

Margin of error — an interval of plausible values

A study rarely reports one number. It reports an estimate together with a margin of error at some confidence level: \[ \text{estimate}\pm\text{margin of error}\quad\longrightarrow\quad \text{interval of plausible values for the \emph{population} value.} \]

What the interval is about, and what it is not about
  • It is about the population parameter — the mean or proportion for everyone.
  • It is not about individual members: the interval does not say that most individuals fall inside it.
  • It is not about the sample: the sample value is known exactly, it is the middle of the interval.
  • It does not say the population value is definitely inside; it says the values inside are the plausible ones at that confidence level.
Ví dụ mẫu

A random sample of the students at a university reported a mean of $12.4$ hours of study per week, with a margin of error of $0.8$ hour at the $95\%$ confidence level. Which conclusion is appropriate?

Thử tự làm ra giấy trước đã.
Giải thích tiếng Việt

Dịch sát nghĩa: “biên sai số $0.8$” nghĩa là con số thật của cả trường có thể là bất kỳ giá trị nào trong khoảng $11.6$–$13.2$. Nó không nói gì về một sinh viên cụ thể, và cũng không nói $12.4$ là sai. Phương án đúng của SAT ở dạng này gần như luôn chứa chữ plausible và một khoảng hai đầu, còn phương án sai thì nói về từng người hoặc nói chắc chắn.

What narrows the interval

At a fixed confidence level, a larger random sample gives a smaller margin of error. Less variability in the data also narrows it. Raising the confidence level widens it. Repeating the study with the same sample size does not narrow it, and neither does surveying a different population.

Ví dụ mẫu

Town A has $40{,}000$ residents and Town B has $400{,}000$. A pollster surveys $1{,}000$ randomly selected residents of each town and reports each estimate at the $95\%$ confidence level. How do the two margins of error compare?

Thử tự làm ra giấy trước đã.
Lỗi thường gặp

Nhầm “cỡ mẫu lớn hơn” với “tổng thể lớn hơn”. Biên sai số phụ thuộc vào số người được hỏi, không phụ thuộc vào số người trong thành phố. Hỏi $1{,}000$ người ở một thành phố nửa triệu dân cho biên sai số gần y hệt như hỏi $1{,}000$ người ở một thành phố năm triệu dân.

Standard deviation, read without computing

Standard deviation measures how far the values sit from their mean. The exam asks for comparisons and for the direction of a change, never for the value itself.

Four rules that settle almost every question
  • Values packed tightly around the mean $\rightarrow$ small standard deviation; values pushed out towards the extremes $\rightarrow$ large.
  • Adding the same number to every value shifts the mean by that number and leaves the standard deviation unchanged.
  • Removing a value far from the mean decreases the standard deviation; adding one far from the mean increases it.
  • Adding a value equal to the mean leaves the mean alone but decreases the standard deviation — the same total spread is now shared among more values.
Ví dụ mẫu

Set M is $10,20,30,40,50$ and Set N is $28,29,30,31,32$. Compare their means and their standard deviations.

Thử tự làm ra giấy trước đã.
Ví dụ mẫu

The values $3,7,8,10,12$ are each multiplied by $3$. What happens to the mean and to the standard deviation?

Thử tự làm ra giấy trước đã.
Giải thích tiếng Việt

Cách nghĩ an toàn: độ lệch chuẩn đo khoảng cách giữa các số với nhau, không đo chúng nằm ở đâu trên trục. Dời cả tập sang phải $6$ đơn vị thì mọi khoảng cách giữ nguyên nên độ lệch chuẩn giữ nguyên. Nhưng kéo giãn tập (nhân mọi giá trị với $2$) thì mọi khoảng cách gấp đôi, độ lệch chuẩn gấp đôi theo.

Which conclusion the study allows

Ví dụ mẫu

A researcher randomly selected $200$ members of a university's chess club and found that they slept a mean of $6.8$ hours per night. Which conclusion is appropriate?

Thử tự làm ra giấy trước đã.

Ba chỗ mất điểm của dạng này

Nhân tỉ lệ với cỡ mẫu thay vì cỡ tổng thể

Cỡ mẫu chỉ có một vai trò: cho phép lập tỉ lệ. Sau đó nó biến mất — phép nhân cuối dùng cỡ tổng thể.

Nhầm "cỡ mẫu lớn hơn" với "tổng thể lớn hơn"

Biên sai số phụ thuộc vào số người được hỏi, không phụ thuộc số người trong thành phố. Hỏi $1000$ người ở một thị trấn nhỏ cho biên sai số hẹp hơn hỏi $400$ người ở một thành phố lớn.

Hiểu khoảng tin cậy là khoảng chứa từng cá nhân

"Trung bình $12{,}4$, biên sai số $0{,}8$" không nghĩa là mọi sinh viên làm việc $11{,}6$–$13{,}2$ giờ. Nó nói về trung bình của cả trường, không nói về ai cả.

Tự kiểm tra — chữa ngay tại chỗ

Câu 1
A random sample of $250$ of the $48{,}000$ residents of a city was surveyed, and $160$ of them said they use the city library. Based on this sample, what is the best estimate of the number of residents who use the library?
Gõ đáp số rồi bấm Kiểm tra.
Câu 2
A random sample of students at a university was surveyed about paid work. The mean was $12.4$ hours per week, with a margin of error of $0.8$ hours. Which is the best interpretation of these results?
  1. Every student at the university works between $11.6$ and $13.2$ hours per week.
  2. The mean number of hours worked by all students at the university is likely between $11.6$ and $13.2$.
  3. About $80\%$ of students work close to $12.4$ hours per week.
  4. The mean of the sample is between $11.6$ and $13.2$ hours.
Chọn một phương án rồi bấm Kiểm tra.
Câu 3
Researchers want to reduce the margin of error in a future survey, keeping the same confidence level. Which change would most likely accomplish this?
  1. Surveying a larger random sample
  2. Conducting the survey in a city with a larger population
  3. Raising the confidence level from $95\%$ to $99\%$
  4. Reporting the median instead of the mean
Chọn một phương án rồi bấm Kiểm tra.

Tóm tắt bỏ túi

  • Câu hỏi đầu tiên: "mẫu được rút ngẫu nhiên từ đâu?" — cụm đó quyết định kết luận đi được xa tới đâu.
  • Ước lượng số lượng: lập tỉ lệ trong mẫu trước, rồi nhân cỡ tổng thể. Cỡ mẫu xong việc là biến mất.
  • Biên sai số nói về con số thật của cả nhóm, không nói về từng cá nhân, cũng không nói về mẫu.
  • Cỡ mẫu lớn hơn → khoảng hẹp hơn. Cỡ tổng thể không ảnh hưởng.
  • Ba thứ làm hẹp khoảng: mẫu lớn hơn · dữ liệu ít phân tán hơn · mức tin cậy thấp hơn.
  • Lặp lại nghiên cứu với mẫu mới thì ước lượng sẽ khác đi một chút — đó là chuyện bình thường, không phải lỗi.
  • Mẫu không ngẫu nhiên (tự nguyện, tiện đâu hỏi đó) thì không suy ra được gì cho tổng thể, dù cỡ mẫu lớn tới đâu.

Làm bài tập của bài này

Bộ M3-09 — 18 câu, chấm ngay trên máy, có đáp án từng câu. Vài câu đầu ở mức nền tảng để chắc tay, phần còn lại từ mức đề thật trở lên.

3 dễ8 khó7 rất khó