What standard deviation measures
The standard deviation, written SD, answers one question: how far from the mean does a typical value sit? Two data sets with the same mean can look completely different, and the SD is the number that separates them.
SD is never negative. SD $=0$ exactly when every value in the data set is identical. The further values sit from the mean, the larger the SD. SD carries the same unit as the data.
Which list has the greater standard deviation? \[ A:\;10,\;20,\;30,\;40,\;50 \qquad\qquad B:\;28,\;29,\;30,\;31,\;32 \]
Both lists have mean $30$. In $A$ the values sit $20,10,0,10,20$ away from the mean; in $B$ they sit only $2,1,0,1,2$ away. Every value of $B$ is closer to the centre, so $\boxed{A}$ has the greater standard deviation. No formula is needed — and none is expected.
Dịch sang tiếng Việt cho dễ nhớ: độ lệch chuẩn là “trung bình của khoảng cách tới trung bình”. Số liệu túm tụm quanh trung bình thì độ lệch chuẩn nhỏ; số liệu tãi ra hai phía thì lớn. Đề SAT gần như không bao giờ bắt tính con số cụ thể — nó bắt so sánh, hoặc hỏi tăng hay giảm. Nếu bạn đang nhớ công thức căn bậc hai thì bạn đang giải sai bài.
Comparing two data sets without computing
Compare by asking where the weight of the data sits relative to the mean. Two traps appear constantly: equal ranges do not force equal SDs, and equal SDs do not force equal ranges.
Both lists below have mean $6$ and range $8$. Which has the greater standard deviation? \[ P:\;2,\;4,\;6,\;8,\;10 \qquad\qquad Q:\;2,\;2,\;6,\;10,\;10 \]
Distances from the mean are $4,2,0,2,4$ for $P$ and $4,4,0,4,4$ for $Q$. Every value of $Q$ that is not at the centre sits at the far end, while $P$ has two values only halfway out. $\boxed{Q}$ has the greater standard deviation, even though the two lists share a mean, a range, a smallest value and a largest value.
Coi khoảng biến thiên là “phiên bản dễ tính” của độ lệch chuẩn. Khoảng biến thiên chỉ nhìn hai giá trị đầu và cuối; độ lệch chuẩn nhìn tất cả. Hai tập cùng khoảng biến thiên hoàn toàn có thể lệch chuẩn khác hẳn nhau, và ngược lại.
A data set of $12$ values has a mean of $7$ and a standard deviation of $0$. What are its median, mode and range?
An SD of $0$ says no value is any distance from the mean, so all twelve values equal $7$. Then the median is $\boxed{7}$, the mode is $\boxed{7}$ and the range is $\boxed{0}$.
Shifting and scaling
- Add $c$ to every value. The whole data set slides along the number line. Mean, median and mode all increase by $c$; SD and range do not change — sliding a picture does not stretch it.
- Multiply every value by $k$. The picture stretches. Mean and median are multiplied by $k$; SD and range are multiplied by $|k|$.
A data set has mean $18$ and standard deviation $4$. Every value is increased by $25$, and then every resulting value is doubled. What are the new mean and standard deviation?
After the shift: mean $18+25=43$, SD still $4$. After doubling: mean $2\times 43=86$, SD $2\times 4=8$. Final answer: mean $\boxed{86}$, SD $\boxed{8}$.
Cộng luôn hằng số vào độ lệch chuẩn: “SD $=4+25=29$”. Cộng cùng một số vào mọi giá trị thì mọi khoảng cách giữa các giá trị giữ nguyên, nên độ phân tán không đổi. Chỉ phép nhân mới kéo giãn được.
Nhớ bằng hình ảnh: cộng hằng số là dời cả tập số liệu, nhân hằng số là phóng to nó. Dời thì trung tâm đổi mà độ rộng giữ nguyên; phóng to thì cả trung tâm lẫn độ rộng cùng nhân lên. Dấu âm chỉ lật hình qua gốc — độ rộng vẫn dương, nên SD nhân với $|k|$ chứ không nhận giá trị âm.
Adding or removing a value
Compare the new value with the current mean:
- A value equal to the mean adds no new distance but adds one more data point, so the SD decreases.
- A value far from the mean adds a large distance, so the SD increases — and the further out it lands, the bigger the jump.
- Removing a value reverses whichever effect adding it would have had.
The list $8,\;10,\;12$ has mean $10$. A fourth value of $10$ is added. Does the standard deviation rise or fall?
The mean stays $10$, and the distances from the mean are $2,0,2$ before and $2,0,0,2$ after. The same total distance is now shared among four values instead of three, so the typical distance per value goes down: the SD $\boxed{\text{decreases}}$.
The list $30,\;34,\;38,\;42,\;46$ has mean $38$. Of the values $38$, $50$, $26$ and $60$, which one — added as a sixth value — would raise the standard deviation the most?
Measure each candidate's distance from the current mean: $38$ is $0$ away, $50$ is $12$ away, $26$ is $12$ away, and $60$ is $22$ away. The furthest one contributes the most extra spread, so $\boxed{60}$ raises the SD the most; $38$ would lower it.
Set $A$ consists of ten values all equal to $30$. Set $B$ consists of ten values all equal to $50$. Set $C$ is $A$ and $B$ combined. Compare the three standard deviations.
$A$ and $B$ each have every value identical, so SD$(A)=$ SD$(B)=0$. Set $C$ has mean $40$, and every one of its twenty values sits $10$ away from that mean — a typical distance of $10$ where both original sets had a typical distance of $0$. Combining two groups whose means differ creates spread that neither group had on its own: $\boxed{\text{SD}(C)>\text{SD}(A)=\text{SD}(B)}$.
Outliers: which summaries survive
An outlier is a value far from the rest of the data. It moves the summaries that use every value (mean, SD, range) and barely touches the ones that use position or frequency (median, mode).
Compare the summaries of \[ 42,\;44,\;45,\;47,\;48,\;50,\;144 \] with those of the same list once $144$ is removed.
With $144$: sum $=420$ over $7$ values, so the mean is $60$; the median is the fourth value, $47$; the range is $144-42=102$.
Without $144$: sum $=276$ over $6$ values, so the mean is $46$; the median is $\frac{45+47}{2}=46$; the range is $50-42=8$.
The mean falls by $14$ and the range by $94$, while the median falls by only $1$. The SD falls sharply as well, since the one value contributing almost all of the spread is gone.
Trung vị và mode được gọi là bền với giá trị lạ: chúng chỉ quan tâm vị trí và số lần xuất hiện, nên một con số cực đoan đẩy chúng đi rất ít. Trung bình, khoảng biến thiên và độ lệch chuẩn thì không bền. Vì vậy khi một bộ dữ liệu có giá trị lạ, trung vị mới là con số mô tả “giá trị điển hình” trung thực hơn — đây chính là lý do báo cáo thu nhập luôn dùng trung vị.
Nghĩ rằng bỏ giá trị lạ thì trung vị cũng nhảy mạnh như trung bình. Bỏ một giá trị làm danh sách ngắn đi một, nên vị trí giữa chỉ dịch nửa bậc — trung vị thường chỉ đổi vài đơn vị, có khi không đổi.