Range, variance, and standard deviation
Definitions & Key takeaways
Range, variance, and standard deviation are all measures of the spread or dispersion of a set of numerical data. The range is the difference between the highest and lowest values in the data set. Variance is a measure of how far each value in a set of data is from the mean value. Variance is calculated by taking the average of the squared differences of each value from the mean. The larger the variance, the more spread out the data is from the mean. Standard deviation is the square root of the variance and it measures the spread of data around the mean. The larger the standard deviation, the more spread out the data is from the mean.
To understand a set of data - having a single number like the mean or median gives us a one number summary, but understanding how the data is distributed is also very important - and that’s where the range, variance, and standard deviation can be helpful.
For example, let’s say we are looking at the weight of 10 people and we divided them into groups A and B. weight of group A(in kg) weight of group B(in kg) 40 45 50 55 60 10 30 50 70 90 .
Mean = (40+45+50+55+60)/5= 50kg Mean = (10+30+50+70+90)/5= 50kg Now, if you calculate the mean weight of group A and group B, you will find both of them have the same value of 50 kg, but the weights of individuals in group A are much more centered around the mean than in group B.
So let’s start by looking at the range, which is the difference between the highest and lowest value in a dataset. In group A, we have (60-40)=20 kg, whereas in group B we have (90-10)=80 kg.
So far so good. But now, let’s say we have decided to include another group called group C.
Weight of Group C (in kg) 10 45 50 55 90. Mean = (10+45+50+55+90)/5= 50 kg Range 90-10 = 80kg So even when we change two data points, group C still has the same mean and range as group B since it depends only on the highest and lowest values, thus it provides no information about how the rest of the data points are distributed.
In this situation, it’s clear that we need a better idea of how all of the values are distributed and to do that we can look at the variance.
To calculate the variance, which is written out as σ2, we take each data point (x), subtract it from the mean (x-bar), and then we square this value so we don’t end up with a negative number.
Next, we add up the squared values and divide that result with the total number of data points (n). So, let’s use this formula to calculate the variance for group A, B, and C where all three had a mean of 50.
So, for group A, we get: (40-50)2+(45-50)2+(50-50)2+(55-50)2+(60-50)2/(5)=50 kg2 For group B, we get: (10-50)2+(30-50)2+(50-50)2+(70-50)2+(90-50)2 /(5) = 800 kg2 And for group C, we get: (10-50)2+(45-50)2+(50-50)2+(55-50)2+(90-50)2 /(5) = 650 kg2 So, with this new value - the variance - we can tell that the data points in group B are generally farther from the mean than in group C because group B has a larger variance.
By comparison, it’s really clear that the points in A have a much smaller spread, compared to both group B and group C. But one problem with the variance is that you end up with a strange unit for the variance which is kg2 which can’t be intuitively compared to the mean or median which is in kg.
So, to deal with this problem, we can take the square root of the variance which gives us the standard deviation. So, the equation for the standard deviation, or σ, ends up in the same units as the mean.
So, according to this equation, the standard deviation for group A is 50 = 7.1kg, for group B is 800= 28.28kg, and for group C is 650 = 25.5kg.
Now, in situations where the data is normally distributed, where most of the data points tends to center around the central value, and there are fewer data points towards the extreme ends, we call it a bell curve.
If we think about the standard deviation in the context of this sort of normally distributed data, it turns out that 68% of the data points are within one standard deviation either above or below the mean, 95% are within two standard deviations, and 99% of the data are within three standard deviations.
So applying this to a population, let’s say we took the weight of adults living in a country and that we found that the data had a normal distribution.
Let’s say that across the country, the average adult weight was 50kg and that the standard deviation across thousands of people was 7.1kg.
That would mean that 68% of the population weighed between 42.9 and 57.1kg, which is one standard deviation above and below the mean.
That 95% of the population weighed between 35.8 and 64.2 kg, which is two standard deviations above and below the mean. And that 99% of the population weighed between 28.7 and 71.3kg, which is three standard deviations above and below the mean.
Alright, as a quick recap, range, variance and standard deviation are three measures of dispersion. The standard deviation gives the average deviation from the mean for all the data points and has the same unit as the mean.
No notes for this video yet
Try adding a note below