Two-way ANOVA
Definitions & Key takeaways
Two-way ANOVA (Analysis of Variance) is a statistical method used to determine whether there are significant differences between two or more groups of data. It involves testing for two or more factors, or variables, that may influence the outcome of an experiment or study.
Two-way ANOVA allows for the examination of the main effects of each factor, as well as any interaction between the factors. The method calculates an F-value, which is then compared to a critical value to determine statistical significance. Two-way ANOVA is commonly used in various research fields, including medicine, psychology, and social sciences, to analyze data and test hypotheses.
Analysis of variance, or simply, ANOVA, is a type of parametric statistical test used to determine if there’s a significant difference between the means or averages of three or more groups.
And significance is normally defined by a p-value of less than 0.05 or 5%. Now when doing any parametric test, there are three key assumptions that we have to make about the population.
First, the sample population must have been recruited randomly. Choosing names randomly ensures that the people included in the study will have similar characteristics to the target population.
This is important because that ensures that the results of the t-test can be applied to the target population - meaning it has good external validity!
The second assumption is that each individual in the sample was recruited independently from other individuals in the sample.
In other words, no individuals influenced whether or not any other individual was included in the study. For example, if two friends decided to get their blood pressures measured on the same day, and they were both included in the study, these two individuals would not be independent of each other and the second assumption would not be met.
Like random sampling, independent recruitment of individuals is important because it ensures that the sample population approximates the target population.
The third assumption is that the sample size is large enough to approximate the target population, which usually means having more than 20 people.
If it’s impossible to get a large sample size, then the sample population must follow a normal bell-shaped distribution for the characteristic being studied because that’s what we would expect to see in the target population.
Okay, now let’s say there are three medications available for lowering systolic blood pressure, and you want to figure out if any of the medications work differently than the others.
Additionally, you want to figure out if the medications work differently for males and females. So, let’s say that you find 6 people - 3 males and 3 females - who take Medication A for 6 weeks, and that afterwards the mean systolic blood pressure is 130 mmHg for males and 125 mmHg for females.
Then, you find another 3 males and 3 females who have been taking Medication B - and afterwards their mean systolic blood pressure is 138 for males and 126 for females, and finally you find 3 males and 3 females who have been taking Medication C - and afterwards their systolic blood pressure is 132 for males and 125 for females.
We can arrange the numbers in a table like this, with sex on the top and medication type on the side, and each cell represents a mean systolic blood pressure value.
To figure out if medication type and sex have an effect on systolic blood pressure, we can use an ANOVA test. Specifically, we would use a two-way ANOVA test, because we are testing two factors - medication type and sex - that each have multiple groups within them.
There are three groups in the medication factor - which are A, B, and C - and we’ll use two groups for sex - male and female.
A two-way ANOVA tests has six hypotheses. The first three are null hypotheses.
The first null hypothesis says that there’s no difference in systolic blood pressure for people taking different medication types.
The second null hypothesis says that there is no difference in systolic blood pressure for males and females. The third null hypothesis says that there is no interaction between medication type and sex.
Interaction means that one factor influences the relationship between the second factor and the outcome. So, in this example, interaction would mean that the effect of medication on systolic blood pressure is different for males and females.
The alternate hypotheses are the opposite of the null hypotheses. The first one says that there is a difference in systolic blood pressure for people taking different medication types; the second one says that there is a difference in systolic blood pressure for males and females; and the third one says that there is an interaction between medication type and sex.
Now, there are seven steps to test these hypotheses. The first step is to calculate the mean of each factor, so the mean for all the medication types and the mean for both sexes.
To find the overall row mean systolic blood pressure for the medication A group, we add up the mean blood pressures for both males and females in that group, and divide by 2.
So, 130 plus 125, divided by 2, is 127.5. We can do the same thing for the medication B and C groups.
So, 138 plus 126, divided by 2, is 132. And 132 plus 125, divided by 2, is 128.5.
Finally, we can find the grand mean - which is the mean blood pressure measurements for all the medication groups - by adding up each row mean and dividing by 3.
So 127.5 plus 132 plus 128.5, divided by 3, gives us a grand mean of 129. To find the overall column means for the second factor, we add up the mean blood pressures in each medication group and divide by 3.
So for males, 130 plus 138 plus 132, divided by 3, is 133.3. And for females, 125 plus 126 plus 125, divided by 3, equals 125.3.
Notice that if we add up the column means for both sexes and divide by 2, it equals 129, which is the grand mean. The second step is to find the total sum of squares, which tells you how much variation there is in the dependent variables.
In other words, it compares every single individual in the study to the grand mean. If the total sum of squares is large, it means that the individuals in the study, at every time point, are very spread out from one another; and if the total sum of squares is small, it means that the individuals in the study, at every time point, are clustered together.
To get the sum of squares, we start by subtracting the grand mean from each individual’s systolic blood pressure measurement, and then squaring it, which is called the squared difference.
Then, we add up all the squared differences to get the total sum of squares. There are 18 people in this study, so we will have 18 squared differences, but to save time, let’s just do the first three.
The first person in the study is a female who took Medication A, and her systolic blood pressure is 124. So, we subtract the grand mean, which is 129, from 124, which equals negative 5.
Then we square it, so 5-squared is 25. The second person is also a female who took Medication A, and her blood pressure is 126.
So, 126 minus 129 is negative 3, squared, is 9. For the third person, who is also a female that took Medication A, we do 125 minus 129, which is negative 4, and then square it, which is 16.
So, if we add up all of the squared differences for everyone in the study, we get a total sum of squares of 743. The third step is to find the between-group variation, which is also called the sum of squares-between, or the SSB.
The sum of squares-between is a measure of how similar each group’s mean is to the grand mean. And, since we have two factors, we have to find the sum of squares-between for each factor and them add them together to get the total sum of squares-between.
To get the sum of squares-between for the medication factor, we start by subtracting the grand mean from each group’s row mean and squaring it, to get the squared difference.
Then, we multiply the squared difference by the number of people in that group. So, let’s start with the medication factor.
For medication 1, the row mean is 127.5. So, the squared difference is 127.5 minus 129, squared, which is 2.25.
Then, since there are 6 people in the medication group, we multiply 2.25 by 6, which equals 13.5. We do that same thing for the other two medication groups - so, 132 minus 129, squared, times 6, is 54; and 128.5 minus 129, squared, times 6, is 1.5.
When we add up all the squared differences for all the medications, we get 13.5 plus 54 plus 1.5, which equals 69. We do a similar process to get the sum of squares-between for the sex factor, but this time we use the column means.
For males, we subtract the grand mean from the column mean, then square it. So, 133.3 minus 129 is 4.3, and 4.3-squared is 18.5.
Then we multiply it by the number of males in the study, which is 9, so 18.5 times 9 is 166.5. For females, the equation is the same, but we substitute in their column mean, which is 125.3.
So 125.3 minus 129, squared, times 9, is 123. And when we add up both the sum of squares-between for males and females we get 289.5.
The fourth step in the ANOVA calculation is to find the within-group variation, which is also called the sum of squares-within, or SSW.
The sum of squares-within is a measure of how similar each individual blood pressure measurement is from its own group mean.
The sum of squares-within is sometimes called the sum of squared-error, because it’s the variation in systolic blood pressure that can’t be explained by either the different medication or the different sex, so it’s sort of this unexplained error in the blood pressure measurement.
Basically, it’s the variation caused by individual characteristics, like age or genetic differences. To find the sum of squares-within, we divide the individual blood pressure measurements into two tables - one for males and one for females.
Now, let’s work with the male table first. For the sum of squares-within, you start by finding the squared differences for each person in one group, and to do this, you take each individual blood pressure measurement and subtract that group’s mean - which is the row mean - and then square it.
For example, let’s just take the first 3 systolic blood pressure, which are the measurements in the Medication A group. So, it’s 122, 136, and 133.
Since the row mean for Medication A is 130, you subtract 130 from each individual measurement, so 122 minus 130 is negative 8, 136 minus 130 is 6, and 133 minus 130 is 3.
Then you square each number and add them all together to get the squared difference. So, when you add up negative 8-squared, or 64, and 6-squared, or 36, and 3-squared, or 9, you get 109.
To save time, let’s just fill in the rest of the values for the table. But, it’s important to notice that the next three rows used the row mean for Medication B and the last three rows used the row mean for Medication C.
When we sum up all of the squared differences for the males, we get a sum of squares-within of 275. Next, we use the same process to find the squared differences for the female table, but the row means that are used this time are 125, 126, and 125 for the Medication A, B, and C groups.
When we add up all the squared differences for the female group, we get 58. And finally, to get the total sum of squares-within for all individuals in the study, we add up the sum of squares-within for both males and females.
So, 275 plus 58 equals 333. For the fifth step, we have to come back to the total sum of squares, which we calculated in step 2.
Conceptually, we can think of the total sum of squares as the aggregate of all the different sum of squares in a study. So, in a two-way ANOVA, the total sum of squares includes the sum of squares-between for the medication factor, the sum of squares-between for the sex factor, the sum of squares-within, and the combined sum of squares for medication and sex - which we haven’t calculated yet.
In fact, the easiest way to calculate the combined sum of squares is to plug in all the other numbers in the equation, so that the only missing number is the combined sum of squares.
So, the total sum of squares, or 743, equals the sum of squares-between for medication, or 69, plus the sum of squares-between for sex, or 289.5, plus the sum of squares-within, or 333, plus the combined sum of squares.
We can reduce this equation down to 743 equals 691.5 plus the combined sum of squares. And if we subtract 691.5 from both sides, we find that the combined sum of squares equals 51.5.
The sixth step is to calculate the mean square for each of the sum of squares. The mean square is similar to the sum of square, but it takes into account the degrees of freedom, which are values that account for the number of groups, number of individuals in the study, or both.
The equation for the mean square is the sum of squares divided by the degrees of freedom. To calculate the mean square for medication, we divide the sum of squares-between for medication, which is 69, by the degrees of freedom, which is the number of groups minus 1.
Since there are 3 medication groups, the mean square is 69 divided by 3 minus 1, which equals 34.5. For the mean square for sex, we divide the sum of squares-between for sex, which is 289.5, by the number of groups minus 1.
Since the number of sex groups is 2, the mean square is 289.5 divided by 1, which is still 289.5. For the combined mean square for medication and sex, we divide the combined sum of squares-between, which is 51.5, by the number of medication groups minus 1, times the number of sex groups minus 1.
So, the combined mean square is 51.5 divided by 3 minus 1, or 2, times 2 minus 1, or 1, equals 25.75. For the mean square for error, we divide the sum of squares-within, which is 333, by the total number of people in the study, minus the number of medication groups times the number of sex groups.
There are 18 people in the study, so 18 minus 3 times 2, or 6, is 12. So, the mean square of error is 333 divided by 12, which equals 27.8.
The seventh, and final step in the ANOVA test is to calculate the F-statistic or F-stat, and to do this, we divide the mean square for medication, the mean square for sex, and the combined mean square each by the mean square for error.
So, the F-stat for medication is 34.5 over 27.8, or 1.2. The F-stat for sex is 289.5 over 27.8, or 10.4.
And the F-stat for both sex and medication is 25.75 over 27.8, or 0.9. To figure out if these are significant F-stats, we have to compare them to the critical values for the study, which are predetermined numbers used to determine whether or not to reject the null hypothesis.
If the value of an F-stat is greater than its critical value, then the null hypothesis for that factor is false. Critical values can be found on an F-distribution table like this one, which has the degrees of freedom for the numerator of the F-stat equation on the top and the degrees of freedom for denominator of the F-stat equation on the side.
Typically, the significance level for the F-distribution table is 0.05, or 5%. For medication, the numerator degrees of freedom is 2 and the denominator degrees of freedom is 12, so we find a critical value of 3.89.
The value of our F-stat is 1.2, which is below the critical value of 3.89, so we cannot reject the null hypothesis, and we conclude that there is no difference in systolic blood pressure for people using different medication types.
For sex, the numerator degrees of freedom is 1 and the denominator degrees of freedom is 12, so the critical value is 4.75.
The value of our F-stat is 10.4, which is above the critical value of 4.75, so we reject the null hypothesis, and we conclude that there is a difference in systolic blood pressure between males and females.
For medication and sex, the numerator degrees of freedom is 2 and the denominator degrees of freedom is 12, so the critical value is 3.89.
The value of our F-stat is 0.9, which is below the critical value of 3.89, so again, we cannot reject the null hypothesis, and we conclude that there is no interaction between sex and medication.
ANOVA tests are most often calculated using statistical software, and the software will often provide p-values for each factor.
And since a two-way ANOVA has three critical values, then the software will provide three p-values. The p-value is the probability of obtaining a given F-stat or a higher F-stat, if the null hypothesis is true.
In short, the p-values cut out the steps of finding the critical values. So, if we use a significance level of 0.05, then a p-value of less than 0.05 will indicate that the null hypothesis is false.
Finally, it’s important to keep in mind that if the assumptions of an ANOVA test are not met, we can’t be sure that the results can be applied to the target population, so an ANOVA shouldn’t be used.
Instead, for a two-way ANOVA, we could use a non-parametric test called the Friedman two-way test, which doesn’t rely on parametric assumptions.
In short, the Friedman test compares the medians of each group and each factor to determine if one of the group’s medians is different than the others.
So in this case, if the median blood pressure is different for people taking different medications and for males or females.
Parametric tests are generally favoured over non-parametric tests, so the results of the Friedman test are not considered as strong of evidence as the ANOVA test.
Alright, as a quick recap, a two-way ANOVA test is a type of parametric test used to compare the means of multiple groups in multiple factors.
A two-way ANOVA test has three null hypotheses and three alternate hypotheses, and to test these hypotheses we calculate an F-statistic for each factor by dividing the mean square of that factor by the mean square of the error.
This F-statistic can be compared to a critical value to determine if the means of the groups are equal or not. There are three assumptions, random sampling, independent recruitment, and large sample size or normal distribution, that must be met in order to complete an ANOVA test.
And if these assumptions are not met, the Friedman two-way test can be used to determine if the medians of the groups are equal or not.
No notes for this video yet
Try adding a note below