Confounding
Definitions & Key takeaways
A confounding variable is a variable that distorts the accurate relationship between exposure and the outcome. For example, suppose you're studying the effects of a new drug. In that case, age might be a confounding variable because the drug may affect people of different ages differently. Confounders can make the outcome and exposure look more or less than they are.
There are many ways to control for confounding variables, such as stratifying your data (i.e., dividing your data into subgroups) or using multivariate analysis. However, there is no perfect way to completely eliminate confounding variables, so always be aware of them and try to account for them as best you can.
A confounder is a variable in a study that distorts the true relationship between an exposure and an outcome, so it looks like the exposure and the outcome are either more associated or less associated than they really are.
For example, let’s say you hear on the news that drinking coffee is associated with developing heart disease, and - because you drink a lot of coffee - you decide to conduct a study to see if this is true.
First, you recruit 100 people that drink coffee and 100 people that don’t drink coffee, follow them for ten years, and then compare the number of people who developed heart disease in each group.
First, off you must really love coffee and be fairly wealthy to spend ten years studying it at the drop of a hat. Now, let’s say that the proportion of people who develop heart disease in the coffee drinking group is - 50 out of 100, or 50% - and proportion of people who develop heart disease in the non-coffee drinking group - is 20 out of 100, or 20%.
Comparing 50% and 20%, you get a relative risk of 2.5, meaning the risk of developing heart disease for people that drink coffee is 2.5 times the risk for people that don’t drink coffee.
The association between coffee drinking and heart disease can be represented by an arrow pointing from the exposure to the outcome.
The arrow represents a potential causal relationship - in other words, coffee drinking potentially causes the development of heart disease.
But does drinking coffee really cause heart disease? Maybe, or maybe there’s a mysterious third variable - like smoking - that’s confounding the relationship, or making it look like there’s an association when there really isn’t one.
To be considered a confounder, two conditions have to be met. The first condition is that a variable has to be associated with the exposure - meaning that the variable is seen to occur significantly more frequently among one group than the other.
So in this case, people that smoke would have to be either more or less likely to drink coffee compared to people that don’t smoke.
For example, 45 people - or 45% - in the coffee drinking group smoked compared to 5 people - or 5% - in the non-coffee drinking group, so people that smoked are 9 times more likely to drink coffee than people that don’t smoke.
Let’s suppose that smoking cigarettes makes people crave coffee, so 90% of people who smoke also drink coffee, while only 50% of people who don’t smoke also drink coffee.
The second condition is that a confounder has to be associated with the outcome, so smoking would have to be associated with developing heart disease.
In our study, of the 50 people that smoked, 40 people - 80% - developed heart disease, and 10 people - 20% - didn’t develop heart disease.
So the risk of heart disease for people that smoke is 4 times higher compared to people that don’t smoke. Since smoking damages the lining of blood vessels, it’s makes sense that people who smoke are more likely to develop heart disease than people that don’t smoke.
This relationship can be represented by drawing an arrow from smoking to heart disease, since an increase in smoking leads to an increase in heart disease.
So, looking at the diagram, we can see that an increase in smoking leads to an increase in coffee drinking and an increase in heart disease.
So even though it looks like there’s a relationship between the two variables, it’s hard to know whether or not heart disease is dependent on coffee, since the risk of heart disease also depends on whether or not a person smokes cigarettes.
In some cases, the mysterious third variable has a different relationship with the exposure - specifically that it’s caused by the exposure.
In that case, the variable may be considered a mediator. For example, let’s say that we wanted to look at the relationship between obesity and heart disease, and let’s say that cholesterol levels are the third variable.
An increase in obesity is typically associated with an increase in cholesterol, and an increase in cholesterol is typically associated with an increased risk of heart disease.
The reason cholesterol isn’t a confounder in this situation is because simply increasing a person’s cholesterol doesn’t necessarily change a person’s weight, so cholesterol has no influence on obesity.
So, here, we could essentially take cholesterol out of the diagram and we’ll still see the true association between obesity and heart disease.
Confounding generally occurs when the two groups being compared aren’t similar to one another - like if there are more smokers in one group than the other.
There are three methods you can use when designing your study to make sure that the two study groups are similar: randomization, restriction of the study population and matching.
Randomized controlled trials use a tool called randomization, meaning individuals get selected to each study group through a process of chance.
Using randomization, there’s a pretty high chance that each group will have similar characteristics. Unfortunately, it may be unethical to assign people into drinking coffee or not drinking coffee if there’s a chance that drinking coffee increases the risk of heart disease, so we can’t use randomization in every study.
Cohort studies and case- control studies don’t randomize individuals to treatment groups, but can avoid confounding in other ways.
For example, they can restrict the study population to only certain characteristics. For example, if you know smoking is a confounder for the relationship between drinking coffee and developing heart disease, then might choose to include only individuals that don’t smoke.
That way, if there’s an association between coffee drinking and heart disease, we know that it’s not due to a difference in smoking habits between the two groups.
Another technique used to control confounders in case- control studies is matching. That’s when the cases - the people that have the outcome, like heart disease - and controls - the people that don’t have the outcome, like no heart disease - are matched based on a certain characteristics, like age, sex, race, socioeconomic status, or occupation.
For example, instead of starting with a group of people who smoke or don’t smoke and following them for ten years, we could start with a group of people who have heart disease or don’t have heart disease, and measure the proportion of people in each group that smoke and don’t smoke.
It’s possible to match individually, like including a person that smokes in the group with heart disease for every person that smokes in the group with no heart disease.
Alternatively you can match by frequency, so that the proportion of controls with a certain characteristic is identical to the proportion of cases with the same characteristic, like if we included 30 men and 20 women in the group with heart disease, we would include 30 men and 20 women in the group without heart disease.
It’s also possible to control for confounders in the analysis stage by matching. For example, you could match each person that smoked in the group that developed heart disease to a person that smoked in the group without heart disease.
This can be a problem though when there are only a few people that meet the criteria in one group, like if the group without heart disease only included 2 people that smoked.
If you match those 2 people with 2 people from the group with heart disease, you would only be able to analyze information from a total of 4 people, which is a very small sample size!
Another way to control for confounding during the analysis stage is by stratification - or doing separate analyses for each strata or level of the confounder.
For example, smoking has two basic strata - yes or no. So, when we analyse the relationship between drinking coffee and heart disease, we could do two separate calculations - one for people that smoke and one for people that don’t smoke.
We can then compare the relative risks of both stratified calculations, as well as the original study’s unstratified relative risk.
As a general rule, if the relative risks in each strata are different than the crude or unstratified relative risk, but the same as each other, then we know that confounding is present.
For example, our crude relative risk was 2.5. Now to figure out the other two relative risks, let’s stratify the results into the 50 people that smoked and the 150 people that didn’t smoke.
Out of the people in the study that smoked, 20 of them drank coffee and 30 of them didn’t drink coffee. Of the coffee drinkers, 10 of them - 50% - developed heart disease, and of the no-coffee drinkers, 12 of them - 40% - also developed heart disease.
Since the proportions are the same, it gives us a relative risk of 1.25. We can do a similar calculation with the other strata, the non-smoking group.
Out of the 150 people in the study that didn’t smoke, 80 of them drank coffee and 70 of them didn’t drink coffee. Out of the 80 that drank coffee, 32 of them - or 40% - developed heart disease and out of the no-coffee drinkers, 21 of them - or 30% - also developed heart disease.
Here, the relative risk is 1.3, which is very similar to the relative risk in the group that didn’t smoke. The average of these two relative risks is 1.29, and this is called the adjusted relative risk.
An adjusted relative risk of 1.29 means that people who drink coffee have 1.29 times the risk of heart disease compared to people who don’t drink coffee, regardless of if they smoke cigarettes or not.
Finally, in some cases, there are multiple variables that confound a relationship. For example, the relationship between coffee drinking and heart disease might be confounded by smoking, and also by age - because older people generally drink more coffee and are at higher risk of heart disease - and by sex - because males tend to drink more coffee than females, and males are at higher risk of heart disease than females.
Stratifying multiple confounders by hand can be very complicated, so this is where statistical models, like linear regression or logistic regression can be really handy.
Alright, as a quick recap, confounders distorts the true relationship between an exposure and an outcome, so it looks like the exposure and the outcome are either more associated or less associated than they really are.
A confounder has to either cause the exposure or be associated with the exposure, and it has to be associated with the outcome.
You can control for confounders in the design stage of a study by randomization in randomized controlled trials, stratification in cohort or case-control studies, and matching in case-control studies.
You can also control for confounders in the analysis stage of a study by matching, stratifying, or adding them to statistical models using computer software.
No notes for this video yet
Try adding a note below