Definitions & Key takeaways

Information bias or measurement bias is a type of bias or error that can occur when researchers are unable to collect accurate data. Information bias can be dangerous because it can lead people to believe they have all the information they need to make a decision, when in fact they may be missing key information, or what they have doesn't correlate with reality. This can cause people to act on inaccurate information, which can lead to costly mistakes.

Information bias or measurement bias is a type of bias or error that can occur when researchers are unable to collect accurate data.
Typically, information can be misclassified in two ways - differential, when information collected from one group is accurate but information collected from the other group is inaccurate, and non- differential, when information collected from both groups is inaccurate.
For example, let’s say you want to figure out if flossing teeth prevents cavities. So you follow 100 people who floss, and 100 people who don’t floss, over the course of ten years, and find out that 30% of the flossers and 60% of the non- flossers ended up getting cavities.
Now, we can divide these proportions, 60% divided by 30%, and conclude that people who don’t floss have 2 times the risk of getting cavities compared to people that do floss.
Now, in this study we assume that every person in the flossing group is going to floss their teeth every single day for ten years, and every person in the non- flossing group is not going to floss their teeth for ten years.
But sometimes people in one of the study groups don’t stick with their exposure for the entire study period. For example, maybe some people in the flossing group stopped flossing halfway through the study, because they ran out of floss and just never bought more.
In this case, they would still be counted by the researchers as part of the flossing group even though they technically switched over to the non- flossing group.
On the other hand, everyone in the non- flossing group stayed in the non-flossing group, meaning that they all really didn’t floss for the entire ten years.
This would cause differential misclassification, because information collected from the non- flossing group would be accurate, but information collected from the flossing group would be inaccurate.
So how does differential misclassification affect the results of a study? Let’s assume that flossing actually does decrease the risk of cavities, so the people from the flossing group who stopped flossing halfway through the study actually had a higher risk of cavities than people who kept flossing for the whole study, which would cause an overall increase in the risk of cavities for the flossing group.
So, going back to our study, we might find that 40% of people in the flossing group got cavities, and we’d be unaware that the true risk is 30%.
Now when we compare the proportions of people who got cavities - 60% in the non- flossing group and divide it by 40% in the flossing group – we end up with 1.5 times the risk of getting cavities for people who don’t floss compared to those that do floss.
In this case the true risk of cavities was underestimated. But although differential misclassification can sometimes underestimate the effect of the exposure on the outcome, it can also overestimate it; so it can have very unpredictable effect on the results.
On the flip side, non- differential misclassification occurs when people in both study groups don’t stick with the exposure, so the information collected from both groups is inaccurate.
For example, just like before, some people in the flossing group stopped flossing halfway through the study because they ran out of floss, which increased their risk of getting cavities from 30% to 40%.
But this time, some people in the non- flossing group also started flossing halfway through the study, maybe because they started dating someone who is picky about oral hygiene, which decreased their risk of cavities from 60% to 50%.
So when we compare the proportion of people in the non- flossing group who got cavities - 50% - to the proportion of people who got cavities in the flossing group - 40% - we find that the risk of cavities for people who don’t floss is 1.25 times the risk of cavities for people who do floss, which is a big underestimation of the true risk.
Non- differential misclassification tends to underestimate the effect of the exposure on the outcome, and as a result researchers are often more concerned about differential misclassification compared to non- differential misclassification.
Now, there are three common types of information biases and they can all lead to differential or non- differential misclassification depending on the situation: surveillance bias, recall bias, and surrogate interview bias.
The first is surveillance bias which occurs when one group is monitored much more closely than another group, which can happen if researchers believe that one group is more at risk for the outcome than another group.
For example, researchers might do a more thorough dental exam for those that don’t floss, because they want to make sure they spot any new cavities that might’ve developed since their last exam.
And they might do a less thorough dental exam for those that do floss, because they don’t expect to find any new cavities.
This could cause an underreporting of cavities - the outcome - in the flossing group, and that would overestimate the protective effect of flossing on cavities.
To avoid surveillance bias, researchers can either blind or mask the individuals who are doing the exams, so that they don’t know who is in the flossing group versus the non-flossing group.
If blinding isn’t possible, researchers can write extremely specific protocols, like how to perform each step of a dental exam, that must be followed for all study participants.
Another type of information bias is recall bias or reporting bias. This can happen when individuals in one group report things differently than those in the other group.
Recall bias is common in case- control studies, which compare the exposure history of those that have a certain outcome, called cases, and those that don’t have a certain outcome, called controls.
For example, individuals with a lot of cavities, the cases, might say they flossed fewer times in the past ten years than they actually did, since they know that not flossing is a risk factor for cavities.
On the other hand, the individuals with very few cavities, the controls, might accurately remember how often they floss, simply because they care a lot about oral hygiene.
In this example, the case group would underreport flossing - the exposure - which will overestimate the protective effect of flossing.
To help reduce recall bias, researchers often use written records, like medical records, to verify the information collected from individual interviews, or they can study exposures that happened in the recent past.
Lastly, information bias can also occur due to a surrogate interview bias. For example, when a study participant has died or is too ill to provide information, researchers sometimes get information from a family member or close friend.
For example, let’s say we want to figure out the effect of drinking alcohol on pancreatic cancer. To do this, we might enroll individuals with pancreatic cancer using hospital records over the past six months.
But since pancreatic cancer has a high fatality rate, some of the individuals on the list may have died. Instead of leaving deceased individuals out of the study, we may try to interview a family member for any deceased individual.
The problem with using surrogate interviews is that surrogates often don’t have accurate information on the participant’s exposure status.
In addition, surrogates tend to recall exposures as being less severe than they really were. For example, a partner of the participant may recall that the participant had 1 to 2 drinks per week, even if in reality the participant had 4 to 5 drinks per week, which would lead to an underestimation of the effect of drinking alcohol on pancreatic cancer.
Typically, it’s tough to avoid surrogate interview bias, but researchers can verify information using written records if they’re available.
Alright, as a quick recap, information bias can be due to non- differential misclassification if both study groups are affected equally or differential misclassification if it affects one study group more than the other.
Typically, differential misclassification is worse since it has the potential to overestimate or underestimate the study results.
Three common types of information bias are surveillance bias, recall bias, and surrogate interview bias, which can be avoided using blinding or strict data collection protocols, or by verifying data with written records.