Selection bias refers to a systematic error that occurs when the sample of individuals being studied is not representative of the population of interest. This can lead to incorrect conclusions and invalid results from those research studies. Several types of selection biases include sampling bias, non-response bias, and Berkson's bias.
Sampling bias occurs when the sample of individuals being studied does not cover or represent the population of interest. In non-response bias, some individuals in a study do not respond to a survey or questionnaire, leading to incorrect conclusions and invalid results. Finally, Berkson's bias occurs when individuals in both case and control groups are recruited based on a characteristic that's associated with the exposure.
Selection bias is a type of bias or error that can occur when researchers choose who will be included in a study. Studies with selection bias might end up having results that can’t be applied to the population outside the study - so lacking external validity.
They may also result in an inaccurate representation of the relationship between an exposure and an outcome - so lacking internal validity.
Typically, the goal of a study is to figure out if an exposure is associated with an outcome in a target population. So ideally a study should be done on a sample population of individuals that is similar to that target population in every meaningful way, which would give the study high external validity.
For example, if you want to figure out how smoking impacts the risk of lung cancer in Portland, Oregon, then people living in Portland are your target population.
Ideally, your sample population would include individuals from Portland. And in addition, the sample population should include people of ages, races, and socioeconomic statuses that reflect the target population as well, because these are all factors that are likely to affect the risk of lung cancer.
If your study only recruits students from one of the local high schools, then your sample population probably won’t represent your target population, since the average age in your study will be younger than the average age in Portland, which is 36 years old.
Now, to make the sample population represent the target population, one tool that can be used is randomization, meaning that individuals get selected to enter the study through a process of chance.
To show how that works, let’s say the researchers put the names of every person in Portland into a brown paper bag, which would have to be pretty big, since there would be over 600,000 names in that bag - probably with a number of repeats.
Then let’s say that you choose a thousand names out of the bag to include in the study - either by simply picking them or by using a computer program to make sure that it’s truly by chance.
That’s randomization. Using randomization, there’s a pretty high chance that the sample population and that target population will be similar, and that the study has high external validity, meaning that any conclusions made about the sample population can be applied to the target population.
Sometimes, even when a population is randomly selected, selection bias can still decrease a study’s external validity. For example, perhaps you decide to randomly choose your sample population from a list of all the house addresses in Portland, or from a list of all the phone numbers in Portland.
In this situation, there’s a high chance of sampling bias, which is a type of selection bias. That’s where some individuals in the target population might have a lower chance or no chance of being selected to join the sample population, because there are some people living in the city that don’t have a fixed address or phone number.
Now, if the people who don’t have a fixed address or phone number have a lower socioeconomic status than those that have a fixed address, then the average income in your study will be higher than the average income of everyone in Portland.
Another common example of sampling bias happens when researchers have the correct address or phone number, but they simply can’t reach that person.
For example, let’s say researchers have a list of names of people in Portland with lung cancer and a list of names of people in Portland without lung cancer, and they want to ask each person about their smoking status in the past ten years.
The researchers decide to call each person on the list between 5pm and 9pm on Wednesdays, since most people are home from work at that time.
But this list excludes people that work during the evenings, like people that work night-shifts like nurses and police officers or have to work multiple jobs to make ends meet.
To avoid sampling bias, researchers can make phone calls at different times during the day and on different days of the week.
In addition, researchers can try to use multiple modes of contact, like emailing or texting the individual or going to their home to see them in person.
Another type of selection bias is called non-response bias - and it’s particularly problematic in studies that take time or effort on the part of the participant - like a survey.
In general, younger people, females, white people, and people with higher education, and higher socioeconomic statuses are most likely to respond to a survey.
Oftentimes, people choose to not complete a survey because they think it will take up too much time or simply because they don’t like answering questions.
So for these types of studies, researchers sometimes use incentives like money or free food to motivate people to participate in an effort to ensure that the sample population accurately reflects the target population.
Now, let’s switch gears and talk about how selection bias can influence a study’s internal validity - or a study’s quality.
Ultimately, to draw conclusions from a study, the key is to make sure that the two groups - for example, the individuals with lung cancer and the individuals without lung cancer - have similar baseline characteristics to one another.
That way the only key difference is the exposure that we’re trying to study, the exposure to smoking. For example, let’s assume that researchers choose individuals with lung cancer from a retirement community, and choose individuals without lung cancer from a college campus.
If that happened, then it would be hard to know if the difference in lung cancer rates between the two groups is due to smoking cigarettes or because of the age difference the two groups of individuals, since most young people - even if they smoke - won’t develop lung cancer until much later in life.
One way to ensure that the two groups in a study are similar, is to recruit individuals from the same place, like recruiting all of your participants from the local police department.
One type of selection bias is called the healthy worker effect. It can happen when researchers recruit their exposed group - the individuals that smoke - from an occupational cohort, like a group of police officers, but they recruit the non-exposed group - the individuals that don’t smoke - from the general population, like if they went door-to-door asking people to join the study.
This type of selection bias is called the healthy worker effect, because people who are less healthy are less likely to be employed.
As a result, the smoking group will include only people with good health - meaning they’ll be overall less likely to develop lung cancer, regardless of smoking status - and the non-smoking group will likely include people with both good and poor health - meaning they’ll be more likely to develop lung cancer, regardless of smoking status; and the overall effect of smoking on lung cancer will be underestimated.
Another type of selection bias is called Berkson’s bias - named after Joseph Berkson, a statistician. Berkson bias happens when individuals in both two groups of a study are recruited based on a characteristic that’s associated with the exposure, and it usually happens in hospital settings.
For example, let’s say we want to look at a group of people with lung cancer - the cases - and a group of people without lung cancer - the controls - and compare how many people smoke in each group.
For our cases, we get a list of 100 people with lung cancer from the local hospital and find out that 90 of them smoke. For our controls, we get a list of 100 people without lung cancer from that same hospital, let’s say it’s a list of individuals with coronary heart disease, and find out that 70 of them smoke.
Using the odds ratio equation, OR = (A*D)/(B*C), it looks like the odds of lung cancer for people who smoke is about 4 times higher than the odds of lung cancer for people who don’t smoke.
The problem here is that smoking is also a risk factor of heart disease, so there is probably a higher percentage of people who smoke in our group of controls compared to a group of controls who were admitted to the hospital for something like a broken bone, since smoking cigarettes isn’t a risk factor for breaking a bone.
So let’s say that instead we use a list of individuals with broken bones to recruit 100 individuals without lung cancer as our controls, and find out that only 20 of them smoke.
Now, the odds of lung cancer for people who smoke is around 36 times higher than the odds of lung cancer for people who don’t smoke.
Since broken bones doesn’t have a known relationship with smoking, this reflects the real odds ratio. So, Berkson’s bias can be avoided by choosing controls that have an outcome - like broken bones - that isn’t associated with the exposure - smoking.
Another type of selection bias occurs when people drop out of the study before the study period is over, and this is called loss to follow-up.
Loss to follow-up can happen when participants die during the course of the study for a reason that’s unrelated to the exposure or the outcome, or drop out of the study because they lose interest, or simply because they move away from a study area without letting the researchers know.
Loss to follow-up can be a problem, especially if the people who are lost to follow-up are the ones most likely to develop the outcome.
For example, out of 200 people - 100 that smoke and 100 that don’t smoke - let’s say that 80 of them develop lung cancer - 60 in the group that smokes and 20 in the group that doesn’t smoke.
But, out of the 60 people that developed lung cancer in the group that smokes, 40 of them are lost to follow-up sometime during the study period, while none of the people in the group that doesn’t smoke are lost to follow-up.
At the end of the study, it now looks like there are 20 people that developed lung cancer in the group that smoked and 20 people that developed lung cancer in the group that didn’t smoke, so it appears that smoking has no effect on lung cancer.
In this example, loss to follow-up underestimates the overall effect of the exposure on the outcome. To help reduce loss to follow-up, researchers often try to minimize the amount of time between contacting participants in their study, like sending them a survey every 2 years instead of every 5 years.
If participants don’t respond to the initial survey, researchers might also try to contact individuals by phone or email, or by getting in touch with a partner or good friend of the participant.
Alright, as a quick recap, selection bias happens when errors in the selection of participants decreases the study’s external validity - how similar the sample population is to the target population - or internal validity - how similar the two groups in the study are to each other.
Sampling bias and non-response bias can decrease the external validity of a study, but can be avoided by using multiple modes of contact and offering incentives to study participants.
The healthy worker effect, Berkson’s bias, and loss to follow-up can all lower a study’s internal validity, but can be avoided by recruiting all participants from the same place, making sure that recruitment isn’t linked to the exposure, and by keeping regular contact with participants or contacting participants in multiple ways.