A case-control study is an observational method used to compare a group of individuals with a particular condition (the cases) to another, a similar group of people without that condition (the controls). The investigation begins after researchers have identified a group of people with the condition they wish to study. A second, comparable group who does not have the condition is then identified from medical records or other sources. Investigators then look back through both groups' records to identify possible causes or risk factors for the condition. Case-control studies can be conducted relatively quickly and cheaply since researchers don't have to follow participants over time as they would in a cohort study. Another advantage is that case-control studies can be used to study rare conditions that would be challenging in a cohort study.
Case-control studies are a type of study design that compares the history of two groups of people - those that have a certain outcome, called cases, and those that don’t have a certain outcome, called controls; to see if they’ve been exposed to different things.
For example, a case-control study might find that the odds of using tanning beds is higher for people with skin cancer, the cases, compared to people without skin cancer, the controls.
Because case-control studies look at the past exposures of people with and without the outcome, this type of study is called retrospective, retro meaning past.
The opposite would be to start with people who either have the exposure or don’t have the exposure, and then follow them over time to see if they develop an outcome in the future.
That would be a prospective study. But, following people over time can take a lot of time and cost money.
By comparison, case-control studies are often quicker and cheaper, since they use data that’s already collected or is relatively easy to collect, like medical records, employment records, or individual interviews and surveys.
For example, researchers might interview a group individuals with skin cancer and a group of individuals without skin cancer to see how many times they used tanning beds in the past five years, then compare the results.
This is particularly true when it comes to rare diseases, but the definition of “rare” varies around the world. For example, in the United States, 1 in 1,500 people is considered rare, but in Japan, 1 in 2,500 people is considered rare.
Either way, it’s easier to recruit individuals with rare diseases into a study, rather than start with a group of healthy individuals and wait to see who develops a rare disease in the future.
For example, in the United States, you’d have to follow a group of 30,000 healthy individuals over time, to identify 20 people who develop a rare disease.
Case-control studies are also useful in situations that may be unethical to test using other approaches. For example, a case-control study could be used to find out if people with skin cancer happened to use tanning beds more than people without skin cancer.
It’s one thing to find out if someone was harmed in the past, but a different thing to cause them harm going forward. Case-control studies can also be used to explore many exposures associated with an outcome, which comes in handy for outbreak investigations.
For example, if there was an E.coli outbreak at a dinner party, researchers might interview everyone who got sick that night.
They could ask both cases who got sick and controls who didn’t, about what they ate. Because case-control studies use data from the past, they can ask about multiple exposures - the fish tacos, the lemon meringue pie, the falafel - all at once, to identify something that the cases might have eaten more of, then the controls to identify the likely source of E.coli.
However, there are some down-sides to case-control studies. In particular, since the study relies on historical exposures, individuals may have a recall bias, meaning that they may overestimate or underestimate the exposure.
For example, they might overestimate how often they go to the gym and underestimate how often they pick their nose. The recall bias really becomes a problem when there’s a long period of time between an exposure and the outcome, like when adults are asked to remember exposures that happened in their childhood.
Oftentimes, individuals that are in the case group remember exposures differently than those that are in the control group.
For example, individuals with skin cancer might say they used a tanning bed more times in the past five years than they actually did, since they know that tanning beds are a risk factor for skin cancer.
On the other hand, individuals without skin cancer might say they used a tanning bed fewer times in the past five years than they actually did, simply because each visit blurs into the next and can be forgotten.
In this example, the case group would have over reporting of tanning bed use, and the control group would have underreporting of tanning bed use, which will make it seem like tanning beds are more harmful than they actually are.
To help reduce recall bias, researchers often use written records, like medical records, to verify the information collected from individual interviews, or they can study exposures that happened in the recent past.
Ultimately, to draw conclusions from a case-control study, the key is to make sure that control groups has similar characteristics to the case group.
That way the only difference is the exposure that we’re trying to study, like the exposure of tanning bed use. For example, the researchers might choose the case group from a neighborhood where most people have light skin tones, and then recruit the control group from a neighborhood where most people have dark skin tones.
This could happen by mistake, if the researchers don’t realize there are demographic differences between the two neighborhoods.
So if that happened, it’s hard to know if the difference in skin cancer rates is due to tanning bed use or a biological difference between the two groups of individuals.
This is called selection bias, which happens when researchers choose two groups that aren’t similar enough, so your results don’t necessarily reflect the real relationship between the outcome and the exposure.
Selection bias can decrease a study’s internal validity – or a study’s quality, and it makes it more difficult to draw conclusions from the results.
One way to avoid selection bias is to use matching. Matching is when the cases and controls are matched based on a certain characteristic, like age, sex, race, socioeconomic status, or occupation.
It’s possible to match individually, like including a person with light skintone in the control group for every light skin tone person in the case group.
Alternatively you can match by frequency, so that the proportion of controls with a certain characteristic is identical to the proportion of cases with the same characteristic, like if we included 20 men and 30 women in the case group, we would include 20 men and 30 women in the control group.
Alright, as a quick recap, case-control studies are a type of retrospective study that’s used to determine if there’s a relationship between an outcome and a past exposure.
Case-control studies are quicker and cheaper than prospective study designs, can be used to investigate multiple exposures for a single outcomes, and are sometimes the only ethical way to an association between an exposure and an outcome.
Case-control studies can have recall bias and selection bias, but these can be reduced by verifying individual responses with written records, choosing recent exposures, and matching.