Definitions & Key takeaways

A cohort study is a study that helps to determine a relationship between an exposure and a future outcome. In cohort studies, a group of people with a specific characteristic, such as exposure to a particular substance, are followed over time to see if they develop a specific disease or health outcome.

Cohort studies can be either prospective or retrospective. In prospective cohort studies, also known as concurrent cohort studies, individuals are followed forward in time, and the number of people who develop a particular outcome gets compared between the two groups. Next, we have retrospective cohort studies, also called historical or non-concurrent cohort studies. In retrospective cohort studies, two groups of individuals are selected in the past and followed up until the present day. Comparing two groups determines the number of individuals in each group who develop a particular outcome.

A group of people who share a common characteristic is called a cohort. For example, people born in the year 1981 make up a birth cohort, and people who work in construction make up an occupational cohort.
Now, cohort studies or longitudinal studies are a type of study design that follows a cohort of people over time to figure out if there’s an association between an exposure and an outcome.
Typically, cohort studies look at individuals in a cohort who have a certain exposure, as well as individuals in a cohort who have not had that exposure, to compare their rates of a certain outcome in the future.
For example, let’s say we want to figure out if there’s a relationship between smoking cigarettes and developing lung cancer.
To do this, we could follow 100,000 individuals that smoke cigarettes, the exposed group, and 100,000 individuals that don’t smoke cigarettes, the non-exposed group, for ten years.
After ten years, let’s say that 82 of the 100,000 people - 0.082% - who smoked developed lung cancer, and only 3 of the 100,000 people - 0.003% - who didn’t smoke developed lung cancer.
We can then compare the groups by dividing the probability of lung cancer for people who smoked - 0.00082 - by the probability of lung cancer for people who didn’t smoke - 0.00003 - and determine that people that smoked had 27 times the risk of developing lung cancer during that ten period.
As it turns out, smoking is the number one risk factor of most types of lung cancer, and people who smoke are 15 to 30 times more likely to develop lung cancer than people who don’t smoke.
Now, there are two main types of cohort studies. The first type is called prospective cohort or concurrent cohort, because individuals are followed forward in time.
An example would be if in 2018 a group of smokers and a group of non-smokers are recruited for the study. Then the two groups are followed for ten years, until 2028, and the number of people who develop lung cancer are compared between the two groups.
Prospective cohort studies are the most common type of cohort study, and are what people usually think of when they hear about a “cohort study”.
The other type of cohort study is a retrospective cohort study, also called historical or non-concurrent cohort study. In retrospective cohort studies, participants are recruited in the past and then followed until the present day.
This can be done either with time-travel or by simply by looking at medical records from the past. For example, you might look at medical records to find a cohort of young adults in 2008 and then divide them into a group of people that smoke cigarettes, the exposed group, and a group of people that don’t smoke cigarettes, the non-exposed group.
Then we can follow the medical records of these individuals over the next ten years, until 2018, and compare the number of individuals in each group who develop lung cancer.
So even though retrospective cohort studies use data from the past, they still have the same basic structure as prospective cohort studies.
In other words, they still start with a cohort of exposed and non-exposed individuals and follow them over time to assess their outcomes.
Now, this is different from case-control studies, which start with individuals who do or don’t have the outcome, and then use data to investigate past exposures.
This can get a bit confusing because case-control studies are sometimes called retrospective studies, and cohort studies can be either prospective cohort or retrospective cohort studies.
One advantage of cohort studies is that they’re able to clearly show the timing or temporality of the relationship between the exposure and the outcome.
For example, out of 200 people, 100 that smoke and 100 that don’t smoke, none of them have lung cancer at the beginning of the study.
But after ten years, more individuals that smoke develop lung cancer compared to individuals that don’t smoke, so it’s pretty clear that smoking happened first and lung cancer happened second.
Because of this, cohort studies are usually preferred over other study designs, like cross-sectional studies, which measure the outcome and the exposure at the exact same time.
For example, a cross-sectional study might start with 100 people who have lung cancer, and 100 people who don’t have lung cancer, and interview all of those people to see how often they currently smoke cigarettes.
The problem here is that people who are diagnosed with lung cancer generally tend to stop smoking, so there might actually be fewer people who currently smoke cigarettes in the lung cancer group compared to the group without lung cancer.
This is an example of reverse causation, where it looks like people who smoke are less likely to get lung cancer, even though the opposite is true.
Cohort studies are also useful in situations that may be unethical to test using other approaches. For example, a cohort study could be used to find out if people that smoke get lung cancer in the future.
But doing a randomized control trial, where individuals are assigned to smoking or not smoking, might be unethical, especially if we think that smoking causes lung cancer.
It’s one thing for someone to choose a harmful exposure for themselves, but it’s a different thing to choose a harmful exposure for someone else - by assigning them to the smoking group.
Cohort studies are also good for figuring out how one exposure might impact many different outcomes, like how smoking might impact lung cancer, asthma, and heart disease.
Cohort studies are useful for looking at common exposures, like smoking cigarettes, but they’re particularly useful for looking at rare exposures, like coal miners who inhale coal dust.
To figure out whether coal dust causes lung disease, we might follow 100 coal miners, the exposed group, and 100 construction workers, the non-exposed group, for ten years and then compare their rates of lung disease.
In this situation, since coal dust exposure is rare, it’s makes more sense to recruit 100 coal-miners into the study, than to start with 100 individuals with lung disease and to try to find out if any had this rare exposure to coal dust.
Because it’s such a rare, rare exposure, there might be only 1 or 2 people out of the 100 individuals with lung disease who were exposed to coal dust.
Now, there are some downsides to cohort studies. Prospective cohort studies often take a lot of time and money, and oftentimes people drop out of the study before it’s over, called loss to follow-up.
Loss to follow-up introduces a type of selection bias, and that can worsen a study’s internal validity - or a study’s quality.
Loss to follow-up can happen when participants die during the course of the study for a reason that’s unrelated to the exposure or the outcome, or drop out of the study because they lose interest, or simply because they move away from a study area without letting the researchers know.
Loss to follow-up can be a problem, especially if the people who are lost to follow-up are the ones most likely to develop the outcome.
For example, out of 200 people - 100 that smoke and 100 that don’t smoke - let’s say that 80 of them develop lung cancer - 60 in the group that smokes and 20 in the group that doesn’t smoke.
But, out of the 60 people that developed lung cancer in the group that smokes, 40 of them are lost to follow-up sometime during the study period, while none of the people in the group that doesn’t smoke are lost to follow-up.
At the end of the study, it now looks like there are 20 people that developed lung cancer in the group that smoked and 20 people that developed lung cancer in the group that didn’t smoke, so it appears that smoking has no effect on lung cancer.
In this example, loss to follow-up underestimates the overall effect of the exposure on the outcome. To help reduce loss to follow-up, researchers often try to minimize the amount of time between contacting participants in their study, like sending them a survey every 2 years instead of every 5 years.
If participants don’t respond to the initial survey, researchers might also try to contact individuals by phone or email, or by getting in touch with a a partner or good friend of the participant.
Alright, as a quick recap, cohort studies are a type of study design used to determine if there’s a relationship between an exposure and a future outcome.
Prospective cohort studies start with a group of individuals in present day and follow them into the future, while retrospective cohort studies start with a group of individuals from the past and follow them to present day.
Both types of cohort studies are useful for determining temporality between an exposure and an outcome, can be used to investigate multiple outcomes for a single exposure, and are sometimes the only ethical way to an association between an exposure and an outcome.
However, cohort studies are expensive and time-consuming, and can be affected by loss to follow-up. Loss to follow-up can be reduced by keeping regular contact with participants and by contacting participants in multiple ways.