Study designs
Definitions & Key takeaways
There are six basic types of epidemiological study designs, and they can each be distinguished using certain criteria. The first criterion for deciding which study design to use is whether you have individual or group data.
So, let’s say that 9 people had migraines. If we have individual data, we can look at the individual characteristics for each of the 9 people that had migraines, like their sex, age, race, or past history of migraines, and we can compare them to the people that didn’t have migraines.
On the other hand, if we have group data, we don’t actually know which specific individuals out of the 100 people had migraines.
So even though we know that 9 people had them, we don’t know which 9 people they were or any of their individual characteristics.
Now, ecological studies are a type of study design that uses group data to figure out if there is a potential association between two variables.
For example, let’s say you want to figure out if people who sleep less are more likely to get migraines. And perhaps you have information about average sleep duration for populations in ten different cities.
You could plot this information on a graph with average sleep duration on the x-axis and the prevalence of migraines—which is the number of people that suffer from migraines, per 100,000 people—on the y-axis.
Generally, we can see that the less sleep a city gets, the higher the prevalence of migraines is for that city. The thing is, we can’t actually say that getting less sleep causes migraines, since we don’t have information about each individual in each city.
All we can say is that there’s an association between sleep duration and prevalence of migraines. Ecological studies are helpful for making hypotheses, though, that can later be tested using individual-level studies.
And, in general, individual-level studies are considered stronger than ecological studies, because knowing individual characteristics can help us determine what risk factors are associated with certain diseases.
So now let’s talk about studies that use individual data. The next criterion we use to decide on a study design is whether or not there’s an intervention, and an intervention is basically just an exposure that the researcher controls.
Studies with interventions are also called experimental studies, randomized controlled trials, or RCTs for short. So, for example, let’s say we want to find out if a newly discovered drug, we’ll call it Drug A, can prevent migraines for up to a year.
In this example, Drug A is the intervention and having a migraine is the outcome. In the most basic RCT, the sample population might be randomly split into two treatment groups, an intervention group that receives Drug A, and a control group that receives a placebo.
The placebo looks and tastes like Drug A but is completely harmless and ineffective - like a tiny capsule filled with water.
After both groups get their treatments, researchers would compare the incidence of migraines in each group—which is the number of individuals in each group who got migraines over the next year.
Determining causality is possible because the intervention group and the control group are randomly selected from the larger target population, so there’s a good chance that people in each group are similar and that the only difference between the two groups is whether or not they were exposed to Drug A.
There are some downsides to RCTs though, mainly that they can sometimes be really expensive, time consuming, and in some cases unethical, depending on the intervention.
Next, let’s talk about studies that don’t have an intervention, and these are called observational studies, because you simply observe what happens to individuals without controlling their exposure.
There are a few different types of observational studies, and the main criterion used to distinguish them is when you measure the exposure.
In other words, whether you measure the exposure before the outcome, after the outcome, or at the same time. The first observational study are cohort or longitudinal studies, and cohort studies measure the exposure before the outcome.
Now, a cohort is simply a group of people who share a common characteristic. So, cohort studies are a type of study design that look at individuals in a cohort who have a certain exposure, as well as individuals in a cohort who have not had that exposure, and then follow both groups over time and compare the incidence of a certain outcome.
For example, let’s say we follow a group of 100 individuals that smoke cigarettes, the exposed group, and 100 people that don’t smoke cigarettes, the unexposed group, and compare the incidence of migraines during the next five years.
Cohort studies are useful when you want to show the timing or temporality of the relationship between the exposure and the outcome.
Cohort studies are also good for looking at rare exposures, like if a certain uncommon medication causes an increased risk of migraines.
It makes more sense to recruit 100 people who were all already using the uncommon medication rather than start with 100 people with migraines and try to figure out if any of them had been exposed to the medication in the past; because since it’s such a rare exposure, there might be only 1 or 2 people out of the 100 individuals with migraines who were exposed to the medication.
A third reason you’d use cohort studies is to look at multiple outcomes from one specific exposure. For example, if we follow 100 people using the uncommon medication, we could also compare the incidence of other side effects, like constipation or nausea, in addition to migraines.
Now, cohort studies are considered to be the strongest type of study besides RCTs, but they have some downsides too. Like RCTs, they can be expensive and time-consuming, and they can also have loss to follow-up, which is when people drop out of the study before it’s over.
The next type of observational study design are case-control studies, which measure the exposure after the outcome is measured.
Case-control studies compare the history of two groups of people—those that have a certain outcome, and those are called the cases, and those that don’t have a certain outcome, which are called the controls—to see if they’ve been exposed to different things.
For example, let’s say we find a group of 100 people with migraines, which are the cases, and 100 people without migraines, which are the controls, and then we compare how many of those people used a certain medication in the past five years.
Now, it’s important to note that you can’t assess incidence in case-control studies, because you’re collecting information on people who already have the outcome of interest.
Case-control studies are particularly helpful when the outcome of interest is rare, because it’s easier to recruit individuals with rare diseases into a study, rather than start with a group of healthy individuals and wait to see who develops the rare disease in the future.
Now, case-control and cohort studies have similar but opposite characteristics, so let’s clarify to keep them straight: Case-control studies are good for rare outcomes, while cohort studies are good for rare exposures.
And case-control studies are also good for looking at multiple exposures for a single outcome, while cohort studies are good for looking at multiple outcomes to a single exposure.
These two things make sense because case-control studies start with the outcome and look back in time to find exposures, while cohort studies start with the exposure and look forward in time to find outcomes.
Now, case-control studies are often preferred when money and time are low, because they’re relatively inexpensive and less time-consuming than cohort studies and RCTs.
But, on the other hand, they’re not able to show temporality or incidence, and they’re also particularly subject to recall bias, which is when people either overestimate or underestimate their past exposure.
Okay, next we have cross-sectional studies, which is a type of study design where an exposure and an outcome are measured at the same time.
For example, let’s say you want to figure out if people who have migraines smoke cigarettes more often than people that don’t have migraines.
To do this, you might look at the medical records of 200 people to measure the outcome and see who has migraines and who doesn’t, and at the same time measure the exposure and see how many people in each group also smoke cigarettes.
You can think about a cross-sectional study like a snapshot of the population at a certain point in time. Since you can only collect the information you see in that one moment, you don’t know what happens before or after the snapshot was taken.
So, we can’t collect any information on incidence, or new cases, and we can only collect information on prevalence, which is the proportion of exposures or outcomes that already exist at a certain time.
Cross-sectional studies are often a cheap, quick, and easy way to collect information about a large number of participants, since all of the information is collected at one time, and typically they are done through surveys.
Also, a lot of information can be collected from each participant, so cross-sectional studies are especially useful for looking at relationships between multiple diseases and multiple outcomes.
For example, in addition to looking at migraines, you could also compare the prevalence of strokes and heart disease in people that smoke cigarettes or don’t.
One downside of cross-sectional studies though, is that you can’t determine causality since you only have information about prevalence and not about incidence.
In other words, since you don’t know anything about what happened to the individuals before they joined the study, it’s impossible to tell if the exposure or the outcome happened first.
Because of this, cross-sectional studies are considered to be weaker than both cohort and case-control studies. Another type of study design is called a case series or clinical series, and these are descriptive studies, meaning they don’t have a comparison group.
This makes them different from other types of observational studies. In a case series, all individuals in the study must have the same particular disease or condition.
For example, let’s say that five 30-year-old men are admitted into a health clinic, and each of them are diagnosed with dementia, which is a disease that most commonly affects older individuals.
A researcher might include these five men in a case series, where they would report information about the men’s characteristics, like their age, occupation, and where they live.
The information from this case series can help give some idea for how the men developed dementia at an early age, and can help generate hypotheses for future research.
But, it’s important to note that information from case series studies can’t be used to determine causality, and they’re generally considered the weakest type of observational study design.
Alright, as a quick recap, there are six basic study designs, and they can each be distinguished using certain criteria.
First, if you only have group level data, you must use an ecological study design. If you have individual level data, you have to decide if the study will include an intervention or if it will be an observation study.
Studies with interventions are generally called randomized controlled trials, or RCTs, and these are considered the gold standard study design.
There are three types of observational studies that have a comparison group. Cohort studies measure the exposure before the outcome, case-control studies measure the exposure after the outcome, and cross-sectional studies measure the exposure and the outcome at the same time.
Observational studies that don’t have a comparison group are called a case series. Generally, we think of cohort studies as the strongest observational study design and case series as the weakest study design.
No notes for this video yet
Try adding a note below