Introduction to biostatistics
Definitions & Key takeaways
Biostatistics refers to the process of collecting, organizing, and analyzing variables collected from living things. Biostatistics involves design studies to answer specific scientific questions, and the skills necessary to properly analyze the data collected from those studies. It also involves effective communication of the results of analyses to scientists and other non-statisticians.
Let’s say you want to figure out if people with high body mass index, or BMI, are at a higher risk of hypertension - or high blood pressure.
Let’s say that you decide to go out and find 100 people with hypertension and 100 people without hypertension and find out the BMI of each person in each group.
You might also collect other information about the individuals in each group, like how old they are, if they smoke cigarettes, or if they drink alcohol, since all of these factors can influence a person’s risk of hypertension.
All of these different pieces of information - called variables - can be put together into a single document or file, called a data set.
A data set usually includes independent variables which are thought to influence or change dependent variables. In our example, the body mass index would be the independent variable and hypertension would be the dependent variable.
The process of collecting, organizing, and analyzing variables in a data set is called statistics, and when the data were collected from living things - like humans, aardvarks, algae, or bacteria - it’s called biostatistics, bio meaning life.
Now, there are two main types of biostatistics. The first type is descriptive statistics, which is used to describe or summarize information about each individual variable in the data set.
Descriptive statistics can be used to find the mean - the average number calculated from a particular variable, the median - the middle number in a variable, and the mode - the number that occurs the most in the variable.
The descriptive statistics of each variable can be calculated for the whole sample - all 200 people - or in each group separately - the 100 people in the group with hypertension or the other 100 people in the group without hypertension.
For example, we might find that the mean body mass index of all people in the study is 24.5, or that the mean body mass index is 28 for the group with hypertension and 21 for the group without hypertension.
We can also use descriptive statistics to find the range, variance, or standard deviation, all of which are ways of understanding how the data are spread out or distributed for a given variable.
For example, we might find that the lowest measured body mass index in the group with hypertension is 23, and the highest is 33, so the range for body mass index in this group is 23 to 33.
Typically, descriptive statistics are reported in a graph or a table. The second type of biostatistics is inferential, which is different from descriptive statistics in two ways.
First, inferential statistics looks at relationships between two or more variables, instead of looking at each individual variable.
For example, we could use inferential statistics to explore the relationship between body mass index and hypertension. We could categorize body mass index into two groups - above 25, or high, and below 25, or low - and we might find that people with high body mass indices have 3 times the odds of hypertension compared to people with low body mass indices.
Typically, inferential statistics are reported by relative risks, attributable risks, odds ratios, or hazard ratios. The goal of descriptive statistics is to describe how similar or different the study groups in a particular sample population are to one another.
For example, let’s say we use descriptive statistics to find that 72% of people in the group with hypertension are male, but only 16% of people in the group without hypertension are male.
This is important finding because men tend to have slightly lower body mass indices than women. As a result, having more men in the group with hypertension, means that the average body mass index in that group will be lower.
Ultimately, if the descriptive statistics find that the study groups are not very similar, we say that the study has low internal validity, and that the results found by inferential statistics may be the result of difference in the two study groups.
On the other hand, the goal of inferential statistics is to apply the results of the sample population to a target population - which is usually just the general population.
So, inferential statistics is concerned about whether or not the two study groups are similar, as well as whether or not the sample population represents the target population.
Ideally, a study should be done on a sample population of individuals that is similar to that target population in every meaningful way.
For example, if your target population is people from Lagos, Nigeria, then ideally your sample population would include people of ages, races, and socioeconomic statuses that reflect the characteristics of people in Lagos.
One tool to make sure the sample population represents the target population is called randomization, meaning that individuals get selected to enter the study through a process of chance.
Let’s say you have a list of everyone in Lagos who has hypertension and a list of everyone in Lagos who doesn’t have hypertension, and you close your eyes and choose 100 names from each list to include in the study.
That’s randomization. Using randomization, there’s a pretty high chance that the sample population and that target population will be similar, and we would say that study has high external validity.
In studies with high external validity, any results from inferential statistics that are made about the sample population can also be applied to the target population.
So, if we find that the odds of hypertension are higher for people with high body mass indices in the sample of 100 people from Lagos, we can conclude that the odds of hypertension are most likely higher for everyone a high body mass index in Lagos.
Typically, inferential statistics start with two hypotheses. The first hypothesis is called the null hypothesis, and it basically says there’s no difference between two variables.
For example, our null hypothesis would state that there’s no difference in the odds of hypertension for people with high body mass indices compared to people with low body mass indices, or simply, that there’s no relationship between hypertension and body mass index.
On the other hand, the alternate hypothesis would state that there is a difference in the odds of hypertension for people with high body mass indices compared to people with low body mass indices, or simply, there is a relationship between hypertension and body mass index.
To test these hypotheses, we generally use a linear or logistic regression model to see if there’s a significant difference between the two groups.
In common terms, “significant” often means important or interesting, but in statistics it means that the relationship between two variables is caused by something other than random chance, and it’s normally defined by a p-value of less than 0.05 or 5%.
For example, let’s say that we use logistic regression to see if there’s a difference in the odds of hypertension for people with high or low body mass indices, and we get an odds ratio of 3, and a p-value of 0.02.
This means that, if the null hypothesis is true, then the probability of getting an odds ratio of 3 - or higher than 3 - simply by chance, is about 2%.
In other words, there’s a very small probability - below 5% - that we would’ve gotten an odds ratio of 3 if the null hypothesis is true!
And because the probability is less than 5%, we can conclude that, most likely, the null hypothesis is false and the alternate hypothesis is true.
And the alternative hypothesis is that there really is a significant difference in the odds of hypertension between the two groups.
Statistical significance is different than clinical significance, which refers to the practical importance of the results.
In other words, a clinically significant result would indicate that the independent variable has a noticeable effect on the dependent variable in the daily life of the participant.
For example, let’s say we found that people with high body mass indices have 2.5 times the odds of hypertension compared to people with low body mass indices, but the p-value was 0.22, which is much higher than 0.05.
In the study, we might conclude that body mass index doesn’t have an statistically significant effect on hypertension. But, in reality, a health professional might conclude that 2.5 times is clinically significant, and they might inform individuals with high body mass indices that they are more likely to have hypertension.
It’s possible to have any combination of statistical and clinical significance - so it’s possible for study results to be statistically significant but not clinically significant, for both statistical and clinical results to be significant, or for both results to not be significant.
Also, there’s no set level of clinical significance - it ultimately depends on the how two variables relate to one another.
Alright, as a quick recap. Descriptive statistics is used to describe the characteristics of each variable within a specific sample, and inferential statistics is used to describe the relationships between two or more variables.
Typically, we use the results of descriptive statistics to make sure the groups within the study are similar to each other, and that the study has high internal validity.
We use inferential statistics to make conclusions about a target population from the sample population, but only if the study has high internal and external validity, meaning the sample population has similar characteristics to the target population.
In inferential statistics, the null and alternate hypotheses are evaluated based on a set p-value, while clinical significance is determined based on how an individual evaluates the relationship between two variables.
No notes for this video yet
Try adding a note below