Definitions & Key takeaways

In testing and measurement, accuracy and precision are two important concepts of the quality of the test results. Accuracy refers to how close the measured value is to the true value. In other words, it reflects the degree to which a test result is correct or exact. A test can be accurate if it uses a tool with high validity.

Precision, on the other hand, refers to the consistency or reproducibility of the results obtained from a test. It reflects the degree of variation or uncertainty in the results. A precise test uses tools with high reliability.

So, a tool with high validity will get results that are close to the true value, and a tool with high reliability will get results that are consistent no matter how many times the measurement is repeated.

Let’s say you want to figure out if eating more daily servings of vegetables will decrease a person’s body mass index (BMI), which is a number calculated by dividing a person’s weight in kilograms by their height in meters squared.
The first step to figuring this out is to collect data about each person in the study, and this is typically done using some type of measurement tool.
For example, we might use a scale to measure a person’s weight, a measuring rod to measure a person’s height, and design a survey to find out how many daily servings of vegetables a person eats.
Now, it’s important to collect high quality data in a study, which means the information collected in the study should accurately reflect what’s really happening.
For example, if a person eats 5 servings of vegetables per day, the data should reflect that they eat 5 servings, instead of 2 servings.
Data quality is determined by the tools used to collect the information, and ideally, these tools have high validity - or accuracy - and high reliability - or repeatability.
A tool with high validity will provide a measurement that’s very close to the true or known value for the thing being measured.
Let’s say we’re going to measure a woman’s weight using two different scales. One scale is a family heirloom that was passed down over multiple generations - so it’s pretty old - and the other scale was a gift from your friend who’s a doctor - so it’s really modern and sophisticated.
The old scale provides a measurement of 80 kilograms, and the modern scale provides a very different measurement of 66 kilograms.
In reality, this woman weighs 65 kilograms, so, since the modern scale provides a measurement that is closer to the woman’s true weight, the modern scale has higher validity.
Using tools with high validity is important for getting correct results in descriptive or inferential statistics. For example, if we used the old scale for all the people in the group with hypertension, but used the new scale for the people in the group without hypertension, then we would think the group with hypertension has a much higher mean body mass index than they really do.
This would lead to an overestimation of the association between body mass index and hypertension. On the other hand, a tool with high reliability will consistently get the same results, no matter how many times the measurement is repeated.
So, let’s say you measure each person’s weight 3 times in a row on each scale. On the old scale, the 3 measurements are 80 kilograms, 81 kilograms, and 80 kilograms, and on the modern scale, the 3 measurements are 66 kilograms, 75 kilograms, and 60 kilograms.
Now, even though the modern scale has higher validity, it actually has lower reliability, because the results of the 3 tests were not consistent with each other.
And, by comparison the old scale has low validity, but it actually has higher reliability, because each of the 3 measurements was very close to one another.
Using tools with high reliability is important because it ensures that any studies repeated at different times will have the same results.
For example, let’s say a study in Kingston and a study in Mexico City both use the same scale to figure out if body mass index is related to hypertension.
If this scale has low reliability, then the study in Kingston might find that higher body mass index increases the risk of hypertension, while the study in Mexico City might find that higher body mass index decreases the risk of hypertension.
In this situation, it’s impossible to tell how body mass index affects hypertension, if at all. It’s possible to have measurement tools that have both high validity and reliability, and this can be represented simply on a histogram - a plot that shows the frequency of certain values - with weight in kilograms on the x-axis and the number of measurements on the y-axis.
Let’s say we measure 100 people, and the mean weight in the population is 65 kilograms. Now, that doesn’t mean that everyone has a weight of 65 kilograms, but the majority of people weigh close to the mean of 65 kilograms, with fewer and fewer people that have extremely high weights or extremely low weights.
This is called a normal distribution, and it creates a curve in the shape of a bell. Now, a measurement tool that has high validity and reliability will create a distribution curve that looks exactly like the true distribution curve.
But if the tool has low validity and high reliability, the curve would be shifted away from the true mean, because the results are not accurate, but the curve might be the same width or even more narrow, because the results were repeatable.
On the other hand, if the tool has high validity and low reliability, the curve would be very close to the true mean, because it is accurate, but the shape would be more flat and spread out, because the results would not be repeatable.
Finally, if the tool has low validity and low reliability, then the curve will be shifted away from the true mean and flattened out, which could definitely provide some unexpected results!
And in that situation, it’s probably time to throw out the tool. Alright, as a quick recap, a measurement tool has two characteristics - validity, or accuracy, and reliability, or repeatability.
A tool with high validity will get results that are close to the true value, and a tool with high reliability will get results that are consistent no matter how many times the measurement is repeated.
Tools with high validity and reliability will get results that are similar to the normal distribution curve, but tools with low validity will shift the curve away from the true mean, and tool with low reliability will flatten the curve.