Definitions & Key takeaways

Positive predictive value (PPV) is a measure of the accuracy of a positive test result in a diagnostic test, calculated by dividing the number of true positives by the sum of true positives and false positives. It is used to quantify the likelihood that a person with a positive test result actually has the condition the test is designed to detect.

On the other hand, there is negative predictive value (NPV), which is a measure of the accuracy of a negative test result in a diagnostic test, calculated by dividing the number of true negatives by the sum of true negatives and false negatives. It is used to quantify the likelihood that a person with a negative test result does not have the condition the test is designed to detect.

Imagine that a person gets the results of a colon cancer screening test. There are two possible scenarios - either the result is positive, indicating that they have colon cancer, or the result is negative, indicating they don’t have colon cancer.
At this point, the person may ask themselves, how worried should I be that it was a positive test result? Or, how reassured should I be that it was a negative test result?
Each test has a positive predictive value, or PPV, which is the probability that people with a positive test result truly have the outcome, and a negative predictive value, or NPV, which is the probability that people with a negative test result truly don’t have the outcome.
Let’s take an example to show how it’s possible to measure a test’s predictive value. Let’s say that we recruit 1000 people - 100 people with colon cancer and 900 people without colon cancer - and then we give them all the same screening test.
That way we can see how many people with positive results actually have colon cancer and how many people with negative results actually don’t have colon cancer.
We can organize the results using a 2 by 2 table, where the true disease status of the individual is on the top of the box, and the results of the screening test are on the side, and each of the cells is labeled a, b, c, or d.
A true positive would be a person who gets a positive test result and has colon cancer. A true negative would be a person who gets a negative test result and doesn’t have colon cancer.
A false positive would be a person who gets a positive test result even though they don’t have colon cancer. And a false negative would be a person who gets a negative test result even though they have colon cancer.
To calculate the positive predictive value, we divide the number of true positives by the total number of people who tested positive - so cell a divided by the sum of cell a and b.
A test with a perfect positive predictive value would have 100 true positives in cell a, because the test would correctly identify everyone who has colon cancer, and zero false positives in cell b.
To calculate negative predictive value, we divide the number of true negatives by the total number of people who tested negative - so cell d divided by the sum of cell c and d.
A test with perfect specificity would have 900 true negatives in cell d, because the test would correctly identify everyone who doesn’t have colon cancer, and zero false negatives, in cell c.
But no test is 100% perfect, so let’s say that cell a contains 90 true positives, cell b contains 50 false positives, cell c contains 30 false negatives, and cell d contains 850 true negatives.
In this situation, the positive predictive value would be 64%, because there are 90 people who are true positives - in cell a - and 140 people who tested positively - cell a plus cell b.
In other words, 64% of people who test positively will actually have colon cancer, while the other 36% of people who test positively will not have colon cancer.
The negative predictive value would be 97%, because there are 850 people - in cell d - who are true negatives and 880 people who tested negatively- cell b plus cell d.
In other words, 97% of people who test negatively will actually not have colon cancer, while the other 3% of people who test negatively will have colon cancer.
Predictive values are commonly confused with sensitivity and specificity. A test with high sensitivity will correctly identify most people who have the condition, and a test with high specificity will correctly identify most people who don’t have the outcome.
The main difference between validity and predictive value is that sensitivity and specificity are fixed characteristics of a test.
On the other hand, predictive values are affected by the prevalence of the outcome. To understand this better, let’s follow another example.
So, imagine we now want to test for colon cancer in a group of 10,000 teenagers. Colon cancer is rare for teenagers, so let’s say the prevalence is only about 1%, in reality, it would be far lower, but this makes the numbers easier to follow.
And let’s also say that the test we use to check for colon cancer has a sensitivity of 99%, meaning it correctly identifies 99% of people who have colon cancer, and a specificity of 95%, so it correctly identifies 95% of people who don’t have colon cancer.
We can put the results of the test in a 2 by 2 table, where 99 people are true positives, 495 people are false positives, 1 person is a false negative, and 9405 people are true negatives.
In this situation, the positive predictive value would be 99 divided by 594, or 17%, and the negative predictive value would be 9405 divided by 9406, or 99%.
In other words, when the prevalence is low, the positive predictive value is low. And that makes sense, because when a disease is rare, it becomes more and more likely that someone that tests positive is actually just a false positive.
On the flip side, let’s think about a scenario where the prevalence of colon cancer is higher, like in a group of 10,000 people over the age of 75, where the prevalence of colon cancer is closer to 5%.
If we had the same sensitivity - 99% - and the same specificity - 95% - we would have 495 true positives, 475 false positives, 5 false negatives, and 9025 true negatives.
The positive predictive value would be 495 divided by 970, or 51%, and the negative predictive value would be 9025 divided by 9030, or 99%.
So, when the prevalence is higher, the positive predictive value is higher. One thing to notice is that the negative predictive value didn’t change much when the prevalence was low or high, and this is because there are a large number of negative tests in a population, even in common conditions.
So, prevalence mainly affects the positive predictive value, and the positive predictive value of a screening test is highest for common diseases.
And, when the prevalence of the outcome is low, the test’s specificity has a big impact on the positive predictive value as well.
For example, let’s say we have a population of 10,000 people where the prevalence of colon cancer is 10%. Keep in mind though that this number isn’t very realistic, since no populations really have a prevalence that high.
This time, we can draw the 2 by 2 table with each cell proportional to the number of people it represents. So, let’s say we’re using a test with 50% sensitivity and 50% specificity, and there are 500 true positives, 4500 false positives, 500 false negatives, and 4500 true negatives.
In this situation, the positive predictive value for this population is 500 divided by 5000, or 10%. So what happens if we increase the test’s specificity to 90% while keeping the sensitivity and the prevalence the same?
Now, there are 500 true positives, 900 false positives, 500 false negatives, and 8100 true negatives, which means the positive predictive value is 500 divided by 140, or 36%.
So, just by increasing the specificity, we increased the positive predictive value by 26%, which is a pretty big difference!
On the other hand, there’s not as big of a difference in the positive predictive value if the sensitivity of the test is increased to 90% while the specificity and prevalence stay the same.
In this situation, there are 900 true positives, 4500 false positives, 100 false negatives, and 4500 true negatives, so the positive predictive value is 900 divided by 5400, or 17% - this is only a 7% increase opposed to the 26% increase we saw when the specificity was increased.
The reason for the difference is that, with a rare outcome, most of the population is on the right side of the table - because most people don’t have the outcome.
So, if we change the numbers on the right side of the table - like when the specificity increases - it has a big impact on the positive predictive value.
On the other hand, there are very few individuals on the left side of the table, because not very many people have the outcome.
So when we change the numbers on the left by changing the sensitivity, it doesn’t have that much of an impact on the overall predictive value.
Alright, as a quick recap, the positive predictive value is the probability that people with a positive test result truly have the outcome, and the negative predictive value is the probability that people with a negative test result truly don’t have the outcome.
Positive predictive value is calculated by dividing the number of true positives by the number of all people who tested positive – both true positives and false positives.
Negative predictive value is calculated by dividing the number of true negatives by the number of all people who tested negative – both true negatives and false negatives.
The positive predictive value is higher when the prevalence of the outcome in the population is higher, but if the prevalence is low, then the positive predictive value increases when the specificity of the test increases.