Types of data
Definitions & Key takeaways
Categorical data includes nominal data in which the order does not matter, such as the hair color of a certain population and ordinal data in which the order is important, such as estimating the degree of pain on a scale from one to ten. Numeric data includes interval data, such as the temperature and ratio data, such as the length of a particular group of students. Numeric data can be continuous, i.e., with intermediary values or discrete, i.e., without intermediary values.
Data are like a set of facts that are measured and recorded, and then summarized to help us make conclusions. Data can either be quantitative or in numbers, like a person’s age, measured in number of years they’ve been alive or it can be qualitative, which is non-numerical, like someone’s blood type {A, B, AB, or O}.
So, data are classified into two main groups - quantitative or numeric data and qualitative or categorical data. Let’s start with categorical data, which involves assigning subjects to a category, ie.
“red” vs “blue” or “high” vs. “low”.
Categorical data can be further broken down into nominal and ordinal data. Nominal data is based on categories that cannot be logically ordered.
For example, blood types - A, B, AB and O are nominal data; There is no logical order or magnitude in blood type. A is not higher than AB, and O is not less than B - they are just different, like apples and oranges.
Now you could say that type AB blood has more antigens than type O blood or that apples are firmer than oranges, but then we’re looking at different data - antigen number and firmness, and not simply the blood type or fruit type.
Other attributes like sex, type of religion, or ethnic background are all examples of nominal data. These attributes are measured in categories, instead of numbers.
Therefore, they don’t have any magnitude; that’s why when you summarize nominal data, you have to use proportions. For example, take a group of 20 classmates: 10 are Blood Type A’s, 5 are Blood Type B’s, and 5 are Blood Type O’s.
You can say that 50% are A, 25% are Type B, and 25% are Type O. And while you can’t calculate a mean or median, you can identify the “mode”, which is the most frequently appearing value of this data: Blood Type A.
Ordinal data are also measured in categories, but unlike nominal data, ordinal data come with a logical order attached. For example, let’s say you want to measure happiness, and you send 100 people a survey that asks: “How happy are you?” They can answer 1 of 4 answer choices: “1.
Sad”, “2. Not happy” “3.
Okay” and “4. Great!”.
Unlike the blood type example, there is a clear logical order here. In ordinal data, categories are ranked as being higher or lower than one another.
As we go from category 1 to category 4, happiness increases. And because there is an order to the data, you can calculate the “median”, which is the middle most value in a dataset arranged in order from the highest to the lowest values, and the “mode” but not the “mean” since we can’t clearly state whether the difference between category 1 and category 2 is quantitatively the same as the difference between category 3 and category 4.
One potential problem with ordinal data is that it can sometimes oversimplify relationships between categories. The jump from category 1 to category 2 may be very small, whereas the jump from category 3 to category 4 may be quite large.
Ordinal data are blind to this nuance, and treat the differences in categories as if they were all the same. In medicine, disease severity is often recorded as ordinal data.
For example, chronic kidney disease is put into 5 stages {Stage 1, Stage 2, Stage 3, Stage 4, and Stage 5}. As the Stage increases, there’s worsening kidney function, and Stage 5 means that there’s almost no kidney function left and requires dialysis.
Here, the difference between Stage 1 and 2 kidney disease is a heck of a lot smaller than the difference between Stage 4 and 5 disease.
That’s why we can’t calculate the “mean” in case of an ordinal data. Let’s move on to numeric data, which can be discrete or continuous data.
Discrete data are measured in whole numbers, meaning they cannot have decimal values. For example, if you measured the number of children in a family, you will only record “Discrete” integer values like {1, 3, 4, or 10} because you can’t have fractions of children.
On the other hand, continuous data, like height and weight, can take on any value, including ones with decimals. For example, a high-tech scale may report a weight of 70.42459310 kg.
Numeric data can also be divided into either interval data or ratio data. Interval data indicates a meaningful quantitative difference between two values and that all the values can be placed in a clear, logical order.
For example, if you measure temperatures with either a Centigrade or Fahrenheit scale, you will get interval data. The difference between 90 and 60 degrees is a measurable 30 degrees, as is the difference between 60 and 30 degrees.
Thus, we can meaningfully add and subtract the data values. Because interval data is numerical, ordered, and there are meaningful differences between values: we can calculate means, medians, and modes.
But we cannot calculate ratios with interval data. For example, stating that 90 degrees F is 3 times as hot as 30 degrees F (90 = 3 x 30) is inaccurate.
That’s because data has to be measured on a scale that includes an absolute value for 0 in order to make relational statements involving multiplication or division.
Temperature data, measured in Celsius or Fahrenheit, do not have a meaningful value for 0, that’s because 0 degrees F or C does not imply that there is no temperature.
Unlike interval data, ratio data does have a meaningful 0. Thus, we can calculate ratios via multiplication and division on top of quantifying differences through addition and subtraction.
As with interval data, we can calculate the mean, median and mode using ratio data. Attributes like age, height, weight, and test scores are very commonly measured ratio data.
Let’s say you have data on 4 students’ test scores and values are {20 40 70 80}. You can state that “80 is 10 points higher than 70”, and because you have ratio data, you can also state that the student with a 40 did twice as well as someone with a 20, and half as well as someone with an 80.
Alright, as a quick recap, data are broadly classified into categorical and numeric data. Categorical data includes nominal data, which have no logical order like blood types, and ordinal data in which order is important, like degrees of happiness.
Numeric data includes interval data where there is no meaningful zero like in the case of temperature, and ratio data which do have a meaningful zero, like in test scores.
Numeric data can also be continuous with decimal values or discrete without decimal values.
No notes for this video yet
Try adding a note below