Introduction to biostatistics

Last updated: November 01, 2022

Introduction to biostatistics

Watch later

Watch later

Human herpesvirus 8 (Kaposi sarcoma)
Herpes simplex virus
Human herpesvirus 6 (Roseola)
Adenovirus
Parvovirus B19
Human papillomavirus
BK virus (Hemorrhagic cystitis)
JC virus (Progressive multifocal leukoencephalopathy)
Poliovirus
Coxsackievirus
Rhinovirus
Hepatitis A and Hepatitis E virus
Influenza virus
Mumps virus
Measles virus
Respiratory syncytial virus
Human parainfluenza viruses
Yellow fever virus
Zika virus
Hepatitis C virus
West Nile virus
Norovirus
Rotavirus
HIV (AIDS)
Rabies virus
Rubella virus
Prions (Spongiform encephalopathy)
Candida
Plasmodium species (Malaria)
Trypanosoma cruzi (Chagas disease)
Protein synthesis inhibitors: Aminoglycosides
Antimetabolites: Sulfonamides and trimethoprim
Antituberculosis medications
Miscellaneous cell wall synthesis inhibitors
Protein synthesis inhibitors: Tetracyclines
Cell wall synthesis inhibitors: Penicillins
Miscellaneous protein synthesis inhibitors
Cell wall synthesis inhibitors: Cephalosporins
DNA synthesis inhibitors: Metronidazole
DNA synthesis inhibitors: Fluoroquinolones
Mechanisms of antibiotic resistance
Integrase and entry inhibitors
Nucleoside reverse transcriptase inhibitors (NRTIs)
Protease inhibitors
Hepatitis medications
Non-nucleoside reverse transcriptase inhibitors (NNRTIs)
Neuraminidase inhibitors
Herpesvirus medications
Azoles
Echinocandins
Miscellaneous antifungal medications
Anthelmintic medications
Antimalarials
Anti-mite and louse medications
Nuclear structure
DNA structure
Transcription of DNA
Translation of mRNA
Gene regulation
Epigenetics
Amino acids and protein folding
Nucleotide metabolism
DNA replication
Lac operon
DNA damage and repair
Cell cycle
Mitosis and meiosis
DNA mutations
Lesch-Nyhan syndrome
Adenosine deaminase deficiency
Purine and pyrimidine synthesis and metabolism disorders: Pathology review
Polymerase chain reaction (PCR) and reverse-transcriptase PCR (RT-PCR)
Gel electrophoresis and genetic testing
ELISA (Enzyme-linked immunosorbent assay)
Karyotyping
DNA cloning
Fluorescence in situ hybridization
Mendelian genetics and punnett squares
Hardy-Weinberg equilibrium
Inheritance patterns
Independent assortment of genes and linkage
Evolution and natural selection
Down syndrome (Trisomy 21)
Edwards syndrome (Trisomy 18)
Patau syndrome (Trisomy 13)
Fragile X syndrome
Huntington disease
Myotonic dystrophy
Friedreich ataxia
Turner syndrome
Klinefelter syndrome
Prader-Willi syndrome
Angelman syndrome
Cri du chat syndrome
Williams syndrome
Alagille syndrome (NORD)
Achondroplasia
Polycystic kidney disease
Familial adenomatous polyposis
Familial hypercholesterolemia
Marfan syndrome
Multiple endocrine neoplasia
Neurofibromatosis
Tuberous sclerosis
von Hippel-Lindau disease
Albinism
Cystic fibrosis
Gaucher disease (NORD)
Glycogen storage disease type I
Glycogen storage disease type II (NORD)
Hemochromatosis
Mucopolysaccharide storage disease type 1 (Hurler syndrome) (NORD)
Leukodystrophy
Niemann-Pick disease types A and B (NORD)
Niemann-Pick disease type C
Phenylketonuria (NORD)
Sickle cell disease (NORD)
Tay-Sachs disease (NORD)
Alpha-thalassemia
Beta-thalassemia
Wilson disease
Alport syndrome
X-linked agammaglobulinemia
Fabry disease (NORD)
Glucose-6-phosphate dehydrogenase (G6PD) deficiency
Hemophilia
Mucopolysaccharide storage disease type 2 (Hunter syndrome) (NORD)
Muscular dystrophy
Wiskott-Aldrich syndrome
Mitochondrial myopathy
Autosomal trisomies: Pathology review
Muscular dystrophies and mitochondrial myopathies: Pathology review
Miscellaneous genetic disorders: Pathology review
Human development days 1-4
Human development days 4-7
Human development week 2
Human development week 3
Ectoderm
Mesoderm
Endoderm
Development of the placenta
Development of the fetal membranes
Development of twins
Hedgehog signaling pathway
Development of the digestive system and body cavities
Development of the umbilical cord
Development of the cardiovascular system
Fetal circulation
Development of the face and palate
Pharyngeal arches, pouches, and clefts
Development of the gastrointestinal system
Development of the teeth
Development of the tongue
Development of the axial skeleton
Development of the muscular system
Development of the renal system
Development of the reproductive system
Development of the respiratory system
Cellular structure and function
Cell membrane
Selective permeability of the cell membrane
Extracellular matrix
Cell-cell junctions
Endocytosis and exocytosis
Osmosis
Resting membrane potential
Nernst equation
Cytoskeleton and intracellular motility
Cell signaling pathways
Adrenoleukodystrophy (NORD)
Zellweger spectrum disorders (NORD)
Ehlers-Danlos syndrome
Peroxisomal disorders: Pathology review
Introduction to biostatistics
Types of data
Probability
Mean, median, and mode
Range, variance, and standard deviation
Standard error of the mean (Central limit theorem)
Normal distribution and z-scores
Paired t-test
Two-sample t-test
Hypothesis testing: One-tailed and two-tailed tests
One-way ANOVA
Two-way ANOVA
Repeated measures ANOVA
Correlation
Methods of regression analysis
Linear regression
Logistic regression
Type I and type II errors
Sensitivity and specificity
Positive and negative predictive value
Test precision and accuracy
Incidence and prevalence
Relative and absolute risk
Odds ratio
Mortality rates and case-fatality
DALY and QALY
Direct standardization
Indirect standardization
Study designs
Ecologic study
Cross sectional study
Case-control study
Cohort study
Randomized control trial
Clinical trials
Sample size
Disease causality
Selection bias
Information bias
Confounding
Interaction
Prevention
Major depressive disorder
Suicide
Bipolar and related disorders
Major depressive disorder with seasonal pattern
Generalized anxiety disorder
Social anxiety disorder
Panic disorder
Phobias
Obsessive-compulsive disorder
Body focused repetitive disorders
Post-traumatic stress disorder
Schizophrenia
Delirium
Amnesia
Dissociative disorders
Anorexia nervosa
Bulimia nervosa
Cluster A personality disorders
Cluster B personality disorders
Cluster C personality disorders
Somatic symptom disorder
Factitious disorder
Tobacco use disorder
Opioid use disorder
Cannabis use disorder
Cocaine use disorder
Alcohol use disorder
Bruxism
Insomnia
Narcolepsy (NORD)
Erectile dysfunction
Attention deficit hyperactivity disorder
Disruptive, impulse control, and conduct disorders
Learning disability
Fetal alcohol syndrome
Tourette syndrome
Autism spectrum disorder
Rett syndrome
Mood disorders: Pathology review
Amnesia, dissociative disorders and delirium: Pathology review
Personality disorders: Pathology review
Eating disorders: Pathology review
Psychological sleep disorders: Pathology review
Psychiatric emergencies: Pathology review
Drug misuse, intoxication and withdrawal: Hallucinogens: Pathology review
Malingering, factitious disorders and somatoform disorders: Pathology review
Trauma- and stress-related disorders: Pathology review
Selective serotonin reuptake inhibitors
Serotonin and norepinephrine reuptake inhibitors
Tricyclic antidepressants
Monoamine oxidase inhibitors
Atypical antidepressants
Typical antipsychotics
Atypical antipsychotics
Lithium
Nonbenzodiazepine anticonvulsants
Anticonvulsants and anxiolytics: Barbiturates
Anticonvulsants and anxiolytics: Benzodiazepines
Psychomotor stimulants

Transcript

Watch video only

Let’s say you want to figure out if people with high body mass index, or BMI, are at a higher risk of hypertension - or high blood pressure.

Let’s say that you decide to go out and find 100 people with hypertension and 100 people without hypertension and find out the BMI of each person in each group.

You might also collect other information about the individuals in each group, like how old they are, if they smoke cigarettes, or if they drink alcohol, since all of these factors can influence a person’s risk of hypertension.

All of these different pieces of information - called variables - can be put together into a single document or file, called a data set.

A data set usually includes independent variables which are thought to influence or change dependent variables.

In our example, the body mass index would be the independent variable and hypertension would be the dependent variable.

The process of collecting, organizing, and analyzing variables in a data set is called statistics, and when the data were collected from living things - like humans, aardvarks, algae, or bacteria - it’s called biostatistics, bio meaning life.

Now, there are two main types of biostatistics.

The first type is descriptive statistics, which is used to describe or summarize information about each individual variable in the data set.

Descriptive statistics can be used to find the mean - the average number calculated from a particular variable, the median - the middle number in a variable, and the mode - the number that occurs the most in the variable.

The descriptive statistics of each variable can be calculated for the whole sample - all 200 people - or in each group separately - the 100 people in the group with hypertension or the other 100 people in the group without hypertension.

For example, we might find that the mean body mass index of all people in the study is 24.5, or that the mean body mass index is 28 for the group with hypertension and 21 for the group without hypertension.

We can also use descriptive statistics to find the range, variance, or standard deviation, all of which are ways of understanding how the data are spread out or distributed for a given variable.

For example, we might find that the lowest measured body mass index in the group with hypertension is 23, and the highest is 33, so the range for body mass index in this group is 23 to 33.

Typically, descriptive statistics are reported in a graph or a table.

The second type of biostatistics is inferential, which is different from descriptive statistics in two ways.

First, inferential statistics looks at relationships between two or more variables, instead of looking at each individual variable.

For example, we could use inferential statistics to explore the relationship between body mass index and hypertension.

We could categorize body mass index into two groups - above 25, or high, and below 25, or low - and we might find that people with high body mass indices have 3 times the odds of hypertension compared to people with low body mass indices.

Typically, inferential statistics are reported by relative risks, attributable risks, odds ratios, or hazard ratios.

The goal of descriptive statistics is to describe how similar or different the study groups in a particular sample population are to one another.

For example, let’s say we use descriptive statistics to find that 72% of people in the group with hypertension are male, but only 16% of people in the group without hypertension are male.

This is an important finding because men tend to have slightly lower body mass indices than women.

As a result, having more men in the group with hypertension, means that the average body mass index in that group will be lower.

Ultimately, if the descriptive statistics find that the study groups are not very similar, we say that the study has low internal validity, and that the results found by inferential statistics may be the result of differences in the two study groups.

On the other hand, the goal of inferential statistics is to apply the results of the sample population to a target population - which is usually just the general population.

So, inferential statistics is concerned about whether or not the two study groups are similar, as well as whether or not the sample population represents the target population.

Ideally, a study should be done on a sample population of individuals that is similar to that target population in every meaningful way.

Key Takeaways

Biostatistics refers to the process of collecting, organizing, and analyzing variables collected from living things. Biostatistics involves design studies to answer specific scientific questions, and the skills necessary to properly analyze the data collected from those studies. It also involves effective communication of the results of analyses to scientists and other non-statisticians.