Which statistical test should I use for my MD/MS thesis?
The test follows from four facts about your primary outcome: its data type (a measurement, a category or a time to an event), how many groups you compare, whether those groups are independent or paired, and, for measurements, whether the data are normally distributed. Answer those four and the tables below name the test.
Four questions that pick the test
Ask them of each objective in your synopsis, starting with the primary one.
- What type of data is the outcome?Continuous: a measurement such as blood pressure, HbA1c or a scale score. Categorical: yes/no (responder, ADR present) or several categories (blood group). Ordinal: ordered categories such as NYHA class or a single Likert item. Time-to-event: time to relapse, to an ADR, or survival. Count: ADRs per patient.
- How many groups?One group against a known value, two groups, or three or more.
- Independent or paired?Different patients in each group are independent. The same patients measured twice (before and after), or matched pairs, are paired.
- Is it normally distributed?For continuous outcomes only. Normal data take the parametric test; clearly non-normal data in a small sample, and ordinal data, take the non-parametric alternative.
The decision tables
| Comparison | Normal data (parametric) | Not normal, or ordinal (non-parametric) |
|---|---|---|
| One group against a known value | One-sample t-test | Wilcoxon signed-rank test |
| Two independent groups (drug A vs drug B) | Independent-samples (Student’s) t-test | Mann-Whitney U test |
| Two paired measurements (before and after, matched pairs) | Paired t-testThe differences must be normal. | Wilcoxon signed-rank test |
| Three or more independent groups | One-way ANOVA | Kruskal-Wallis test |
| Three or more measurements on the same patients | Repeated-measures ANOVA, or a mixed model | Friedman test |
| Association between two continuous variables | Pearson correlation (r)Both variables normal. | Spearman correlation (ρ) |
| Comparison | Test |
|---|---|
| Two independent groups, comparing proportions | Chi-square test, if every expected cell count is 5 or more; Fisher’s exact test if not |
| Two paired measurements (the same patients before and after) | McNemar’s test |
| Three or more groups, or more than two categories | Chi-square test for an r×c table |
| Three or more paired measurements of a yes/no outcome | Cochran’s Q test |
| A trend across ordered groups (such as age bands) | Chi-square test for trend (Cochran-Armitage) |
| Ordinal outcome (NYHA class, pain bands, a Likert item) | Mann-Whitney U for two groups, Kruskal-Wallis for three or more, Wilcoxon signed-rank if paired |
| Question | Method |
|---|---|
| Describe survival, or time to an event, over follow-up | Kaplan-Meier curve |
| Compare time to an event between groups | Log-rank test |
| Time to an event, adjusting for age, sex and other factors | Cox proportional hazards regression (hazard ratio) |
| Continuous outcome, adjusting for other factors | Multiple linear regression |
| Yes/no outcome, adjusting for other factors | Logistic regression (odds ratio) |
| Count outcome (number of ADRs per patient) | Poisson regression; negative binomial regression if the variance is much larger than the mean |
Survival, regression and repeated-measures models have assumptions of their own to check. For these, involve a statistician.
Is my data normal?
- Test it. Use the Shapiro-Wilk test for samples under 50, and the Kolmogorov-Smirnov test with the Lilliefors correction above 50. A p-value above 0.05 means no evidence against normality.
- Look at it. Always draw a histogram and a Q-Q plot; points lying close to the straight line suggest normal data. The plot tells you more than the p-value alone.
- Check the right thing. Check each group separately for an independent comparison, and the before-minus-after differences for a paired one.
- If it is not normal. With fewer than about 30 per group, take the non-parametric column. With more than 30, t-tests and ANOVA tolerate mild non-normality (the central limit theorem). Heavily skewed lab values such as ALT, AST or viral load can often be log-transformed and analysed as normal.
Which groups differ? Post-hoc tests
A significant ANOVA or Kruskal-Wallis test says the groups are not all alike. It does not say which ones differ. Run one global test first, then a post-hoc test:
- After ANOVA: Tukey HSD by default; Dunnett when each group is compared with one control; Bonferroni when you want to be conservative, or Holm-Bonferroni, which keeps the same error control with more power.
- After Kruskal-Wallis: Dunn’s test.
Running a separate t-test or chi-square test for every pair instead inflates the chance of a false positive (type 1 error).
What to report
Descriptive statistics
Normally distributed data as mean ± SD; skewed or ordinal data as median (IQR); categories as number (%). Match the summary to the test: mean ± SD beside a t-test, median (IQR) beside a Mann-Whitney U test.
P-values with confidence intervals and effect sizes
A p-value says how surprising the difference would be by chance alone, not how big it is. Report the size of the effect with its 95% confidence interval:
| Analysis | Effect size |
|---|---|
| Two means | Mean difference with 95% CI; Cohen’s d |
| Two proportions | Risk difference, relative risk or odds ratio, with 95% CI |
| Correlation | r (Pearson) or ρ (Spearman) |
| Chi-square test | Cramér’s V, or phi for a 2×2 table |
| ANOVA | η² (eta squared) or partial η² |
| Mann-Whitney U test | r = Z/√n, or the rank-biserial correlation |
| Cox regression | Hazard ratio with 95% CI |
Large samples make trivial differences “significant”, and small samples miss important ones. The confidence interval shows both, so judge clinical relevance as well as p.
Common mistakes examiners spot
- “Appropriate statistical tests will be used.” Name the test for each objective in the synopsis.
- A t-test on a single Likert item. It is ordinal; use Mann-Whitney U. Summed scale scores can sometimes be treated as continuous.
- An unpaired test on before-and-after data. Measurements on the same patients need the paired test.
- Repeated pairwise tests instead of one global test with post-hoc comparisons.
- A p-value with no effect size or confidence interval.
- “Post-hoc power” to explain a non-significant result. Report the observed effect and its confidence interval instead.
Software residents use
- SPSS: menu-driven, and the package many departments and guides already know.
- JASP and Jamovi: point-and-click programs built on R, with tables laid out close to publication style; Jamovi also has a power and sample-size module.
- R: a scripting language. Every step is saved as code, so the analysis can be rerun exactly, at the cost of a steeper start.
- Epi Info and OpenEpi: epidemiology tools, Epi Info from the US CDC and OpenEpi in the browser, handy for 2×2 tables, proportions and sample size.
Excel suits the master chart (one row per participant, one column per variable, coded consistently), but run the tests in a statistics package and name it, with its version, in your methods.
Worked examples: from objective to test
These are illustrations of the method, not results or topics to copy.
- Objective
- To compare the fall in HbA1c at 12 weeks between adults with type 2 diabetes on metformin plus sitagliptin and those on metformin plus glimepiride.
- Four answers
- Continuous outcome (change in HbA1c); two groups; independent between groups, paired before and after within each group; normality to check.
- Test
- Check the change in each group with Shapiro-Wilk and a Q-Q plot. Between groups: independent-samples t-test if normal, Mann-Whitney U if not. Within a group, before against after: paired t-test, or Wilcoxon signed-rank.
- Report
- Mean ± SD, or median (IQR), in each group, and the mean difference between groups with its 95% CI.
- Objective
- To compare the incidence of postoperative nausea and vomiting (PONV) in the first 24 hours between patients given ondansetron and those given dexamethasone before laparoscopic cholecystectomy.
- Four answers
- Categorical outcome (PONV yes/no); two groups; independent; normality does not apply.
- Test
- Chi-square test if every expected count in the 2×2 table is 5 or more; Fisher’s exact test if any is smaller, which is likely when PONV is uncommon and the groups are small. A secondary objective comparing nausea severity on a 0–3 scale is ordinal: Mann-Whitney U.
- Report
- Number (%) with PONV in each group, and the risk difference or relative risk with its 95% CI.
Want your analysis done, and a statistician’s certificate?
Descriptive and inferential statistics for every objective, with tables, charts and a note on the test used for each, for your viva. The statistician review certificate many colleges ask for at IEC submission is also available. We report what your data shows.
See statistics support →Need a sample size first?
The test and the sample size are planned together, in the synopsis. Work out a first figure with the free calculator, then confirm the assumptions with your guide.
Open the calculator →Need help with more than the statistics? See thesis support.
Questions
Which test compares two groups?
When do I use a non-parametric test?
Chi-square or Fisher’s exact test?
Do I need a statistician for my thesis?
Sources
- Indrayan A, Malhotra RK. Medical Biostatistics. 4th ed. Boca Raton: CRC Press; 2018.
- Mahajan BK. Methods in Biostatistics for Medical Students and Research Workers. New Delhi: Jaypee Brothers.
- Whitley E, Ball J. Statistics review 6: Nonparametric methods. Crit Care. 2002;6(6):509–13. PubMed 12493072
- Bewick V, Cheek L, Ball J. Statistics review 8: Qualitative data – tests of association. Crit Care. 2004;8(1):46–53. PubMed 14975045
- Bewick V, Cheek L, Ball J. Statistics review 9: One-way analysis of variance. Crit Care. 2004;8(2):130–6. PubMed 15025774
- Ghasemi A, Zahediasl S. Normality tests for statistical analysis: a guide for non-statisticians. Int J Endocrinol Metab. 2012;10(2):486–9. PubMed 23843808