Sample size calculator for your MD/MS thesis

Free, no sign-up, and it runs in your browser. Enter the values from an earlier published study and get the sample size rounded up, with the formula and your numbers written out for your synopsis.

Updated MD · MS · DNB

Calculate your sample size

1. What does your study’s main objective do?
2. Enter your values

From an earlier published study. If no study reports it, 50 gives the largest sample.

Precision

5 means the estimate should fall within ±5 percentage points of the true value.

For a comparison this is 1 − α, two-sided: 95% means α = 0.05.

Only when the whole population is small and known, such as every student in one college. Leave it blank otherwise.

10 adds enough participants to cover 10% who do not respond or drop out.

3. Your sample size

Sample size

–

    How the sample size is calculated

    The calculator uses the standard formula for the objective you choose, puts your values into it, and rounds the result up to the next whole participant. The expected proportion or standard deviation must come from an earlier published study that you cite in the synopsis; the precision, or the difference worth detecting, is a decision you make with your guide. Every formula here is two-sided, and the two-group formulas assume two groups of equal size.

    Estimate one proportion

    n = Z² × p × (1 − p) / d²

    p is the expected proportion and d the absolute precision. With relative precision ε, d = ε × p.

    Estimate one mean

    n = Z² × σ² / d²

    σ is the expected standard deviation and d the precision, both in the outcome’s units.

    Compare two proportions

    n per group = (Zα + Zβ)² × [p₁(1 − p₁) + p₂(1 − p₂)] / (p₁ − p₂)²

    p1 and p2 are the proportions expected in the two groups.

    Compare two means

    n per group = 2 × (Zα + Zβ)² × σ² / Δ²

    σ is the common standard deviation and Δ the smallest difference between the means that matters clinically.

    Finite population correction

    n′ = n / (1 + (n − 1) / N)

    For the one-group modes, when the whole population of size N is small and known.

    Non-response or drop-out

    n″ = n′ / (1 − r)

    r is the proportion expected not to respond or to drop out, for example 0.1 for 10%.

    Confidence levelZ (Zα, two-sided)PowerZβ
    90%1.64580%0.84
    95%1.9690%1.28
    99%2.58

    These are the rounded Z values thesis protocols conventionally write down. Each step carries the unrounded result forward, and only the final number is rounded, always up: 384.16 becomes 385, never 384. Two-group formulas give the number for each group, so the total is twice that.

    Worked examples

    Round, illustrative numbers, not taken from any real study. Each one loads into the calculator.

    Example · illustrative numbers

    Estimate one proportion

    A cross-sectional study of a condition that earlier studies put at about 30%, with 95% confidence and a precision of ±5 percentage points.

    1. n = 1.96² × 0.3 × (1 − 0.3) / 0.05² = 3.8416 × 0.21 / 0.0025 = 322.69
    2. Rounded up: 323 participants
    3. With 10% non-response: 322.69 / (1 − 0.1) = 358.55, so 359
    4. If the whole population were 1,000 people: 322.69 / (1 + (322.69 − 1) / 1000) = 244.15, so 245
    Example · illustrative numbers

    Estimate one mean

    Estimating a mean when earlier studies report a standard deviation of 10 (in mmHg, say), with 95% confidence and a precision of ±2 mmHg.

    1. n = 1.96² × 10² / 2² = 3.8416 × 100 / 4 = 96.04
    2. Rounded up: 97 participants
    Example · illustrative numbers

    Compare two proportions

    A trial in which the outcome is expected in 40% of one group and 20% of the other, with a two-sided α of 0.05 (95%) and 80% power.

    1. n per group = (1.96 + 0.84)² × [0.4 × (1 − 0.4) + 0.2 × (1 − 0.2)] / (0.4 − 0.2)² = 7.84 × 0.4 / 0.04 = 78.4
    2. Rounded up: 79 per group, 158 in total
    Example · illustrative numbers

    Compare two means

    Two groups with an expected standard deviation of 10, where a difference of 5 between the means matters clinically, with a two-sided α of 0.05 and 80% power.

    1. n per group = 2 × (1.96 + 0.84)² × 10² / 5² = 2 × 7.84 × 100 / 25 = 62.72
    2. Rounded up: 63 per group, 126 in total

    Which mode fits your study

    Size the study for its primary objective, the one question the study is built to answer. If two objectives need different formulas, calculate both and use the larger number.

    Your main objective asksTypical designMode
    What proportion of people have a condition, follow a practice or know a fact?Cross-sectional study: prevalence, KAP or prescription-audit surveyEstimate one proportion
    What is the average value of a measurement?Cross-sectional descriptive studyEstimate one mean
    Does a yes-or-no outcome differ between two groups?Randomised trial or comparative cohort with a binary outcome; case-control study, with p1 and p2 as the proportions exposed among cases and controlsCompare two proportions
    Does a measured outcome differ between two groups?Randomised trial or comparative study with a continuous outcomeCompare two means

    This calculator does not cover the designs below. They need a different formula, so have a statistician do or check the calculation:

    • paired or before-and-after measurements, matched case-control studies, or unequal group sizes
    • more than two groups, a correlation, or a regression with several variables
    • a diagnostic test’s sensitivity and specificity, time to an event, or a non-inferiority or equivalence question
    • sampling in clusters such as wards, villages or schools, which needs a design effect

    Before the number goes in your synopsis

    A starting point, not the final word

    The result is a starting point. Your protocol’s design decides the formula, and your guide and, where your college asks for one, a statistician should confirm the formula and every assumption before the synopsis goes to the IEC.

    State the formula, every value you used and the published study each value came from. A reviewer checks those assumptions more closely than the arithmetic. For a fuller explanation, read how sample size is worked out for an MD/MS thesis.

    Questions residents ask

    Which sample size formula do I use for a cross-sectional study?
    If the main objective estimates a prevalence or another proportion, use n = Z² × p × (1 − p) / d², where p is the proportion expected from an earlier published study and d is the absolute precision. With p = 50%, d = 5 percentage points and 95% confidence (Z = 1.96), n is 385. If the main objective estimates an average, use n = Z² × σ² / d² instead.
    What value of p do I use if no earlier study reports it?
    Use p = 50%, which gives the largest sample for a given precision, and say so in the synopsis. A published study in a similar population, or a small pilot study, gives a better estimate when one is available.
    Why does OpenEpi or G*Power give a slightly different number?
    They use exact Z values rather than the rounded 1.96, 0.84 and 1.28, and for two proportions OpenEpi’s formulas use a pooled proportion, with an optional continuity correction. This calculator uses the rounded values and the formulas most thesis protocols write down, so the answers can differ by a few participants. Either is acceptable if you name the formula or the tool you used.
    Is the calculator’s answer enough for my IEC submission?
    It is a starting point. Your study design decides the formula, and your guide and, where your college asks for one, a statistician should confirm the formula and every assumption before the synopsis goes to the IEC. State the formula, the values used and the study each value came from.

    Sources

    Need it done and certified?

    We review the sample size in your synopsis and the assumptions behind it, and issue the statistician review certificate many colleges ask for at IEC submission. Analysis of your data, when it is collected, is on the same page.